{
  "id": 324223,
  "title": "10th place solution",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/writeups/flg2k-10th-place-solution",
  "author_name": "",
  "post_date": "2022-05-10T15:19:01.825471Z",
  "votes": 25,
  "comment_count": 3,
  "views": 0,
  "content": "<p>First, thanks for the great competition to H&amp;M and Kaggle. I enjoyed it a lot, it's a very nice kind of challenge, as it's not heavily compute-gated and allows for a wide range of solutions.</p>\n<p>Many Kagglers already shared  their solution and since I followed the same framework as most, I will only mention some differences.</p>\n<h6>CV-LB correlation</h6>\n<ol>\n<li>Many people reported a gap between CV and LB or a \"slowdown\" when their models improved. I never experienced that, a linear fit between my CV and LB gives: f(x) = 0.793x + c with R²=98.96%. </li>\n<li>The problem is most likely caused by information leaking from the future. For example when I train a CF-model on all data up to 09-15 and later use this to make predictions for the week 09-01, I will boost my CV by implicitly using purchase information from the future.</li>\n<li>I avoided this by training a full set of models for every week. I tagged the models with the last date used to create it. That way it was impossible to leak any information.</li>\n<li>While this sounds good, it also was my \"downfall\". I simply did not have enough compute to do this for many different models and weeks. Which is why my recall methods were much less diverse than those of the top solutions and which is why I only used four weeks of training data in the end.</li>\n</ol>\n<h6>Retrieval strategy / Features</h6>\n<ol>\n<li>In addition to what many people used (rebuys, I2I, CF, Top-Sellers etc.), I created a simple sales prediction model. Based on recent sales of an article it predicted next weeks sales.</li>\n<li>The main use of this model was helping to detect products that are phasing in or out (since we don't have that data).</li>\n<li>Example: an article that is only available for a short time (special collaboration etc.) might have daily sales like: 0 0 28 431 389 105 32 11. Using recent popularity will massively overestimate next weeks sales while the prediction model learned the article is phasing out (out of stock).</li>\n<li>It was in the end one of the most used features, probably since it covers popularity and (future) availability in one.</li>\n<li>My work horse for similarity was LightFM (BPR and WARP, with and without item features). I used it for recommendation and item-to-item similarity (by comparing feature embeddings).</li>\n</ol>\n<h6>Hardware / Optimization</h6>\n<ol>\n<li>Desktop with 12C and 64GB Ram.</li>\n<li>I used feature storing liberally.</li>\n<li>To decouple training data size from Ram available, I trained my LightGBM model from HDF5.</li>\n<li>Final model was an ensemble of four LightGBM models. But the improvement was rather small, ~1% compared to the single best model.</li>\n</ol>",
  "messages": [
    {
      "id": "1783676",
      "postDate": "05/10/2022 15:19:01",
      "content": "<p>First, thanks for the great competition to H&amp;M and Kaggle. I enjoyed it a lot, it's a very nice kind of challenge, as it's not heavily compute-gated and allows for a wide range of solutions.</p>\n<p>Many Kagglers already shared  their solution and since I followed the same framework as most, I will only mention some differences.</p>\n<h6>CV-LB correlation</h6>\n<ol>\n<li>Many people reported a gap between CV and LB or a \"slowdown\" when their models improved. I never experienced that, a linear fit between my CV and LB gives: f(x) = 0.793x + c with R²=98.96%. </li>\n<li>The problem is most likely caused by information leaking from the future. For example when I train a CF-model on all data up to 09-15 and later use this to make predictions for the week 09-01, I will boost my CV by implicitly using purchase information from the future.</li>\n<li>I avoided this by training a full set of models for every week. I tagged the models with the last date used to create it. That way it was impossible to leak any information.</li>\n<li>While this sounds good, it also was my \"downfall\". I simply did not have enough compute to do this for many different models and weeks. Which is why my recall methods were much less diverse than those of the top solutions and which is why I only used four weeks of training data in the end.</li>\n</ol>\n<h6>Retrieval strategy / Features</h6>\n<ol>\n<li>In addition to what many people used (rebuys, I2I, CF, Top-Sellers etc.), I created a simple sales prediction model. Based on recent sales of an article it predicted next weeks sales.</li>\n<li>The main use of this model was helping to detect products that are phasing in or out (since we don't have that data).</li>\n<li>Example: an article that is only available for a short time (special collaboration etc.) might have daily sales like: 0 0 28 431 389 105 32 11. Using recent popularity will massively overestimate next weeks sales while the prediction model learned the article is phasing out (out of stock).</li>\n<li>It was in the end one of the most used features, probably since it covers popularity and (future) availability in one.</li>\n<li>My work horse for similarity was LightFM (BPR and WARP, with and without item features). I used it for recommendation and item-to-item similarity (by comparing feature embeddings).</li>\n</ol>\n<h6>Hardware / Optimization</h6>\n<ol>\n<li>Desktop with 12C and 64GB Ram.</li>\n<li>I used feature storing liberally.</li>\n<li>To decouple training data size from Ram available, I trained my LightGBM model from HDF5.</li>\n<li>Final model was an ensemble of four LightGBM models. But the improvement was rather small, ~1% compared to the single best model.</li>\n</ol>",
      "rawMarkdown": "First, thanks for the great competition to H&M and Kaggle. I enjoyed it a lot, it's a very nice kind of challenge, as it's not heavily compute-gated and allows for a wide range of solutions.\n\nMany Kagglers already shared  their solution and since I followed the same framework as most, I will only mention some differences.\n\n###### CV-LB correlation\n1. Many people reported a gap between CV and LB or a \"slowdown\" when their models improved. I never experienced that, a linear fit between my CV and LB gives: f(x) = 0.793x + c with R²=98.96%. \n2. The problem is most likely caused by information leaking from the future. For example when I train a CF-model on all data up to 09-15 and later use this to make predictions for the week 09-01, I will boost my CV by implicitly using purchase information from the future.\n3. I avoided this by training a full set of models for every week. I tagged the models with the last date used to create it. That way it was impossible to leak any information.\n4. While this sounds good, it also was my \"downfall\". I simply did not have enough compute to do this for many different models and weeks. Which is why my recall methods were much less diverse than those of the top solutions and which is why I only used four weeks of training data in the end.\n\n###### Retrieval strategy / Features\n1. In addition to what many people used (rebuys, I2I, CF, Top-Sellers etc.), I created a simple sales prediction model. Based on recent sales of an article it predicted next weeks sales.\n2. The main use of this model was helping to detect products that are phasing in or out (since we don't have that data).\n3. Example: an article that is only available for a short time (special collaboration etc.) might have daily sales like: 0 0 28 431 389 105 32 11. Using recent popularity will massively overestimate next weeks sales while the prediction model learned the article is phasing out (out of stock).\n4. It was in the end one of the most used features, probably since it covers popularity and (future) availability in one.\n5. My work horse for similarity was LightFM (BPR and WARP, with and without item features). I used it for recommendation and item-to-item similarity (by comparing feature embeddings).\n\n###### Hardware / Optimization\n1. Desktop with 12C and 64GB Ram.\n2. I used feature storing liberally.\n3. To decouple training data size from Ram available, I trained my LightGBM model from HDF5.\n4. Final model was an ensemble of four LightGBM models. But the improvement was rather small, ~1% compared to the single best model.",
      "votes": null
    },
    {
      "id": "1783753",
      "postDate": "05/10/2022 16:38:28",
      "content": "<p>The prediction model to try to capture out of stock items seems a very good idea. Thanks for sharing and congratulations for gold.</p>",
      "rawMarkdown": "The prediction model to try to capture out of stock items seems a very good idea. Thanks for sharing and congratulations for gold.",
      "votes": null
    },
    {
      "id": "1783948",
      "postDate": "05/10/2022 19:40:07",
      "content": "<p>Congrats for achieving that position/Gold medal and thank you for sharing your valuable solution.</p>",
      "rawMarkdown": "Congrats for achieving that position/Gold medal and thank you for sharing your valuable solution.",
      "votes": null
    },
    {
      "id": "1784103",
      "postDate": "05/10/2022 23:56:55",
      "content": "<p>Great solution! You have inspired me to try a similar technique for the next competition :)</p>",
      "rawMarkdown": "Great solution! You have inspired me to try a similar technique for the next competition :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1783753,
      "author_name": "igorkf",
      "author_url": "",
      "post_date": "05/10/2022 16:38:28",
      "content": "<p>The prediction model to try to capture out of stock items seems a very good idea. Thanks for sharing and congratulations for gold.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1783948,
      "author_name": "mpwolke",
      "author_url": "",
      "post_date": "05/10/2022 19:40:07",
      "content": "<p>Congrats for achieving that position/Gold medal and thank you for sharing your valuable solution.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1784103,
      "author_name": "",
      "author_url": "",
      "post_date": "05/10/2022 23:56:55",
      "content": "<p>Great solution! You have inspired me to try a similar technique for the next competition :)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1783676": "First, thanks for the great competition to H&M and Kaggle. I enjoyed it a lot, it's a very nice kind of challenge, as it's not heavily compute-gated and allows for a wide range of solutions.\n\nMany Kagglers already shared  their solution and since I followed the same framework as most, I will only mention some differences.\n\n###### CV-LB correlation\n1. Many people reported a gap between CV and LB or a \"slowdown\" when their models improved. I never experienced that, a linear fit between my CV and LB gives: f(x) = 0.793x + c with R²=98.96%. \n2. The problem is most likely caused by information leaking from the future. For example when I train a CF-model on all data up to 09-15 and later use this to make predictions for the week 09-01, I will boost my CV by implicitly using purchase information from the future.\n3. I avoided this by training a full set of models for every week. I tagged the models with the last date used to create it. That way it was impossible to leak any information.\n4. While this sounds good, it also was my \"downfall\". I simply did not have enough compute to do this for many different models and weeks. Which is why my recall methods were much less diverse than those of the top solutions and which is why I only used four weeks of training data in the end.\n\n###### Retrieval strategy / Features\n1. In addition to what many people used (rebuys, I2I, CF, Top-Sellers etc.), I created a simple sales prediction model. Based on recent sales of an article it predicted next weeks sales.\n2. The main use of this model was helping to detect products that are phasing in or out (since we don't have that data).\n3. Example: an article that is only available for a short time (special collaboration etc.) might have daily sales like: 0 0 28 431 389 105 32 11. Using recent popularity will massively overestimate next weeks sales while the prediction model learned the article is phasing out (out of stock).\n4. It was in the end one of the most used features, probably since it covers popularity and (future) availability in one.\n5. My work horse for similarity was LightFM (BPR and WARP, with and without item features). I used it for recommendation and item-to-item similarity (by comparing feature embeddings).\n\n###### Hardware / Optimization\n1. Desktop with 12C and 64GB Ram.\n2. I used feature storing liberally.\n3. To decouple training data size from Ram available, I trained my LightGBM model from HDF5.\n4. Final model was an ensemble of four LightGBM models. But the improvement was rather small, ~1% compared to the single best model.",
    "1783753": "The prediction model to try to capture out of stock items seems a very good idea. Thanks for sharing and congratulations for gold.",
    "1783948": "Congrats for achieving that position/Gold medal and thank you for sharing your valuable solution.",
    "1784103": "Great solution! You have inspired me to try a similar technique for the next competition :)"
  },
  "source": "meta"
}