{
  "id": 324224,
  "title": "22th solution (Moro part)",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/discussion/324224",
  "author_name": "",
  "post_date": "2022-05-10T15:29:31.343322100Z",
  "votes": 18,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Thank you to the organizers for the fun competition and everyone who participated.<br>\nAnd thank you to my teammates ( <a href=\"https://www.kaggle.com/narimatsu\" target=\"_blank\">@narimatsu</a> <a href=\"https://www.kaggle.com/haradataman\" target=\"_blank\">@haradataman</a>  <a href=\"https://www.kaggle.com/taku37\" target=\"_blank\">@taku37</a> <a href=\"https://www.kaggle.com/iwatatakuya\" target=\"_blank\">@iwatatakuya</a> )</p>\n<p>I share my solution (a part of team solution).</p>\n<h1>1. dataset/validation</h1>\n<ul>\n<li>make 4 data -&gt;  train=data1+2+3, valid=data0<br>\n[data0]  target: 9/16-9/22, variable: -9/15<br>\n[data1] target: 9/09-9/15, variable: -9/08<br>\n[data2] target: 9/02-9/08, variable: -9/01<br>\n[data3] target: 8/26-9/01, variable: -8/25</li>\n</ul>\n<h1>2. make candidates</h1>\n<ul>\n<li>recall method -&gt; rate=31%<br>\n(1) popular items<br>\n(2) recent purchase items<br>\n(3) different color items<br>\n(4) together purchase items<br>\n(5) popular items of same user attr<br>\n(6) similar items using cosine-similarity between vectors (word2vec from purchase history)</li>\n</ul>\n<h1>3. feature engineering</h1>\n<ul>\n<li>total number: 118</li>\n<li>user feature</li>\n<li>item feature </li>\n<li>transaction feature<ul>\n<li>count-encoding with decay over time (decay=1/(1+days//7))</li>\n<li>key: user, item, userxitem</li>\n<li>agg: mean, std, min, max, count, sum</li></ul></li>\n<li>flag of recall method</li>\n</ul>\n<h1>4. model</h1>\n<ul>\n<li>LightGBM (objective=lambdarank)</li>\n<li>concat data1, data2, data3 -&gt; train one model (with weight=7:2:1)</li>\n<li>local=0.03353，private=0.03045</li>\n</ul>\n<h1>I tried GNN-model (But didn't work)</h1>\n<ul>\n<li>I tried GNN(Link-prediction) first in this competion. Because I was interested in GNN. I spent most of this competition creating GNN-model.</li>\n<li>I used dgl library.  I referred to the official page.<ul>\n<li>Define hetero-graph: 2type-nodes(user, item), edge(purchase)</li>\n<li>GNN-model: input(graph) -&gt; RGCN(2-layers) &gt; User embedding / Item embedding -&gt; edge-score</li></ul></li>\n<li>I could define and train GNN-model, but score is very low,,, I gave up on this attempt.  I didn't have enough skills. So I switched to LightGBM model later in the competition.</li>\n<li>It didn't work, but I think I learned a lot (I want to think so)</li>\n</ul>\n<p><strong>If anyone can predict well with GNN, please let me know !</strong></p>\n<p>Thank you for reading.</p>",
  "messages": [
    {
      "id": "1783683",
      "postDate": "05/10/2022 15:29:31",
      "content": "<p>Thank you to the organizers for the fun competition and everyone who participated.<br>\nAnd thank you to my teammates ( <a href=\"https://www.kaggle.com/narimatsu\" target=\"_blank\">@narimatsu</a> <a href=\"https://www.kaggle.com/haradataman\" target=\"_blank\">@haradataman</a>  <a href=\"https://www.kaggle.com/taku37\" target=\"_blank\">@taku37</a> <a href=\"https://www.kaggle.com/iwatatakuya\" target=\"_blank\">@iwatatakuya</a> )</p>\n<p>I share my solution (a part of team solution).</p>\n<h1>1. dataset/validation</h1>\n<ul>\n<li>make 4 data -&gt;  train=data1+2+3, valid=data0<br>\n[data0]  target: 9/16-9/22, variable: -9/15<br>\n[data1] target: 9/09-9/15, variable: -9/08<br>\n[data2] target: 9/02-9/08, variable: -9/01<br>\n[data3] target: 8/26-9/01, variable: -8/25</li>\n</ul>\n<h1>2. make candidates</h1>\n<ul>\n<li>recall method -&gt; rate=31%<br>\n(1) popular items<br>\n(2) recent purchase items<br>\n(3) different color items<br>\n(4) together purchase items<br>\n(5) popular items of same user attr<br>\n(6) similar items using cosine-similarity between vectors (word2vec from purchase history)</li>\n</ul>\n<h1>3. feature engineering</h1>\n<ul>\n<li>total number: 118</li>\n<li>user feature</li>\n<li>item feature </li>\n<li>transaction feature<ul>\n<li>count-encoding with decay over time (decay=1/(1+days//7))</li>\n<li>key: user, item, userxitem</li>\n<li>agg: mean, std, min, max, count, sum</li></ul></li>\n<li>flag of recall method</li>\n</ul>\n<h1>4. model</h1>\n<ul>\n<li>LightGBM (objective=lambdarank)</li>\n<li>concat data1, data2, data3 -&gt; train one model (with weight=7:2:1)</li>\n<li>local=0.03353，private=0.03045</li>\n</ul>\n<h1>I tried GNN-model (But didn't work)</h1>\n<ul>\n<li>I tried GNN(Link-prediction) first in this competion. Because I was interested in GNN. I spent most of this competition creating GNN-model.</li>\n<li>I used dgl library.  I referred to the official page.<ul>\n<li>Define hetero-graph: 2type-nodes(user, item), edge(purchase)</li>\n<li>GNN-model: input(graph) -&gt; RGCN(2-layers) &gt; User embedding / Item embedding -&gt; edge-score</li></ul></li>\n<li>I could define and train GNN-model, but score is very low,,, I gave up on this attempt.  I didn't have enough skills. So I switched to LightGBM model later in the competition.</li>\n<li>It didn't work, but I think I learned a lot (I want to think so)</li>\n</ul>\n<p><strong>If anyone can predict well with GNN, please let me know !</strong></p>\n<p>Thank you for reading.</p>",
      "rawMarkdown": "Thank you to the organizers for the fun competition and everyone who participated.\nAnd thank you to my teammates ( @narimatsu @haradataman  @taku37 @iwatatakuya )\n\nI share my solution (a part of team solution).\n\n# 1. dataset/validation\n- make 4 data ->  train=data1+2+3, valid=data0\n  [data0]  target: 9/16-9/22, variable: -9/15\n  [data1] target: 9/09-9/15, variable: -9/08\n  [data2] target: 9/02-9/08, variable: -9/01\n  [data3] target: 8/26-9/01, variable: -8/25\n\n# 2. make candidates\n- recall method -> rate=31%\n  (1) popular items\n  (2) recent purchase items\n  (3) different color items\n  (4) together purchase items\n  (5) popular items of same user attr\n  (6) similar items using cosine-similarity between vectors (word2vec from purchase history)\n\n# 3. feature engineering\n- total number: 118\n- user feature\n- item feature \n- transaction feature\n   - count-encoding with decay over time (decay=1/(1+days//7))\n   - key: user, item, userxitem\n   - agg: mean, std, min, max, count, sum\n- flag of recall method\n\n# 4. model\n- LightGBM (objective=lambdarank)\n- concat data1, data2, data3 -> train one model (with weight=7:2:1)\n- local=0.03353，private=0.03045\n\n# I tried GNN-model (But didn't work)\n- I tried GNN(Link-prediction) first in this competion. Because I was interested in GNN. I spent most of this competition creating GNN-model.\n- I used dgl library.  I referred to the official page.\n    - Define hetero-graph: 2type-nodes(user, item), edge(purchase)\n    - GNN-model: input(graph) -> RGCN(2-layers) > User embedding / Item embedding -> edge-score\n- I could define and train GNN-model, but score is very low,,, I gave up on this attempt.  I didn't have enough skills. So I switched to LightGBM model later in the competition.\n- It didn't work, but I think I learned a lot (I want to think so)\n\n**If anyone can predict well with GNN, please let me know !**\n\nThank you for reading.",
      "votes": null
    },
    {
      "id": "1783784",
      "postDate": "05/10/2022 17:06:38",
      "content": "<p>Congratulations for nice work.</p>\n<blockquote>\n  <p>concat data1, data2, data3 -&gt; train one model (with weight=7:2:1)</p>\n</blockquote>\n<p>For \"weight=7:2:1\" you mean sample weights? For data1 samples give weight=7, for data2 samples give weight=2 and so on?</p>",
      "rawMarkdown": "Congratulations for nice work.\n> concat data1, data2, data3 -> train one model (with weight=7:2:1)\n\nFor \"weight=7:2:1\" you mean sample weights? For data1 samples give weight=7, for data2 samples give weight=2 and so on?",
      "votes": null
    },
    {
      "id": "1784023",
      "postDate": "05/10/2022 21:14:06",
      "content": "<p>Thank you for question.</p>\n<p>yes.<br>\nI used lgb.Dataset().  So I set sample weight (w_train) in weight parameter.</p>\n<pre><code>lgb_train = lgb.Dataset(data=x_train, label=y_train, weight=w_train, group=group_train)\nlgb_valid = lgb.Dataset(data=x_valid, label=y_valid, weight=None, group=group_valid)\nmodel = lgb.train(params,\n                  lgb_train,\n                  valid_sets=[lgb_train, lgb_valid],\n                  valid_names=[\"train\", \"valid\"],\n                  early_stopping_rounds=200,\n                  verbose_eval=20,\n                 )\n</code></pre>",
      "rawMarkdown": "Thank you for question.\n\nyes.\nI used lgb.Dataset().  So I set sample weight (w_train) in weight parameter.\n\n```\nlgb_train = lgb.Dataset(data=x_train, label=y_train, weight=w_train, group=group_train)\nlgb_valid = lgb.Dataset(data=x_valid, label=y_valid, weight=None, group=group_valid)\nmodel = lgb.train(params,\n                  lgb_train,\n                  valid_sets=[lgb_train, lgb_valid],\n                  valid_names=[\"train\", \"valid\"],\n                  early_stopping_rounds=200,\n                  verbose_eval=20,\n                 )\n```",
      "votes": null
    },
    {
      "id": "1784056",
      "postDate": "05/10/2022 22:25:19",
      "content": "<p>Thanks for answering.</p>",
      "rawMarkdown": "Thanks for answering.",
      "votes": null
    },
    {
      "id": "1784104",
      "postDate": "05/10/2022 23:58:12",
      "content": "<p>+1 - Making GNNs work for recommendation systems is an unsolved problem and it's fantastic to see someone giving it a go. I think you're right that you learn a lot even if it doesn't work out in the end.</p>",
      "rawMarkdown": "1 - Making GNNs work for recommendation systems is an unsolved problem and it's fantastic to see someone giving it a go. I think you're right that you learn a lot even if it doesn't work out in the end.",
      "votes": null
    },
    {
      "id": "1784742",
      "postDate": "05/11/2022 12:35:54",
      "content": "<p>Thank you for comment. I'd like to continue to try it.</p>",
      "rawMarkdown": "Thank you for comment. I'd like to continue to try it.",
      "votes": null
    },
    {
      "id": "1809703",
      "postDate": "06/03/2022 00:56:44",
      "content": "<p>Hey, post competition I was also trying to solve this using GNN as link prediction problem…I am getting MAP@12 ~ 3*10^-6 which is basically 0. <a href=\"https://www.kaggle.com/mrugankakarte/h-m-personalization-using-gnn\" target=\"_blank\">https://www.kaggle.com/mrugankakarte/h-m-personalization-using-gnn</a> Here's how I tried, let me know your thoughts on this. </p>",
      "rawMarkdown": "Hey, post competition I was also trying to solve this using GNN as link prediction problem...I am getting MAP@12 ~ 3*10^-6 which is basically 0. https://www.kaggle.com/mrugankakarte/h-m-personalization-using-gnn Here's how I tried, let me know your thoughts on this.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1783784,
      "author_name": "igorkf",
      "author_url": "",
      "post_date": "05/10/2022 17:06:38",
      "content": "<p>Congratulations for nice work.</p>\n<blockquote>\n  <p>concat data1, data2, data3 -&gt; train one model (with weight=7:2:1)</p>\n</blockquote>\n<p>For \"weight=7:2:1\" you mean sample weights? For data1 samples give weight=7, for data2 samples give weight=2 and so on?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1784023,
          "author_name": "moromoromoro",
          "author_url": "",
          "post_date": "05/10/2022 21:14:06",
          "content": "<p>Thank you for question.</p>\n<p>yes.<br>\nI used lgb.Dataset().  So I set sample weight (w_train) in weight parameter.</p>\n<pre><code>lgb_train = lgb.Dataset(data=x_train, label=y_train, weight=w_train, group=group_train)\nlgb_valid = lgb.Dataset(data=x_valid, label=y_valid, weight=None, group=group_valid)\nmodel = lgb.train(params,\n                  lgb_train,\n                  valid_sets=[lgb_train, lgb_valid],\n                  valid_names=[\"train\", \"valid\"],\n                  early_stopping_rounds=200,\n                  verbose_eval=20,\n                 )\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1784056,
          "author_name": "igorkf",
          "author_url": "",
          "post_date": "05/10/2022 22:25:19",
          "content": "<p>Thanks for answering.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1784104,
      "author_name": "",
      "author_url": "",
      "post_date": "05/10/2022 23:58:12",
      "content": "<p>+1 - Making GNNs work for recommendation systems is an unsolved problem and it's fantastic to see someone giving it a go. I think you're right that you learn a lot even if it doesn't work out in the end.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1784742,
          "author_name": "moromoromoro",
          "author_url": "",
          "post_date": "05/11/2022 12:35:54",
          "content": "<p>Thank you for comment. I'd like to continue to try it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1809703,
      "author_name": "mrugankakarte",
      "author_url": "",
      "post_date": "06/03/2022 00:56:44",
      "content": "<p>Hey, post competition I was also trying to solve this using GNN as link prediction problem…I am getting MAP@12 ~ 3*10^-6 which is basically 0. <a href=\"https://www.kaggle.com/mrugankakarte/h-m-personalization-using-gnn\" target=\"_blank\">https://www.kaggle.com/mrugankakarte/h-m-personalization-using-gnn</a> Here's how I tried, let me know your thoughts on this. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1783683": "Thank you to the organizers for the fun competition and everyone who participated.\nAnd thank you to my teammates ( @narimatsu @haradataman  @taku37 @iwatatakuya )\n\nI share my solution (a part of team solution).\n\n# 1. dataset/validation\n- make 4 data ->  train=data1+2+3, valid=data0\n  [data0]  target: 9/16-9/22, variable: -9/15\n  [data1] target: 9/09-9/15, variable: -9/08\n  [data2] target: 9/02-9/08, variable: -9/01\n  [data3] target: 8/26-9/01, variable: -8/25\n\n# 2. make candidates\n- recall method -> rate=31%\n  (1) popular items\n  (2) recent purchase items\n  (3) different color items\n  (4) together purchase items\n  (5) popular items of same user attr\n  (6) similar items using cosine-similarity between vectors (word2vec from purchase history)\n\n# 3. feature engineering\n- total number: 118\n- user feature\n- item feature \n- transaction feature\n   - count-encoding with decay over time (decay=1/(1+days//7))\n   - key: user, item, userxitem\n   - agg: mean, std, min, max, count, sum\n- flag of recall method\n\n# 4. model\n- LightGBM (objective=lambdarank)\n- concat data1, data2, data3 -> train one model (with weight=7:2:1)\n- local=0.03353，private=0.03045\n\n# I tried GNN-model (But didn't work)\n- I tried GNN(Link-prediction) first in this competion. Because I was interested in GNN. I spent most of this competition creating GNN-model.\n- I used dgl library.  I referred to the official page.\n    - Define hetero-graph: 2type-nodes(user, item), edge(purchase)\n    - GNN-model: input(graph) -> RGCN(2-layers) > User embedding / Item embedding -> edge-score\n- I could define and train GNN-model, but score is very low,,, I gave up on this attempt.  I didn't have enough skills. So I switched to LightGBM model later in the competition.\n- It didn't work, but I think I learned a lot (I want to think so)\n\n**If anyone can predict well with GNN, please let me know !**\n\nThank you for reading.",
    "1783784": "Congratulations for nice work.\n> concat data1, data2, data3 -> train one model (with weight=7:2:1)\n\nFor \"weight=7:2:1\" you mean sample weights? For data1 samples give weight=7, for data2 samples give weight=2 and so on?",
    "1784023": "Thank you for question.\n\nyes.\nI used lgb.Dataset().  So I set sample weight (w_train) in weight parameter.\n\n```\nlgb_train = lgb.Dataset(data=x_train, label=y_train, weight=w_train, group=group_train)\nlgb_valid = lgb.Dataset(data=x_valid, label=y_valid, weight=None, group=group_valid)\nmodel = lgb.train(params,\n                  lgb_train,\n                  valid_sets=[lgb_train, lgb_valid],\n                  valid_names=[\"train\", \"valid\"],\n                  early_stopping_rounds=200,\n                  verbose_eval=20,\n                 )\n```",
    "1784056": "Thanks for answering.",
    "1784104": "1 - Making GNNs work for recommendation systems is an unsolved problem and it's fantastic to see someone giving it a go. I think you're right that you learn a lot even if it doesn't work out in the end.",
    "1784742": "Thank you for comment. I'd like to continue to try it.",
    "1809703": "Hey, post competition I was also trying to solve this using GNN as link prediction problem...I am getting MAP@12 ~ 3*10^-6 which is basically 0. https://www.kaggle.com/mrugankakarte/h-m-personalization-using-gnn Here's how I tried, let me know your thoughts on this."
  },
  "source": "meta"
}