{
  "id": 324205,
  "title": "46th place solution",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/writeups/nadare-46th-place-solution",
  "author_name": "",
  "post_date": "2022-06-06T14:49:40.980Z",
  "votes": 14,
  "comment_count": 6,
  "views": 0,
  "content": "<h1>46th place solution (Item2Vec/Transformer)</h1>\n<h1>Introduction</h1>\n<p>Thank you to H &amp; M, kaggle, and my rivals for the exciting competition.<br>\nThis post is a brief summary of my solution.<br>\nSince it is English written by google translate, I think that there are some parts that are not connected, but if you have any questions, please feel free to ask.</p>\n<h1>Overview</h1>\n<p>This time, I designed the model using the 2-tower model using NN.<br>\nThe 2-tower model provides high performance recommendations by using two parts: the retrieval part, which narrows down candidates with a lightweight model, and the ranking part, which makes predictions with a highly accurate model.<br>\nSpecifically, the flow of prediction / learning is as follows.</p>\n<ul>\n<li><p>1st stage (retrieval part) Public: 0.02137 Private: 0.02296<br>\nItem2Vec is used to narrow down the ranking of the number of articles appearing around the predicted date. (1)<br>\nExtract items with a history in the past that have an interaction with another user most recently on a rule basis. (2)</p></li>\n<li><p>2nd stage (ranking part) Public: 0.2779 Private: 0.02846<br>\nRerank (1) and (2) using Transformer.</p></li>\n</ul>\n<h1>Solution</h1>\n<p>The model will be explained in the following three parts.</p>\n<ul>\n<li>Embedding part common to retrieval / ranking</li>\n<li>Retrieval part</li>\n<li>ranking part</li>\n</ul>\n<h2>Embedding part</h2>\n<p>In the Embedding part, the pre-calculated image / natural language model embedding and the category embedding initialized with word2vec of gensim are mixed with <a href=\"https://arxiv.org/abs/2008.13535\" target=\"_blank\">DCNV2</a>.<br>\nEach embedding was obtained as follows.</p>\n<ul>\n<li><p>image</p>\n<ul>\n<li>swin_large_patch4_window12_384_in22k + TruncatedSVD (1024 dim)</li>\n<li>EfficientNetV2L + TruncatedSVD (1024 dim)</li></ul></li>\n<li><p>Natural language</p>\n<ul>\n<li>ELECTRA LARGE (Enter by concatenating detail_desc and category name)</li>\n<li>Sentence-T5 st5-11b (Enter by concatenating detail_desc and category name)</li>\n<li>BERT Tokenizer + Tf-Idf + PCA (128 dim)</li></ul></li>\n<li><p>Category</p></li>\n<li><p>Word2vec by gensim</p></li>\n<li><p>Table data</p>\n<ul>\n<li>price</li>\n<li>Total quantity (same products are combined into one record for each day)</li>\n<li>How many interactions for that user</li>\n<li>Price statistics when compared to other days</li></ul></li>\n</ul>\n<p>Since these features were introduced before the implementation of other methods with greater effects, we have not verified the individual effects, but I feel that they were not very effective.</p>\n<h2>Retrieval part</h2>\n<p>In the Retrieval part, I narrowed down by the search vector that took the weighted average of the embedding of the article and the cosine similarity of the embedding of the product.</p>\n<h3>Model</h3>\n<h4>Target embedding</h4>\n<p>I put the embedding of the article into several fully connected layers and converted it.</p>\n<h4>Query embedding</h4>\n<p>I used a weighted average weighted by time etc. for the embedding of the article.<br>\nAs for how to add weight, the weight was calculated by <code>(1/2) ^ relu (edge_feature)</code> with the feature such as elapsed time as edge_feature, and the weighted average was calculated.</p>\n<p>I also considered Transformer using self attention, but it was meaningless in this part.</p>\n<h3>Loss</h3>\n<p>We calculated <a href=\"https://arxiv.org/abs/2004.11362\" target=\"_blank\">SupervisedContrastiveLoss</a> with 1 as the one with interaction and 0 as the negative example without it.<br>\nAt the same time, I used <a href=\"https://arxiv.org/abs/1905.00292\" target=\"_blank\">AdaCos</a>. The margin parameter is 0.</p>\n<p>When the margin is greater than 0, all cosine similarity diverges negatively.<br>\nAlso, I was monitoring the cosine similarity of the regular case, but in the end it was about 0.3.<br>\nIn terms of ranking, the correct example out of 4096 samples was about 600th, and I felt that this task was difficult.</p>\n<h3>Subsampling</h3>\n<p>Batch segmented users by age 10 years, ensured that targets of the same age and same day were in the batch, and trained in chronological order.<br>\nDuring training, I sampled up to 4096 articles from the top 10,000 appearances of articles in the last week of the group segmented by 10 years old.<br>\nSampling used <a href=\"https://stackoverflow.com/questions/43310075/sampling-without-replacement-from-a-given-non-uniform-distribution-in-tensorflow\" target=\"_blank\">Gumbel-max trick</a>.</p>\n<p>At the time of inference, we calculated the cosine similarity for 10,000 items that appeared frequently in the last 90 days, which had somebody's purchase history in the last 7 days. (I also tried an approximate neighborhood search like ScaNN, but the time impact was negligible)</p>\n<h2>Ranking part</h2>\n<p>In the ranking, rank learning was performed using Transformer.</p>\n<h3>Model</h3>\n<p>We further used Transformer with the recommended symmetric article as Query and the user's browsing history as Key &amp; Value. Again, there was no need for a self-attention type Encoder or a Decoder with more than one layer.<br>\nEmbedding obtained by this and the result of the intermediate layer calculated so far were put into DCNV2 and binary classification was performed.</p>\n<h3>Loss</h3>\n<p>loss is based on <a href=\"https://www.tensorflow.org/ranking/api_docs/python/tfr/keras/losses/ApproxNDCGLoss\" target=\"_blank\">ApproxNDCGLoss</a> and uses a unique loss with a penalty for each rank. ..<br>\nWe compared ApproxNDCGLoss (original), rank difference ** 1, and rank difference ** 2, and the rank difference ** 1 was the best result.</p>\n<h3>Subsampling</h3>\n<p>At the time of training, a maximum of 128 negative cases were created by Gumbel-max trick with the ratio of the cube of the cosine similarity as the probability. This gave good feedback to the model as it could take a reasonably weaker negative example than taking the top as it is.<br>\nAt the time of inference, we predicted the top 128 of the Retrieveval part and the past browsing history.</p>\n<h1>Machine &amp; Experiments</h1>\n<p>I used a core i9-12900 CPU, 64GB RAM, and an RTX 3090 GPU.<br>\nOf that, CPU and RAM were used about half, learning using all data was 1 epoch 9 hours, and the highest score was 2 epoch.</p>\n<p>As an aside, the experiment number on the notebook was 159 at the end, and I feel that I spent a lot of effort.</p>\n<p>I will update it from time to time. Feel free to ask questions!</p>\n<h1>UPDATE</h1>\n<p>I published a more detailed solution on the technical blog of my company. The explanation for those who did not participate in the competition and the part that I could not write in the discussion due to my lack of English ability are written. If you can read Japanese, please read it.<br>\n<a href=\"https://future-architect.github.io/articles/20220602b/\" target=\"_blank\">H&amp;M Personalized Fashion Recommendations 参加記 (46th/2952)</a></p>\n<p><a href=\"https://future-architect.github.io/images/20220602b/H&amp;M_46th_solution_overview.drawio.png\" target=\"_blank\">image of solution overview</a></p>",
  "messages": [
    {
      "id": "1783591",
      "postDate": "05/10/2022 14:01:16",
      "content": "<h1>46th place solution (Item2Vec/Transformer)</h1>\n<h1>Introduction</h1>\n<p>Thank you to H &amp; M, kaggle, and my rivals for the exciting competition.<br>\nThis post is a brief summary of my solution.<br>\nSince it is English written by google translate, I think that there are some parts that are not connected, but if you have any questions, please feel free to ask.</p>\n<h1>Overview</h1>\n<p>This time, I designed the model using the 2-tower model using NN.<br>\nThe 2-tower model provides high performance recommendations by using two parts: the retrieval part, which narrows down candidates with a lightweight model, and the ranking part, which makes predictions with a highly accurate model.<br>\nSpecifically, the flow of prediction / learning is as follows.</p>\n<ul>\n<li><p>1st stage (retrieval part) Public: 0.02137 Private: 0.02296<br>\nItem2Vec is used to narrow down the ranking of the number of articles appearing around the predicted date. (1)<br>\nExtract items with a history in the past that have an interaction with another user most recently on a rule basis. (2)</p></li>\n<li><p>2nd stage (ranking part) Public: 0.2779 Private: 0.02846<br>\nRerank (1) and (2) using Transformer.</p></li>\n</ul>\n<h1>Solution</h1>\n<p>The model will be explained in the following three parts.</p>\n<ul>\n<li>Embedding part common to retrieval / ranking</li>\n<li>Retrieval part</li>\n<li>ranking part</li>\n</ul>\n<h2>Embedding part</h2>\n<p>In the Embedding part, the pre-calculated image / natural language model embedding and the category embedding initialized with word2vec of gensim are mixed with <a href=\"https://arxiv.org/abs/2008.13535\" target=\"_blank\">DCNV2</a>.<br>\nEach embedding was obtained as follows.</p>\n<ul>\n<li><p>image</p>\n<ul>\n<li>swin_large_patch4_window12_384_in22k + TruncatedSVD (1024 dim)</li>\n<li>EfficientNetV2L + TruncatedSVD (1024 dim)</li></ul></li>\n<li><p>Natural language</p>\n<ul>\n<li>ELECTRA LARGE (Enter by concatenating detail_desc and category name)</li>\n<li>Sentence-T5 st5-11b (Enter by concatenating detail_desc and category name)</li>\n<li>BERT Tokenizer + Tf-Idf + PCA (128 dim)</li></ul></li>\n<li><p>Category</p></li>\n<li><p>Word2vec by gensim</p></li>\n<li><p>Table data</p>\n<ul>\n<li>price</li>\n<li>Total quantity (same products are combined into one record for each day)</li>\n<li>How many interactions for that user</li>\n<li>Price statistics when compared to other days</li></ul></li>\n</ul>\n<p>Since these features were introduced before the implementation of other methods with greater effects, we have not verified the individual effects, but I feel that they were not very effective.</p>\n<h2>Retrieval part</h2>\n<p>In the Retrieval part, I narrowed down by the search vector that took the weighted average of the embedding of the article and the cosine similarity of the embedding of the product.</p>\n<h3>Model</h3>\n<h4>Target embedding</h4>\n<p>I put the embedding of the article into several fully connected layers and converted it.</p>\n<h4>Query embedding</h4>\n<p>I used a weighted average weighted by time etc. for the embedding of the article.<br>\nAs for how to add weight, the weight was calculated by <code>(1/2) ^ relu (edge_feature)</code> with the feature such as elapsed time as edge_feature, and the weighted average was calculated.</p>\n<p>I also considered Transformer using self attention, but it was meaningless in this part.</p>\n<h3>Loss</h3>\n<p>We calculated <a href=\"https://arxiv.org/abs/2004.11362\" target=\"_blank\">SupervisedContrastiveLoss</a> with 1 as the one with interaction and 0 as the negative example without it.<br>\nAt the same time, I used <a href=\"https://arxiv.org/abs/1905.00292\" target=\"_blank\">AdaCos</a>. The margin parameter is 0.</p>\n<p>When the margin is greater than 0, all cosine similarity diverges negatively.<br>\nAlso, I was monitoring the cosine similarity of the regular case, but in the end it was about 0.3.<br>\nIn terms of ranking, the correct example out of 4096 samples was about 600th, and I felt that this task was difficult.</p>\n<h3>Subsampling</h3>\n<p>Batch segmented users by age 10 years, ensured that targets of the same age and same day were in the batch, and trained in chronological order.<br>\nDuring training, I sampled up to 4096 articles from the top 10,000 appearances of articles in the last week of the group segmented by 10 years old.<br>\nSampling used <a href=\"https://stackoverflow.com/questions/43310075/sampling-without-replacement-from-a-given-non-uniform-distribution-in-tensorflow\" target=\"_blank\">Gumbel-max trick</a>.</p>\n<p>At the time of inference, we calculated the cosine similarity for 10,000 items that appeared frequently in the last 90 days, which had somebody's purchase history in the last 7 days. (I also tried an approximate neighborhood search like ScaNN, but the time impact was negligible)</p>\n<h2>Ranking part</h2>\n<p>In the ranking, rank learning was performed using Transformer.</p>\n<h3>Model</h3>\n<p>We further used Transformer with the recommended symmetric article as Query and the user's browsing history as Key &amp; Value. Again, there was no need for a self-attention type Encoder or a Decoder with more than one layer.<br>\nEmbedding obtained by this and the result of the intermediate layer calculated so far were put into DCNV2 and binary classification was performed.</p>\n<h3>Loss</h3>\n<p>loss is based on <a href=\"https://www.tensorflow.org/ranking/api_docs/python/tfr/keras/losses/ApproxNDCGLoss\" target=\"_blank\">ApproxNDCGLoss</a> and uses a unique loss with a penalty for each rank. ..<br>\nWe compared ApproxNDCGLoss (original), rank difference ** 1, and rank difference ** 2, and the rank difference ** 1 was the best result.</p>\n<h3>Subsampling</h3>\n<p>At the time of training, a maximum of 128 negative cases were created by Gumbel-max trick with the ratio of the cube of the cosine similarity as the probability. This gave good feedback to the model as it could take a reasonably weaker negative example than taking the top as it is.<br>\nAt the time of inference, we predicted the top 128 of the Retrieveval part and the past browsing history.</p>\n<h1>Machine &amp; Experiments</h1>\n<p>I used a core i9-12900 CPU, 64GB RAM, and an RTX 3090 GPU.<br>\nOf that, CPU and RAM were used about half, learning using all data was 1 epoch 9 hours, and the highest score was 2 epoch.</p>\n<p>As an aside, the experiment number on the notebook was 159 at the end, and I feel that I spent a lot of effort.</p>\n<p>I will update it from time to time. Feel free to ask questions!</p>\n<h1>UPDATE</h1>\n<p>I published a more detailed solution on the technical blog of my company. The explanation for those who did not participate in the competition and the part that I could not write in the discussion due to my lack of English ability are written. If you can read Japanese, please read it.<br>\n<a href=\"https://future-architect.github.io/articles/20220602b/\" target=\"_blank\">H&amp;M Personalized Fashion Recommendations 参加記 (46th/2952)</a></p>\n<p><a href=\"https://future-architect.github.io/images/20220602b/H&amp;M_46th_solution_overview.drawio.png\" target=\"_blank\">image of solution overview</a></p>",
      "rawMarkdown": "46th place solution (Item2Vec/Transformer)\n=================================\n\n#Introduction\nThank you to H & M, kaggle, and my rivals for the exciting competition.\nThis post is a brief summary of my solution.\nSince it is English written by google translate, I think that there are some parts that are not connected, but if you have any questions, please feel free to ask.\n\n#Overview\nThis time, I designed the model using the 2-tower model using NN.\nThe 2-tower model provides high performance recommendations by using two parts: the retrieval part, which narrows down candidates with a lightweight model, and the ranking part, which makes predictions with a highly accurate model.\nSpecifically, the flow of prediction / learning is as follows.\n\n* 1st stage (retrieval part) Public: 0.02137 Private: 0.02296\n  Item2Vec is used to narrow down the ranking of the number of articles appearing around the predicted date. (1)\nExtract items with a history in the past that have an interaction with another user most recently on a rule basis. (2)\n\n* 2nd stage (ranking part) Public: 0.2779 Private: 0.02846\n  Rerank (1) and (2) using Transformer.\n\n#Solution\nThe model will be explained in the following three parts.\n\n* Embedding part common to retrieval / ranking\n* Retrieval part\n* ranking part\n\n## Embedding part\nIn the Embedding part, the pre-calculated image / natural language model embedding and the category embedding initialized with word2vec of gensim are mixed with [DCNV2] (https://arxiv.org/abs/2008.13535).\nEach embedding was obtained as follows.\n\n* image\n  * swin_large_patch4_window12_384_in22k + TruncatedSVD (1024 dim)\n  * EfficientNetV2L + TruncatedSVD (1024 dim)\n\n* Natural language\n  * ELECTRA LARGE (Enter by concatenating detail_desc and category name)\n  * Sentence-T5 st5-11b (Enter by concatenating detail_desc and category name)\n  * BERT Tokenizer + Tf-Idf + PCA (128 dim)\n\n* Category\n* Word2vec by gensim\n\n* Table data\n  * price\n  * Total quantity (same products are combined into one record for each day)\n  * How many interactions for that user\n  * Price statistics when compared to other days\n\nSince these features were introduced before the implementation of other methods with greater effects, we have not verified the individual effects, but I feel that they were not very effective.\n\n## Retrieval part\nIn the Retrieval part, I narrowed down by the search vector that took the weighted average of the embedding of the article and the cosine similarity of the embedding of the product.\n\n### Model\n\n#### Target embedding\nI put the embedding of the article into several fully connected layers and converted it.\n\n#### Query embedding\nI used a weighted average weighted by time etc. for the embedding of the article.\nAs for how to add weight, the weight was calculated by `(1/2) ^ relu (edge_feature)` with the feature such as elapsed time as edge_feature, and the weighted average was calculated.\n\nI also considered Transformer using self attention, but it was meaningless in this part.\n\n### Loss\nWe calculated [SupervisedContrastiveLoss] (https://arxiv.org/abs/2004.11362) with 1 as the one with interaction and 0 as the negative example without it.\nAt the same time, I used [AdaCos] (https://arxiv.org/abs/1905.00292). The margin parameter is 0.\n\nWhen the margin is greater than 0, all cosine similarity diverges negatively.\nAlso, I was monitoring the cosine similarity of the regular case, but in the end it was about 0.3.\nIn terms of ranking, the correct example out of 4096 samples was about 600th, and I felt that this task was difficult.\n\n### Subsampling\nBatch segmented users by age 10 years, ensured that targets of the same age and same day were in the batch, and trained in chronological order.\nDuring training, I sampled up to 4096 articles from the top 10,000 appearances of articles in the last week of the group segmented by 10 years old.\nSampling used [Gumbel-max trick] (https://stackoverflow.com/questions/43310075/sampling-without-replacement-from-a-given-non-uniform-distribution-in-tensorflow).\n\nAt the time of inference, we calculated the cosine similarity for 10,000 items that appeared frequently in the last 90 days, which had somebody's purchase history in the last 7 days. (I also tried an approximate neighborhood search like ScaNN, but the time impact was negligible)\n\n\n## Ranking part\nIn the ranking, rank learning was performed using Transformer.\n\n### Model\nWe further used Transformer with the recommended symmetric article as Query and the user's browsing history as Key & Value. Again, there was no need for a self-attention type Encoder or a Decoder with more than one layer.\nEmbedding obtained by this and the result of the intermediate layer calculated so far were put into DCNV2 and binary classification was performed.\n\n### Loss\nloss is based on [ApproxNDCGLoss] (https://www.tensorflow.org/ranking/api_docs/python/tfr/keras/losses/ApproxNDCGLoss) and uses a unique loss with a penalty for each rank. ..\nWe compared ApproxNDCGLoss (original), rank difference ** 1, and rank difference ** 2, and the rank difference ** 1 was the best result.\n\n### Subsampling\nAt the time of training, a maximum of 128 negative cases were created by Gumbel-max trick with the ratio of the cube of the cosine similarity as the probability. This gave good feedback to the model as it could take a reasonably weaker negative example than taking the top as it is.\nAt the time of inference, we predicted the top 128 of the Retrieveval part and the past browsing history.\n\n#Machine & Experiments\nI used a core i9-12900 CPU, 64GB RAM, and an RTX 3090 GPU.\nOf that, CPU and RAM were used about half, learning using all data was 1 epoch 9 hours, and the highest score was 2 epoch.\n\nAs an aside, the experiment number on the notebook was 159 at the end, and I feel that I spent a lot of effort.\n\n\nI will update it from time to time. Feel free to ask questions!\n\n# UPDATE\nI published a more detailed solution on the technical blog of my company. The explanation for those who did not participate in the competition and the part that I could not write in the discussion due to my lack of English ability are written. If you can read Japanese, please read it.\n[H&M Personalized Fashion Recommendations 参加記 (46th/2952)](https://future-architect.github.io/articles/20220602b/)\n\n[image of solution overview](https://future-architect.github.io/images/20220602b/H&M_46th_solution_overview.drawio.png)",
      "votes": null
    },
    {
      "id": "1783594",
      "postDate": "05/10/2022 14:02:34",
      "content": "<p>Japanese version (same as English)</p>\n<p>46th place solution (Item2Vec/Transformer)</p>\n<h1>はじめに</h1>\n<p>H&amp;Mとkaggle、そしてライバルの皆さん、楽しいコンペをありがとうございました。<br>\nこの投稿は自身の解法の簡単なまとめです。<br>\ngoogle翻訳で書いた英語なのでつたない部分もあると思いますが、わからない点があれば気軽に質問してください。</p>\n<h1>Overview</h1>\n<p>私は今回、NNを用いた2-towerモデルを用いてモデルを設計しました。<br>\n2-towerモデルは軽量のモデルで候補の絞り込みを行うretrievalパートと、精度の高いモデルで予測を行うrankingパートの2つを用いることで高パフォーマンスの推薦を提供します。<br>\n具体的には予測・学習の流れは以下のようになります。</p>\n<ul>\n<li><p>1st stage (retrieval part) Public: 0.02137 Private: 0.02296<br>\n予測日周辺でのarticle出現数のランキング上位に対しItem2Vecで絞り込みを行う。(1)<br>\n　直近でほかのユーザのinteractionがある、過去に履歴のあるアイテムをルールベースで取り出す。(2)</p></li>\n<li><p>2nd stage (ranking part) Public: 0.2779 Private: 0.02846<br>\n(1)、(2)に対しTransformerを用いてリランキングを行う。</p></li>\n</ul>\n<h1>Solution</h1>\n<p>モデルについては以下の三つに分けて説明をします。</p>\n<ul>\n<li>retrieval/rankingで共通のEmbeddingパート</li>\n<li>Retrievalパート</li>\n<li>rankingパート</li>\n</ul>\n<h2>Embedding part</h2>\n<p>Embeddingパートでは事前計算した画像・自然言語モデルのembeddingと、gensimのword2vecで初期化したカテゴリembeddingを<a href=\"https://arxiv.org/abs/2008.13535\" target=\"_blank\">DCNV2</a>でミックスしました。<br>\n各embeddingは以下のように取得しました。</p>\n<ul>\n<li><p>画像</p>\n<ul>\n<li>swin_large_patch4_window12_384_in22k + TruncatedSVD (1024 dim)</li>\n<li>EfficientNetV2L + TruncatedSVD(1024 dim)</li></ul></li>\n<li><p>自然言語</p>\n<ul>\n<li>ELECTRA LARGE (detail_descとカテゴリ名を連結して入力)</li>\n<li>Sentence-T5 st5-11b (detail_descとカテゴリ名を連結して入力)</li>\n<li>BERT Tokenizer + Tf-Idf + PCA(128 dim)</li></ul></li>\n<li><p>カテゴリ</p>\n<ul>\n<li>gensimによるword2vec</li></ul></li>\n<li><p>テーブルデータ</p>\n<ul>\n<li>price</li>\n<li>合計個数(同じ商品は日ごとに一つのレコードにまとめた)</li>\n<li>そのユーザーにとって何回目のinteractionか</li>\n<li>ほかの日と比較した際の値段の統計値</li></ul></li>\n</ul>\n<p>これらの特徴はほかのより大きな効果を持つ手法の実装前に導入したので個々の効果は検証していませんが、体感としてあまり効果はなかったように感じています。</p>\n<h2>Retrieval part</h2>\n<p>Retrievalパートではarticleのembeddingの加重平均をとった検索ベクトルと商品のembeddingのコサイン類似度により絞り込みを行いました。</p>\n<h3>Model</h3>\n<h4>Target embedding</h4>\n<p>Articleのembeddingを何層かの全結合層に入れ変換しました。</p>\n<h4>Query embedding</h4>\n<p>Articleのembeddingについて時間等で重みづけした加重平均を使いました。<br>\n重みのつけ方は経過時間等の特徴をedge_featureとして<code>(1/2) ^ relu(edge_feature)</code>で重みを計算し加重平均を計算しました。</p>\n<p>self attentionを用いたTransformerも検討したのですが、このパートでは無意味でした。</p>\n<h3>Loss</h3>\n<p>interactionのあったものを1、そうでない負例を0として、<a href=\"https://arxiv.org/abs/2004.11362\" target=\"_blank\">SupervisedContrastiveLoss</a>を計算しました。<br>\nまた、同時に<a href=\"https://arxiv.org/abs/1905.00292\" target=\"_blank\">AdaCos</a>を用いました。マージンパラメータは0です。</p>\n<p>マージンは0より大きくするとコサイン類似度がすべて負に発散しました。<br>\nまた、正例のコサイン類似度をモニターしていましたが、最終的には0.3程度になりました。<br>\n順位にすると4096個のサンプル中正例はおよそ600位になり、今回のタスクが難しかったことを感じました。</p>\n<h3>Subsampling</h3>\n<p>バッチは、ユーザーを10歳毎にセグメントし、同じ年代かつ同じ日のターゲットがバッチ内に入るようにし、時系列にそって訓練しました。<br>\n訓練時は10歳ごとにセグメントしたグループの直近一週間でのarticleの出現上位10000位から最大4096個をサンプリングしました。<br>\nサンプリングは<a href=\"https://stackoverflow.com/questions/43310075/sampling-without-replacement-from-a-given-non-uniform-distribution-in-tensorflow\" target=\"_blank\">Gumbel-max trick</a>を用いました。</p>\n<p>推論時は直近7日間で誰かしらの購入履歴がある、直近90日で良く出現した10000個のアイテムに対して、コサイン類似度を総当たりで計算しました。(ScaNNのような近似近傍探索も試しましたが、時間的な影響はごくわずかでした)</p>\n<h2>Ranking part</h2>\n<p>ランキングではTransformerを用いランク学習を行いました。</p>\n<h3>Model</h3>\n<p>推薦対称のarticleをQuery、ユーザーの閲覧履歴をKey &amp; ValueとしたTransformerを一層用いました。ここでもself-attention型のEncoderや、二層以上のDecoderは不要でした。<br>\nこれによって得たEmbeddingと、それまでに計算した中間層の結果をDCNV2に入れて、二値分類を行いました。</p>\n<h3>Loss</h3>\n<p>lossは<a href=\"https://www.tensorflow.org/ranking/api_docs/python/tfr/keras/losses/ApproxNDCGLoss\" target=\"_blank\">ApproxNDCGLoss</a>を元に、順位ごとのペナルティを順位差にした独自のlossを用いました。<br>\nApproxNDCGLoss(original)、順位差の一乗、順位差の二乗を比較しましたが、順位差の一乗が最も良い結果でした。</p>\n<h3>Subsampling</h3>\n<p>訓練時はコサイン類似度の3乗の比を確率としてGumbel-max trickで負例を最大128個作成しました。これは上位をそのままとってくるよりも程よく弱い負例をとることができモデルに良いフィードバックを与えました。<br>\n推論時はRetrievalパートの上位128件と過去の閲覧履歴について予測を行いました。</p>\n<h1>Machine &amp; Experiments</h1>\n<p>私はcorei9-12900のCPU、64GBのRAM、RTX3090のGPUを用いました。<br>\nそのうちCPUとRAMは半分程度の使用量で、全データを用いた学習は1epoch9時間、最高スコアは2epoch目で出ました。</p>\n<p>余談ですが、notebookの実験番号は最後で159になり、非常に労力を費やしたなと感じています。</p>\n<p>随時更新します。気軽に質問してください！</p>",
      "rawMarkdown": "Japanese version (same as English)\n\n46th place solution (Item2Vec/Transformer)\n\n# はじめに\nH&Mとkaggle、そしてライバルの皆さん、楽しいコンペをありがとうございました。\nこの投稿は自身の解法の簡単なまとめです。\ngoogle翻訳で書いた英語なのでつたない部分もあると思いますが、わからない点があれば気軽に質問してください。\n\n# Overview\n私は今回、NNを用いた2-towerモデルを用いてモデルを設計しました。\n2-towerモデルは軽量のモデルで候補の絞り込みを行うretrievalパートと、精度の高いモデルで予測を行うrankingパートの2つを用いることで高パフォーマンスの推薦を提供します。\n具体的には予測・学習の流れは以下のようになります。\n\n* 1st stage (retrieval part) Public: 0.02137 Private: 0.02296\n  予測日周辺でのarticle出現数のランキング上位に対しItem2Vecで絞り込みを行う。(1)\n　直近でほかのユーザのinteractionがある、過去に履歴のあるアイテムをルールベースで取り出す。(2)\n\n* 2nd stage (ranking part) Public: 0.2779 Private: 0.02846\n  (1)、(2)に対しTransformerを用いてリランキングを行う。\n\n# Solution\nモデルについては以下の三つに分けて説明をします。\n\n* retrieval/rankingで共通のEmbeddingパート\n* Retrievalパート\n* rankingパート\n\n## Embedding part\nEmbeddingパートでは事前計算した画像・自然言語モデルのembeddingと、gensimのword2vecで初期化したカテゴリembeddingを[DCNV2](https://arxiv.org/abs/2008.13535)でミックスしました。\n各embeddingは以下のように取得しました。\n\n* 画像\n  * swin_large_patch4_window12_384_in22k + TruncatedSVD (1024 dim)\n  * EfficientNetV2L + TruncatedSVD(1024 dim)\n\n* 自然言語\n  * ELECTRA LARGE (detail_descとカテゴリ名を連結して入力)\n  * Sentence-T5 st5-11b (detail_descとカテゴリ名を連結して入力)\n  * BERT Tokenizer + Tf-Idf + PCA(128 dim)\n\n* カテゴリ\n  * gensimによるword2vec\n\n* テーブルデータ\n  * price\n  * 合計個数(同じ商品は日ごとに一つのレコードにまとめた)\n  * そのユーザーにとって何回目のinteractionか\n  * ほかの日と比較した際の値段の統計値\n\nこれらの特徴はほかのより大きな効果を持つ手法の実装前に導入したので個々の効果は検証していませんが、体感としてあまり効果はなかったように感じています。\n\n## Retrieval part\nRetrievalパートではarticleのembeddingの加重平均をとった検索ベクトルと商品のembeddingのコサイン類似度により絞り込みを行いました。\n\n### Model\n\n#### Target embedding\nArticleのembeddingを何層かの全結合層に入れ変換しました。\n\n#### Query embedding\nArticleのembeddingについて時間等で重みづけした加重平均を使いました。\n重みのつけ方は経過時間等の特徴をedge_featureとして`(1/2) ^ relu(edge_feature)`で重みを計算し加重平均を計算しました。\n\nself attentionを用いたTransformerも検討したのですが、このパートでは無意味でした。\n\n### Loss\ninteractionのあったものを1、そうでない負例を0として、[SupervisedContrastiveLoss](https://arxiv.org/abs/2004.11362)を計算しました。\nまた、同時に[AdaCos](https://arxiv.org/abs/1905.00292)を用いました。マージンパラメータは0です。\n\nマージンは0より大きくするとコサイン類似度がすべて負に発散しました。\nまた、正例のコサイン類似度をモニターしていましたが、最終的には0.3程度になりました。\n順位にすると4096個のサンプル中正例はおよそ600位になり、今回のタスクが難しかったことを感じました。\n\n### Subsampling\nバッチは、ユーザーを10歳毎にセグメントし、同じ年代かつ同じ日のターゲットがバッチ内に入るようにし、時系列にそって訓練しました。\n訓練時は10歳ごとにセグメントしたグループの直近一週間でのarticleの出現上位10000位から最大4096個をサンプリングしました。\nサンプリングは[Gumbel-max trick](https://stackoverflow.com/questions/43310075/sampling-without-replacement-from-a-given-non-uniform-distribution-in-tensorflow)を用いました。\n\n推論時は直近7日間で誰かしらの購入履歴がある、直近90日で良く出現した10000個のアイテムに対して、コサイン類似度を総当たりで計算しました。(ScaNNのような近似近傍探索も試しましたが、時間的な影響はごくわずかでした)\n\n\n## Ranking part\nランキングではTransformerを用いランク学習を行いました。\n\n### Model\n推薦対称のarticleをQuery、ユーザーの閲覧履歴をKey & ValueとしたTransformerを一層用いました。ここでもself-attention型のEncoderや、二層以上のDecoderは不要でした。\nこれによって得たEmbeddingと、それまでに計算した中間層の結果をDCNV2に入れて、二値分類を行いました。\n\n### Loss\nlossは[ApproxNDCGLoss](https://www.tensorflow.org/ranking/api_docs/python/tfr/keras/losses/ApproxNDCGLoss)を元に、順位ごとのペナルティを順位差にした独自のlossを用いました。\nApproxNDCGLoss(original)、順位差の一乗、順位差の二乗を比較しましたが、順位差の一乗が最も良い結果でした。\n\n### Subsampling\n訓練時はコサイン類似度の3乗の比を確率としてGumbel-max trickで負例を最大128個作成しました。これは上位をそのままとってくるよりも程よく弱い負例をとることができモデルに良いフィードバックを与えました。\n推論時はRetrievalパートの上位128件と過去の閲覧履歴について予測を行いました。\n\n# Machine & Experiments\n私はcorei9-12900のCPU、64GBのRAM、RTX3090のGPUを用いました。\nそのうちCPUとRAMは半分程度の使用量で、全データを用いた学習は1epoch9時間、最高スコアは2epoch目で出ました。\n\n余談ですが、notebookの実験番号は最後で159になり、非常に労力を費やしたなと感じています。\n\n\n随時更新します。気軽に質問してください！",
      "votes": null
    },
    {
      "id": "1784099",
      "postDate": "05/10/2022 23:49:22",
      "content": "<p>This is a great summary of a very effective solution! Thank you for sharing your process and results.</p>",
      "rawMarkdown": "This is a great summary of a very effective solution! Thank you for sharing your process and results.",
      "votes": null
    },
    {
      "id": "1784115",
      "postDate": "05/11/2022 00:14:10",
      "content": "<p><a href=\"https://www.kaggle.com/nadare\" target=\"_blank\">@nadare</a> congratulations! thank you for the detailed write up</p>",
      "rawMarkdown": "nadare congratulations! thank you for the detailed write up",
      "votes": null
    },
    {
      "id": "1810979",
      "postDate": "06/04/2022 06:09:17",
      "content": "<p>Can you share the solution code? Many of us would like to learn the details.</p>\n<p>Kaggle notebook or github repo or anyway is fine. </p>",
      "rawMarkdown": "Can you share the solution code? Many of us would like to learn the details.\n\nKaggle notebook or github repo or anyway is fine.",
      "votes": null
    },
    {
      "id": "1813100",
      "postDate": "06/06/2022 14:42:56",
      "content": "<p>Thanks for your comment.<br>\nIt's difficult to publish everything because the code isn't managed properly and only I can understand it.<br>\nHowever, if there is a part that you are interested in, cut out that part and publish it (if possible).</p>",
      "rawMarkdown": "Thanks for your comment.\nIt's difficult to publish everything because the code isn't managed properly and only I can understand it.\nHowever, if there is a part that you are interested in, cut out that part and publish it (if possible).",
      "votes": null
    },
    {
      "id": "2636998",
      "postDate": "02/05/2024 12:27:42",
      "content": "<p>I don't image DataFrame after you selected items each users.</p>\n<ul>\n<li>vertical holding(?)  </li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>user</th>\n<th>item</th>\n<th>feature…</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>A</td>\n<td>item1</td>\n<td>…</td>\n</tr>\n<tr>\n<td>A</td>\n<td>item2</td>\n<td>…</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li>horizontal holding(?). </li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>user</th>\n<th>item candidate1</th>\n<th>item candidate2</th>\n<th>….</th>\n<th>feature…</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>A</td>\n<td>item1</td>\n<td>item2</td>\n<td>….</td>\n<td></td>\n</tr>\n<tr>\n<td>B</td>\n<td>item2</td>\n<td>…</td>\n<td>….</td>\n<td></td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "I don't image DataFrame after you selected items each users.\n- vertical holding(?)  \n\n|user  | item |feature...|\n| --- | --- | --- |\n| A | item1 |...|\n| A | item2 |...|\n- horizontal holding(?). \n\n|user  | item candidate1 |item candidate2 |....|feature...|\n| --- | --- | --- | --- | --- |\n| A | item1 |item2|....|\n| B | item2 |...|....|",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1783594,
      "author_name": "nadare",
      "author_url": "",
      "post_date": "05/10/2022 14:02:34",
      "content": "<p>Japanese version (same as English)</p>\n<p>46th place solution (Item2Vec/Transformer)</p>\n<h1>はじめに</h1>\n<p>H&amp;Mとkaggle、そしてライバルの皆さん、楽しいコンペをありがとうございました。<br>\nこの投稿は自身の解法の簡単なまとめです。<br>\ngoogle翻訳で書いた英語なのでつたない部分もあると思いますが、わからない点があれば気軽に質問してください。</p>\n<h1>Overview</h1>\n<p>私は今回、NNを用いた2-towerモデルを用いてモデルを設計しました。<br>\n2-towerモデルは軽量のモデルで候補の絞り込みを行うretrievalパートと、精度の高いモデルで予測を行うrankingパートの2つを用いることで高パフォーマンスの推薦を提供します。<br>\n具体的には予測・学習の流れは以下のようになります。</p>\n<ul>\n<li><p>1st stage (retrieval part) Public: 0.02137 Private: 0.02296<br>\n予測日周辺でのarticle出現数のランキング上位に対しItem2Vecで絞り込みを行う。(1)<br>\n　直近でほかのユーザのinteractionがある、過去に履歴のあるアイテムをルールベースで取り出す。(2)</p></li>\n<li><p>2nd stage (ranking part) Public: 0.2779 Private: 0.02846<br>\n(1)、(2)に対しTransformerを用いてリランキングを行う。</p></li>\n</ul>\n<h1>Solution</h1>\n<p>モデルについては以下の三つに分けて説明をします。</p>\n<ul>\n<li>retrieval/rankingで共通のEmbeddingパート</li>\n<li>Retrievalパート</li>\n<li>rankingパート</li>\n</ul>\n<h2>Embedding part</h2>\n<p>Embeddingパートでは事前計算した画像・自然言語モデルのembeddingと、gensimのword2vecで初期化したカテゴリembeddingを<a href=\"https://arxiv.org/abs/2008.13535\" target=\"_blank\">DCNV2</a>でミックスしました。<br>\n各embeddingは以下のように取得しました。</p>\n<ul>\n<li><p>画像</p>\n<ul>\n<li>swin_large_patch4_window12_384_in22k + TruncatedSVD (1024 dim)</li>\n<li>EfficientNetV2L + TruncatedSVD(1024 dim)</li></ul></li>\n<li><p>自然言語</p>\n<ul>\n<li>ELECTRA LARGE (detail_descとカテゴリ名を連結して入力)</li>\n<li>Sentence-T5 st5-11b (detail_descとカテゴリ名を連結して入力)</li>\n<li>BERT Tokenizer + Tf-Idf + PCA(128 dim)</li></ul></li>\n<li><p>カテゴリ</p>\n<ul>\n<li>gensimによるword2vec</li></ul></li>\n<li><p>テーブルデータ</p>\n<ul>\n<li>price</li>\n<li>合計個数(同じ商品は日ごとに一つのレコードにまとめた)</li>\n<li>そのユーザーにとって何回目のinteractionか</li>\n<li>ほかの日と比較した際の値段の統計値</li></ul></li>\n</ul>\n<p>これらの特徴はほかのより大きな効果を持つ手法の実装前に導入したので個々の効果は検証していませんが、体感としてあまり効果はなかったように感じています。</p>\n<h2>Retrieval part</h2>\n<p>Retrievalパートではarticleのembeddingの加重平均をとった検索ベクトルと商品のembeddingのコサイン類似度により絞り込みを行いました。</p>\n<h3>Model</h3>\n<h4>Target embedding</h4>\n<p>Articleのembeddingを何層かの全結合層に入れ変換しました。</p>\n<h4>Query embedding</h4>\n<p>Articleのembeddingについて時間等で重みづけした加重平均を使いました。<br>\n重みのつけ方は経過時間等の特徴をedge_featureとして<code>(1/2) ^ relu(edge_feature)</code>で重みを計算し加重平均を計算しました。</p>\n<p>self attentionを用いたTransformerも検討したのですが、このパートでは無意味でした。</p>\n<h3>Loss</h3>\n<p>interactionのあったものを1、そうでない負例を0として、<a href=\"https://arxiv.org/abs/2004.11362\" target=\"_blank\">SupervisedContrastiveLoss</a>を計算しました。<br>\nまた、同時に<a href=\"https://arxiv.org/abs/1905.00292\" target=\"_blank\">AdaCos</a>を用いました。マージンパラメータは0です。</p>\n<p>マージンは0より大きくするとコサイン類似度がすべて負に発散しました。<br>\nまた、正例のコサイン類似度をモニターしていましたが、最終的には0.3程度になりました。<br>\n順位にすると4096個のサンプル中正例はおよそ600位になり、今回のタスクが難しかったことを感じました。</p>\n<h3>Subsampling</h3>\n<p>バッチは、ユーザーを10歳毎にセグメントし、同じ年代かつ同じ日のターゲットがバッチ内に入るようにし、時系列にそって訓練しました。<br>\n訓練時は10歳ごとにセグメントしたグループの直近一週間でのarticleの出現上位10000位から最大4096個をサンプリングしました。<br>\nサンプリングは<a href=\"https://stackoverflow.com/questions/43310075/sampling-without-replacement-from-a-given-non-uniform-distribution-in-tensorflow\" target=\"_blank\">Gumbel-max trick</a>を用いました。</p>\n<p>推論時は直近7日間で誰かしらの購入履歴がある、直近90日で良く出現した10000個のアイテムに対して、コサイン類似度を総当たりで計算しました。(ScaNNのような近似近傍探索も試しましたが、時間的な影響はごくわずかでした)</p>\n<h2>Ranking part</h2>\n<p>ランキングではTransformerを用いランク学習を行いました。</p>\n<h3>Model</h3>\n<p>推薦対称のarticleをQuery、ユーザーの閲覧履歴をKey &amp; ValueとしたTransformerを一層用いました。ここでもself-attention型のEncoderや、二層以上のDecoderは不要でした。<br>\nこれによって得たEmbeddingと、それまでに計算した中間層の結果をDCNV2に入れて、二値分類を行いました。</p>\n<h3>Loss</h3>\n<p>lossは<a href=\"https://www.tensorflow.org/ranking/api_docs/python/tfr/keras/losses/ApproxNDCGLoss\" target=\"_blank\">ApproxNDCGLoss</a>を元に、順位ごとのペナルティを順位差にした独自のlossを用いました。<br>\nApproxNDCGLoss(original)、順位差の一乗、順位差の二乗を比較しましたが、順位差の一乗が最も良い結果でした。</p>\n<h3>Subsampling</h3>\n<p>訓練時はコサイン類似度の3乗の比を確率としてGumbel-max trickで負例を最大128個作成しました。これは上位をそのままとってくるよりも程よく弱い負例をとることができモデルに良いフィードバックを与えました。<br>\n推論時はRetrievalパートの上位128件と過去の閲覧履歴について予測を行いました。</p>\n<h1>Machine &amp; Experiments</h1>\n<p>私はcorei9-12900のCPU、64GBのRAM、RTX3090のGPUを用いました。<br>\nそのうちCPUとRAMは半分程度の使用量で、全データを用いた学習は1epoch9時間、最高スコアは2epoch目で出ました。</p>\n<p>余談ですが、notebookの実験番号は最後で159になり、非常に労力を費やしたなと感じています。</p>\n<p>随時更新します。気軽に質問してください！</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1784099,
      "author_name": "",
      "author_url": "",
      "post_date": "05/10/2022 23:49:22",
      "content": "<p>This is a great summary of a very effective solution! Thank you for sharing your process and results.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1784115,
      "author_name": "lachlangillian",
      "author_url": "",
      "post_date": "05/11/2022 00:14:10",
      "content": "<p><a href=\"https://www.kaggle.com/nadare\" target=\"_blank\">@nadare</a> congratulations! thank you for the detailed write up</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1810979,
      "author_name": "jupyternotebook",
      "author_url": "",
      "post_date": "06/04/2022 06:09:17",
      "content": "<p>Can you share the solution code? Many of us would like to learn the details.</p>\n<p>Kaggle notebook or github repo or anyway is fine. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1813100,
          "author_name": "nadare",
          "author_url": "",
          "post_date": "06/06/2022 14:42:56",
          "content": "<p>Thanks for your comment.<br>\nIt's difficult to publish everything because the code isn't managed properly and only I can understand it.<br>\nHowever, if there is a part that you are interested in, cut out that part and publish it (if possible).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2636998,
      "author_name": "yyyymmdd1031",
      "author_url": "",
      "post_date": "02/05/2024 12:27:42",
      "content": "<p>I don't image DataFrame after you selected items each users.</p>\n<ul>\n<li>vertical holding(?)  </li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>user</th>\n<th>item</th>\n<th>feature…</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>A</td>\n<td>item1</td>\n<td>…</td>\n</tr>\n<tr>\n<td>A</td>\n<td>item2</td>\n<td>…</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li>horizontal holding(?). </li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>user</th>\n<th>item candidate1</th>\n<th>item candidate2</th>\n<th>….</th>\n<th>feature…</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>A</td>\n<td>item1</td>\n<td>item2</td>\n<td>….</td>\n<td></td>\n</tr>\n<tr>\n<td>B</td>\n<td>item2</td>\n<td>…</td>\n<td>….</td>\n<td></td>\n</tr>\n</tbody>\n</table>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1783591": "46th place solution (Item2Vec/Transformer)\n=================================\n\n#Introduction\nThank you to H & M, kaggle, and my rivals for the exciting competition.\nThis post is a brief summary of my solution.\nSince it is English written by google translate, I think that there are some parts that are not connected, but if you have any questions, please feel free to ask.\n\n#Overview\nThis time, I designed the model using the 2-tower model using NN.\nThe 2-tower model provides high performance recommendations by using two parts: the retrieval part, which narrows down candidates with a lightweight model, and the ranking part, which makes predictions with a highly accurate model.\nSpecifically, the flow of prediction / learning is as follows.\n\n* 1st stage (retrieval part) Public: 0.02137 Private: 0.02296\n  Item2Vec is used to narrow down the ranking of the number of articles appearing around the predicted date. (1)\nExtract items with a history in the past that have an interaction with another user most recently on a rule basis. (2)\n\n* 2nd stage (ranking part) Public: 0.2779 Private: 0.02846\n  Rerank (1) and (2) using Transformer.\n\n#Solution\nThe model will be explained in the following three parts.\n\n* Embedding part common to retrieval / ranking\n* Retrieval part\n* ranking part\n\n## Embedding part\nIn the Embedding part, the pre-calculated image / natural language model embedding and the category embedding initialized with word2vec of gensim are mixed with [DCNV2] (https://arxiv.org/abs/2008.13535).\nEach embedding was obtained as follows.\n\n* image\n  * swin_large_patch4_window12_384_in22k + TruncatedSVD (1024 dim)\n  * EfficientNetV2L + TruncatedSVD (1024 dim)\n\n* Natural language\n  * ELECTRA LARGE (Enter by concatenating detail_desc and category name)\n  * Sentence-T5 st5-11b (Enter by concatenating detail_desc and category name)\n  * BERT Tokenizer + Tf-Idf + PCA (128 dim)\n\n* Category\n* Word2vec by gensim\n\n* Table data\n  * price\n  * Total quantity (same products are combined into one record for each day)\n  * How many interactions for that user\n  * Price statistics when compared to other days\n\nSince these features were introduced before the implementation of other methods with greater effects, we have not verified the individual effects, but I feel that they were not very effective.\n\n## Retrieval part\nIn the Retrieval part, I narrowed down by the search vector that took the weighted average of the embedding of the article and the cosine similarity of the embedding of the product.\n\n### Model\n\n#### Target embedding\nI put the embedding of the article into several fully connected layers and converted it.\n\n#### Query embedding\nI used a weighted average weighted by time etc. for the embedding of the article.\nAs for how to add weight, the weight was calculated by `(1/2) ^ relu (edge_feature)` with the feature such as elapsed time as edge_feature, and the weighted average was calculated.\n\nI also considered Transformer using self attention, but it was meaningless in this part.\n\n### Loss\nWe calculated [SupervisedContrastiveLoss] (https://arxiv.org/abs/2004.11362) with 1 as the one with interaction and 0 as the negative example without it.\nAt the same time, I used [AdaCos] (https://arxiv.org/abs/1905.00292). The margin parameter is 0.\n\nWhen the margin is greater than 0, all cosine similarity diverges negatively.\nAlso, I was monitoring the cosine similarity of the regular case, but in the end it was about 0.3.\nIn terms of ranking, the correct example out of 4096 samples was about 600th, and I felt that this task was difficult.\n\n### Subsampling\nBatch segmented users by age 10 years, ensured that targets of the same age and same day were in the batch, and trained in chronological order.\nDuring training, I sampled up to 4096 articles from the top 10,000 appearances of articles in the last week of the group segmented by 10 years old.\nSampling used [Gumbel-max trick] (https://stackoverflow.com/questions/43310075/sampling-without-replacement-from-a-given-non-uniform-distribution-in-tensorflow).\n\nAt the time of inference, we calculated the cosine similarity for 10,000 items that appeared frequently in the last 90 days, which had somebody's purchase history in the last 7 days. (I also tried an approximate neighborhood search like ScaNN, but the time impact was negligible)\n\n\n## Ranking part\nIn the ranking, rank learning was performed using Transformer.\n\n### Model\nWe further used Transformer with the recommended symmetric article as Query and the user's browsing history as Key & Value. Again, there was no need for a self-attention type Encoder or a Decoder with more than one layer.\nEmbedding obtained by this and the result of the intermediate layer calculated so far were put into DCNV2 and binary classification was performed.\n\n### Loss\nloss is based on [ApproxNDCGLoss] (https://www.tensorflow.org/ranking/api_docs/python/tfr/keras/losses/ApproxNDCGLoss) and uses a unique loss with a penalty for each rank. ..\nWe compared ApproxNDCGLoss (original), rank difference ** 1, and rank difference ** 2, and the rank difference ** 1 was the best result.\n\n### Subsampling\nAt the time of training, a maximum of 128 negative cases were created by Gumbel-max trick with the ratio of the cube of the cosine similarity as the probability. This gave good feedback to the model as it could take a reasonably weaker negative example than taking the top as it is.\nAt the time of inference, we predicted the top 128 of the Retrieveval part and the past browsing history.\n\n#Machine & Experiments\nI used a core i9-12900 CPU, 64GB RAM, and an RTX 3090 GPU.\nOf that, CPU and RAM were used about half, learning using all data was 1 epoch 9 hours, and the highest score was 2 epoch.\n\nAs an aside, the experiment number on the notebook was 159 at the end, and I feel that I spent a lot of effort.\n\n\nI will update it from time to time. Feel free to ask questions!\n\n# UPDATE\nI published a more detailed solution on the technical blog of my company. The explanation for those who did not participate in the competition and the part that I could not write in the discussion due to my lack of English ability are written. If you can read Japanese, please read it.\n[H&M Personalized Fashion Recommendations 参加記 (46th/2952)](https://future-architect.github.io/articles/20220602b/)\n\n[image of solution overview](https://future-architect.github.io/images/20220602b/H&M_46th_solution_overview.drawio.png)",
    "1783594": "Japanese version (same as English)\n\n46th place solution (Item2Vec/Transformer)\n\n# はじめに\nH&Mとkaggle、そしてライバルの皆さん、楽しいコンペをありがとうございました。\nこの投稿は自身の解法の簡単なまとめです。\ngoogle翻訳で書いた英語なのでつたない部分もあると思いますが、わからない点があれば気軽に質問してください。\n\n# Overview\n私は今回、NNを用いた2-towerモデルを用いてモデルを設計しました。\n2-towerモデルは軽量のモデルで候補の絞り込みを行うretrievalパートと、精度の高いモデルで予測を行うrankingパートの2つを用いることで高パフォーマンスの推薦を提供します。\n具体的には予測・学習の流れは以下のようになります。\n\n* 1st stage (retrieval part) Public: 0.02137 Private: 0.02296\n  予測日周辺でのarticle出現数のランキング上位に対しItem2Vecで絞り込みを行う。(1)\n　直近でほかのユーザのinteractionがある、過去に履歴のあるアイテムをルールベースで取り出す。(2)\n\n* 2nd stage (ranking part) Public: 0.2779 Private: 0.02846\n  (1)、(2)に対しTransformerを用いてリランキングを行う。\n\n# Solution\nモデルについては以下の三つに分けて説明をします。\n\n* retrieval/rankingで共通のEmbeddingパート\n* Retrievalパート\n* rankingパート\n\n## Embedding part\nEmbeddingパートでは事前計算した画像・自然言語モデルのembeddingと、gensimのword2vecで初期化したカテゴリembeddingを[DCNV2](https://arxiv.org/abs/2008.13535)でミックスしました。\n各embeddingは以下のように取得しました。\n\n* 画像\n  * swin_large_patch4_window12_384_in22k + TruncatedSVD (1024 dim)\n  * EfficientNetV2L + TruncatedSVD(1024 dim)\n\n* 自然言語\n  * ELECTRA LARGE (detail_descとカテゴリ名を連結して入力)\n  * Sentence-T5 st5-11b (detail_descとカテゴリ名を連結して入力)\n  * BERT Tokenizer + Tf-Idf + PCA(128 dim)\n\n* カテゴリ\n  * gensimによるword2vec\n\n* テーブルデータ\n  * price\n  * 合計個数(同じ商品は日ごとに一つのレコードにまとめた)\n  * そのユーザーにとって何回目のinteractionか\n  * ほかの日と比較した際の値段の統計値\n\nこれらの特徴はほかのより大きな効果を持つ手法の実装前に導入したので個々の効果は検証していませんが、体感としてあまり効果はなかったように感じています。\n\n## Retrieval part\nRetrievalパートではarticleのembeddingの加重平均をとった検索ベクトルと商品のembeddingのコサイン類似度により絞り込みを行いました。\n\n### Model\n\n#### Target embedding\nArticleのembeddingを何層かの全結合層に入れ変換しました。\n\n#### Query embedding\nArticleのembeddingについて時間等で重みづけした加重平均を使いました。\n重みのつけ方は経過時間等の特徴をedge_featureとして`(1/2) ^ relu(edge_feature)`で重みを計算し加重平均を計算しました。\n\nself attentionを用いたTransformerも検討したのですが、このパートでは無意味でした。\n\n### Loss\ninteractionのあったものを1、そうでない負例を0として、[SupervisedContrastiveLoss](https://arxiv.org/abs/2004.11362)を計算しました。\nまた、同時に[AdaCos](https://arxiv.org/abs/1905.00292)を用いました。マージンパラメータは0です。\n\nマージンは0より大きくするとコサイン類似度がすべて負に発散しました。\nまた、正例のコサイン類似度をモニターしていましたが、最終的には0.3程度になりました。\n順位にすると4096個のサンプル中正例はおよそ600位になり、今回のタスクが難しかったことを感じました。\n\n### Subsampling\nバッチは、ユーザーを10歳毎にセグメントし、同じ年代かつ同じ日のターゲットがバッチ内に入るようにし、時系列にそって訓練しました。\n訓練時は10歳ごとにセグメントしたグループの直近一週間でのarticleの出現上位10000位から最大4096個をサンプリングしました。\nサンプリングは[Gumbel-max trick](https://stackoverflow.com/questions/43310075/sampling-without-replacement-from-a-given-non-uniform-distribution-in-tensorflow)を用いました。\n\n推論時は直近7日間で誰かしらの購入履歴がある、直近90日で良く出現した10000個のアイテムに対して、コサイン類似度を総当たりで計算しました。(ScaNNのような近似近傍探索も試しましたが、時間的な影響はごくわずかでした)\n\n\n## Ranking part\nランキングではTransformerを用いランク学習を行いました。\n\n### Model\n推薦対称のarticleをQuery、ユーザーの閲覧履歴をKey & ValueとしたTransformerを一層用いました。ここでもself-attention型のEncoderや、二層以上のDecoderは不要でした。\nこれによって得たEmbeddingと、それまでに計算した中間層の結果をDCNV2に入れて、二値分類を行いました。\n\n### Loss\nlossは[ApproxNDCGLoss](https://www.tensorflow.org/ranking/api_docs/python/tfr/keras/losses/ApproxNDCGLoss)を元に、順位ごとのペナルティを順位差にした独自のlossを用いました。\nApproxNDCGLoss(original)、順位差の一乗、順位差の二乗を比較しましたが、順位差の一乗が最も良い結果でした。\n\n### Subsampling\n訓練時はコサイン類似度の3乗の比を確率としてGumbel-max trickで負例を最大128個作成しました。これは上位をそのままとってくるよりも程よく弱い負例をとることができモデルに良いフィードバックを与えました。\n推論時はRetrievalパートの上位128件と過去の閲覧履歴について予測を行いました。\n\n# Machine & Experiments\n私はcorei9-12900のCPU、64GBのRAM、RTX3090のGPUを用いました。\nそのうちCPUとRAMは半分程度の使用量で、全データを用いた学習は1epoch9時間、最高スコアは2epoch目で出ました。\n\n余談ですが、notebookの実験番号は最後で159になり、非常に労力を費やしたなと感じています。\n\n\n随時更新します。気軽に質問してください！",
    "1784099": "This is a great summary of a very effective solution! Thank you for sharing your process and results.",
    "1784115": "nadare congratulations! thank you for the detailed write up",
    "1810979": "Can you share the solution code? Many of us would like to learn the details.\n\nKaggle notebook or github repo or anyway is fine.",
    "1813100": "Thanks for your comment.\nIt's difficult to publish everything because the code isn't managed properly and only I can understand it.\nHowever, if there is a part that you are interested in, cut out that part and publish it (if possible).",
    "2636998": "I don't image DataFrame after you selected items each users.\n- vertical holding(?)  \n\n|user  | item |feature...|\n| --- | --- | --- |\n| A | item1 |...|\n| A | item2 |...|\n- horizontal holding(?). \n\n|user  | item candidate1 |item candidate2 |....|feature...|\n| --- | --- | --- | --- | --- |\n| A | item1 |item2|....|\n| B | item2 |...|....|"
  },
  "source": "meta"
}