{
  "id": 324310,
  "title": "12th place Writeup",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/writeups/kazuki-2-shimacos-1e-3-12th-place-writeup",
  "author_name": "",
  "post_date": "2022-05-11T16:48:49.523Z",
  "votes": 39,
  "comment_count": 9,
  "views": 0,
  "content": "<p>First of all, I would like to thank the competition hosts.<br>\nHere I will try to summarize some of the main points of our solution.</p>\n<p><img src=\"https://i.ibb.co/vxjNPwv/2022-05-10-21-39-35.png\" alt=\"2022-05-10-21-39-35\"></p>\n<h3><strong>GBDT part ( <a href=\"https://www.kaggle.com/onodera\" target=\"_blank\">@onodera</a> )</strong></h3>\n<p><a href=\"https://www.kaggle.com/onodera\" target=\"_blank\">@onodera</a> tried to solve this recommendation task by splitting into 3 models with GBDT(XGB).</p>\n<ol>\n<li>Reorder<br>\nPredict whether the customer will purchase the article which was purchased by the customer before</li>\n<li>new item top1k<br>\nPredict whether the customer will purchase the most purchased article(top1k) in last 1 week</li>\n<li>Other Color<br>\nPredicting whether a customer will purchase a different color of a product they have previously purchased</li>\n</ol>\n<p>And then got <strong>0.03242</strong> on both LB by blending above outputs.</p>\n<h3><strong>NN part ( KF )</strong></h3>\n<p>I used E2E NN approach based on the great <a href=\"https://www.kaggle.com/code/aerdem4/h-m-pure-pytorch-baseline\" target=\"_blank\">baseline notebook</a> published by <a href=\"https://www.kaggle.com/aerdem4\" target=\"_blank\">@aerdem4</a>.</p>\n<ol>\n<li>Learning metrics between candidate articles and transaction articles<ul>\n<li>Assign the following knowledge to the article's embedding vectors<ul>\n<li>Article meta information ID (e.g. product code)</li>\n<li>EfficientNet image features compressed by PCA</li>\n<li>Graph embedding using Deep Walk and Word2Vec</li></ul></li>\n<li>Training with the following classification losses.</li></ul></li>\n<li>Binary classification of whether a candidate article will be purchased or not<ul>\n<li>Top 2~5 transaction article features similar to each candidate articles<ul>\n<li>Similarity scores computed by metric learning module</li>\n<li>Statistical features for time elapsed since last purchase</li></ul></li>\n<li>Optimize BCELoss &amp; DiceLoss</li></ul></li>\n</ol>\n<p>And then got <strong>0.298~0.303</strong> on both LB for each single model.</p>\n<h3><strong>NN Stacking part ( <a href=\"https://www.kaggle.com/shimacos\" target=\"_blank\">@shimacos</a> )</strong></h3>\n<p><a href=\"https://www.kaggle.com/shimacos\" target=\"_blank\">@shimacos</a> built 2 Stacking model from KF's many NN models with LightGBM.<br>\nBy including these models, the result of weighted blend improved from 0.03377 to 0.03394.</p>\n<h3><strong>Weighted blending ( KF )</strong></h3>\n<p>I blended GBDT, NN, and stacked NN models using Optuna to optimize map12 directly.<br>\nAnd then got CV: <strong>0.04047</strong>, Public: <strong>0.03389</strong>, and Private: <strong>0.03394</strong>.</p>",
  "messages": [
    {
      "id": "1784169",
      "postDate": "05/11/2022 01:52:39",
      "content": "<p>First of all, I would like to thank the competition hosts.<br>\nHere I will try to summarize some of the main points of our solution.</p>\n<p><img src=\"https://i.ibb.co/vxjNPwv/2022-05-10-21-39-35.png\" alt=\"2022-05-10-21-39-35\"></p>\n<h3><strong>GBDT part ( <a href=\"https://www.kaggle.com/onodera\" target=\"_blank\">@onodera</a> )</strong></h3>\n<p><a href=\"https://www.kaggle.com/onodera\" target=\"_blank\">@onodera</a> tried to solve this recommendation task by splitting into 3 models with GBDT(XGB).</p>\n<ol>\n<li>Reorder<br>\nPredict whether the customer will purchase the article which was purchased by the customer before</li>\n<li>new item top1k<br>\nPredict whether the customer will purchase the most purchased article(top1k) in last 1 week</li>\n<li>Other Color<br>\nPredicting whether a customer will purchase a different color of a product they have previously purchased</li>\n</ol>\n<p>And then got <strong>0.03242</strong> on both LB by blending above outputs.</p>\n<h3><strong>NN part ( KF )</strong></h3>\n<p>I used E2E NN approach based on the great <a href=\"https://www.kaggle.com/code/aerdem4/h-m-pure-pytorch-baseline\" target=\"_blank\">baseline notebook</a> published by <a href=\"https://www.kaggle.com/aerdem4\" target=\"_blank\">@aerdem4</a>.</p>\n<ol>\n<li>Learning metrics between candidate articles and transaction articles<ul>\n<li>Assign the following knowledge to the article's embedding vectors<ul>\n<li>Article meta information ID (e.g. product code)</li>\n<li>EfficientNet image features compressed by PCA</li>\n<li>Graph embedding using Deep Walk and Word2Vec</li></ul></li>\n<li>Training with the following classification losses.</li></ul></li>\n<li>Binary classification of whether a candidate article will be purchased or not<ul>\n<li>Top 2~5 transaction article features similar to each candidate articles<ul>\n<li>Similarity scores computed by metric learning module</li>\n<li>Statistical features for time elapsed since last purchase</li></ul></li>\n<li>Optimize BCELoss &amp; DiceLoss</li></ul></li>\n</ol>\n<p>And then got <strong>0.298~0.303</strong> on both LB for each single model.</p>\n<h3><strong>NN Stacking part ( <a href=\"https://www.kaggle.com/shimacos\" target=\"_blank\">@shimacos</a> )</strong></h3>\n<p><a href=\"https://www.kaggle.com/shimacos\" target=\"_blank\">@shimacos</a> built 2 Stacking model from KF's many NN models with LightGBM.<br>\nBy including these models, the result of weighted blend improved from 0.03377 to 0.03394.</p>\n<h3><strong>Weighted blending ( KF )</strong></h3>\n<p>I blended GBDT, NN, and stacked NN models using Optuna to optimize map12 directly.<br>\nAnd then got CV: <strong>0.04047</strong>, Public: <strong>0.03389</strong>, and Private: <strong>0.03394</strong>.</p>",
      "rawMarkdown": "First of all, I would like to thank the competition hosts.\nHere I will try to summarize some of the main points of our solution.\n\n<img src=\"https://i.ibb.co/vxjNPwv/2022-05-10-21-39-35.png\" alt=\"2022-05-10-21-39-35\" border=\"0\">\n\n### **GBDT part ( @onodera )**\n\n@onodera tried to solve this recommendation task by splitting into 3 models with GBDT(XGB).\n1. Reorder\n    Predict whether the customer will purchase the article which was purchased by the customer before\n2. new item top1k\n    Predict whether the customer will purchase the most purchased article(top1k) in last 1 week\n3. Other Color\n    Predicting whether a customer will purchase a different color of a product they have previously purchased\n\nAnd then got **0.03242** on both LB by blending above outputs.\n\n### **NN part ( KF )**\n\nI used E2E NN approach based on the great [baseline notebook](https://www.kaggle.com/code/aerdem4/h-m-pure-pytorch-baseline) published by @aerdem4.\n\n1. Learning metrics between candidate articles and transaction articles\n  - Assign the following knowledge to the article's embedding vectors\n      - Article meta information ID (e.g. product code)\n      - EfficientNet image features compressed by PCA\n      - Graph embedding using Deep Walk and Word2Vec\n  - Training with the following classification losses.\n2. Binary classification of whether a candidate article will be purchased or not\n  - Top 2~5 transaction article features similar to each candidate articles\n      - Similarity scores computed by metric learning module\n      - Statistical features for time elapsed since last purchase\n  - Optimize BCELoss & DiceLoss\n\nAnd then got **0.298~0.303** on both LB for each single model.\n\n### **NN Stacking part ( @shimacos )**\n\n@shimacos built 2 Stacking model from KF's many NN models with LightGBM.\nBy including these models, the result of weighted blend improved from 0.03377 to 0.03394.\n\n### **Weighted blending ( KF )**\n\nI blended GBDT, NN, and stacked NN models using Optuna to optimize map12 directly.\nAnd then got CV: **0.04047**, Public: **0.03389**, and Private: **0.03394**.",
      "votes": null
    },
    {
      "id": "1784179",
      "postDate": "05/11/2022 02:12:15",
      "content": "<p>Congratulations! In the NN part, I'm interested in your methodology for creating embeddings. I was surprised to see \"Efficientnet image features compressed by PCA\". This sounds very cool (purely for image tasks), and I would appreciate if you could either talk about or share some resources on how PCA could be applied to feature embeddings of images. </p>",
      "rawMarkdown": "Congratulations! In the NN part, I'm interested in your methodology for creating embeddings. I was surprised to see \"Efficientnet image features compressed by PCA\". This sounds very cool (purely for image tasks), and I would appreciate if you could either talk about or share some resources on how PCA could be applied to feature embeddings of images.",
      "votes": null
    },
    {
      "id": "1784370",
      "postDate": "05/11/2022 06:24:25",
      "content": "<p>“Efficientnet image features compressed by PCA\" is just computed by timm and cuml.PCA, and then combined with other categorical / numerical features as shown in the figure.<br>\n<img src=\"https://i.ibb.co/KsWSxBc/2022-05-11-15-21-05.png\" alt=\"figure\"></p>",
      "rawMarkdown": "“Efficientnet image features compressed by PCA\" is just computed by timm and cuml.PCA, and then combined with other categorical / numerical features as shown in the figure.\n![figure](https://i.ibb.co/KsWSxBc/2022-05-11-15-21-05.png)",
      "votes": null
    },
    {
      "id": "1784540",
      "postDate": "05/11/2022 08:33:59",
      "content": "<p>Great job on your 12th place finish! This is a really thorough write-up of your solution. It's great to see the different approaches you took and how you blended them together. Thanks for sharing!</p>",
      "rawMarkdown": "Great job on your 12th place finish! This is a really thorough write-up of your solution. It's great to see the different approaches you took and how you blended them together. Thanks for sharing!",
      "votes": null
    },
    {
      "id": "1784552",
      "postDate": "05/11/2022 08:44:50",
      "content": "<p><a href=\"https://www.kaggle.com/kfujikawa\" target=\"_blank\">@kfujikawa</a> <br>\nCongratulations for your gold medal!!<br>\nI have a question about NN part.<br>\nWhat do you mean when you say “ Learning metrics between candidate articles and transaction articles”? </p>",
      "rawMarkdown": "kfujikawa \nCongratulations for your gold medal!!\nI have a question about NN part.\nWhat do you mean when you say “ Learning metrics between candidate articles and transaction articles”?",
      "votes": null
    },
    {
      "id": "1784907",
      "postDate": "05/11/2022 15:17:46",
      "content": "<p>The binary classifier takes as input the features for the transaction articles that has the top2~5 cosine similarity to the candidate articles.<br>\nThis cosine similarity is calculated by article embeddings that can be learned by BCELoss, so I understood that this is a kind of metric learning.<br>\n(Sorry if my understanding is wrong.)</p>",
      "rawMarkdown": "The binary classifier takes as input the features for the transaction articles that has the top2~5 cosine similarity to the candidate articles.\nThis cosine similarity is calculated by article embeddings that can be learned by BCELoss, so I understood that this is a kind of metric learning.\n(Sorry if my understanding is wrong.)",
      "votes": null
    },
    {
      "id": "1785566",
      "postDate": "05/12/2022 08:05:59",
      "content": "<p>This is a really good write-up of your solution. Congratulations on your gold medal👍👍</p>",
      "rawMarkdown": "This is a really good write-up of your solution. Congratulations on your gold medal👍👍",
      "votes": null
    },
    {
      "id": "1785649",
      "postDate": "05/12/2022 09:55:15",
      "content": "<p><a href=\"https://www.kaggle.com/kfujikawa\" target=\"_blank\">@kfujikawa</a> <br>\nThank you! I understood your point.<br>\nI only know the concept of metric learning, but I think you're right.</p>\n<p>Could I ask you another question?<br>\nYou said \"Assign the following knowledge to the article's embedding vectors.\"<br>\nIf you don't mind, please show me some real code or dummy code to do so. I'm a beginner of PyTorch, I don't know how I could apply it in code.</p>\n<pre><code>class EmbeddingLayer(nn.Module):\n def __init__():\n    self.article_id_embedding = torch.nn.Embedding(\n        ARTICLE_ID_NUM, embedding_dim=EMBEDDING_DIM)\n    self.efficient_net_embedding = torch.nn.Linear(\n        EFFICIENT_NET_DIM, EMBEDDING_DIM)\n\n def forward(article_history):\n    efficient_net_history = get_efficient_net(article_history)\n    embedded_history = self.article_id_embedding(article_history) + self.efficient_net_embedding(efficient_net_history)\n\n    return embedded_history\n</code></pre>\n<p>I tried to write code following your figure in another comment. Is it what you meant?</p>",
      "rawMarkdown": "kfujikawa \nThank you! I understood your point.\nI only know the concept of metric learning, but I think you're right.\n\nCould I ask you another question?\nYou said \"Assign the following knowledge to the article's embedding vectors.\"\nIf you don't mind, please show me some real code or dummy code to do so. I'm a beginner of PyTorch, I don't know how I could apply it in code.\n\n```\nclass EmbeddingLayer(nn.Module):\n def __init__():\n    self.article_id_embedding = torch.nn.Embedding(\n        ARTICLE_ID_NUM, embedding_dim=EMBEDDING_DIM)\n    self.efficient_net_embedding = torch.nn.Linear(\n        EFFICIENT_NET_DIM, EMBEDDING_DIM)\n\n def forward(article_history):\n    efficient_net_history = get_efficient_net(article_history)\n    embedded_history = self.article_id_embedding(article_history) + self.efficient_net_embedding(efficient_net_history)\n    \n    return embedded_history\n```\nI tried to write code following your figure in another comment. Is it what you meant?",
      "votes": null
    },
    {
      "id": "1786447",
      "postDate": "05/12/2022 22:45:51",
      "content": "<p>You are right. Real code is here:</p>\n<pre><code>class AdditiveFeatureEmbedding(torch.nn.Module):\n    def __init__(\n        self,\n        categorical_embeddings: Optional[dict] = None,\n        numerical_embeddings: Optional[dict] = None,\n    ):\n        super().__init__()\n        self.categorical_embeddings = torch.nn.ModuleDict(\n            categorical_embeddings or {}\n        )\n        self.numerical_embeddings = torch.nn.ModuleDict(\n            numerical_embeddings or {}\n        )\n\n    @property\n    def columns(self):\n        return [\n            *self.categorical_embeddings.keys(),\n            *self.numerical_embeddings.keys(),\n        ]\n\n    def __call__(self, **inputs):\n        h1 = sum([\n            f(inputs[k])\n            for k, f in self.categorical_embeddings.items()\n        ])\n        h2 = sum([\n            f(inputs[k])\n            for k, f in self.numerical_embeddings.items()\n        ])\n        h = h1 + h2\n        return h\n</code></pre>\n<p>And then</p>\n<pre><code>AdditiveFeatureEmbedding(\n    categorical_embeddings={\n        \"article_embed_ids\": torch.nn.Embedding(\n            num_embeddings=NUM_EMBEDDINGS,\n            embedding_dim=HIDDEN_DIM,\n        ),\n        ...\n    },\n    numerical_embeddings={\n        \"efficientnet_pca\": torch.nn.LazyLinear(\n            out_features =HIDDEN_DIM\n        ),\n        ...\n    }\n)\n</code></pre>",
      "rawMarkdown": "You are right. Real code is here:\n```\nclass AdditiveFeatureEmbedding(torch.nn.Module):\n    def __init__(\n        self,\n        categorical_embeddings: Optional[dict] = None,\n        numerical_embeddings: Optional[dict] = None,\n    ):\n        super().__init__()\n        self.categorical_embeddings = torch.nn.ModuleDict(\n            categorical_embeddings or {}\n        )\n        self.numerical_embeddings = torch.nn.ModuleDict(\n            numerical_embeddings or {}\n        )\n\n    @property\n    def columns(self):\n        return [\n            *self.categorical_embeddings.keys(),\n            *self.numerical_embeddings.keys(),\n        ]\n\n    def __call__(self, **inputs):\n        h1 = sum([\n            f(inputs[k])\n            for k, f in self.categorical_embeddings.items()\n        ])\n        h2 = sum([\n            f(inputs[k])\n            for k, f in self.numerical_embeddings.items()\n        ])\n        h = h1 + h2\n        return h\n```\n\nAnd then\n\n```\nAdditiveFeatureEmbedding(\n    categorical_embeddings={\n        \"article_embed_ids\": torch.nn.Embedding(\n            num_embeddings=NUM_EMBEDDINGS,\n            embedding_dim=HIDDEN_DIM,\n        ),\n        ...\n    },\n    numerical_embeddings={\n        \"efficientnet_pca\": torch.nn.LazyLinear(\n            out_features =HIDDEN_DIM\n        ),\n        ...\n    }\n)\n```",
      "votes": null
    },
    {
      "id": "1786579",
      "postDate": "05/13/2022 03:29:12",
      "content": "<p><a href=\"https://www.kaggle.com/kfujikawa\" target=\"_blank\">@kfujikawa</a> <br>\nThank you for sharing real code!!<br>\nI can learn from it.</p>",
      "rawMarkdown": "kfujikawa \nThank you for sharing real code!!\nI can learn from it.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1784179,
      "author_name": "anandparthiban",
      "author_url": "",
      "post_date": "05/11/2022 02:12:15",
      "content": "<p>Congratulations! In the NN part, I'm interested in your methodology for creating embeddings. I was surprised to see \"Efficientnet image features compressed by PCA\". This sounds very cool (purely for image tasks), and I would appreciate if you could either talk about or share some resources on how PCA could be applied to feature embeddings of images. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1784370,
          "author_name": "kfujikawa",
          "author_url": "",
          "post_date": "05/11/2022 06:24:25",
          "content": "<p>“Efficientnet image features compressed by PCA\" is just computed by timm and cuml.PCA, and then combined with other categorical / numerical features as shown in the figure.<br>\n<img src=\"https://i.ibb.co/KsWSxBc/2022-05-11-15-21-05.png\" alt=\"figure\"></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1784540,
      "author_name": "",
      "author_url": "",
      "post_date": "05/11/2022 08:33:59",
      "content": "<p>Great job on your 12th place finish! This is a really thorough write-up of your solution. It's great to see the different approaches you took and how you blended them together. Thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1784552,
      "author_name": "hanejiyuto",
      "author_url": "",
      "post_date": "05/11/2022 08:44:50",
      "content": "<p><a href=\"https://www.kaggle.com/kfujikawa\" target=\"_blank\">@kfujikawa</a> <br>\nCongratulations for your gold medal!!<br>\nI have a question about NN part.<br>\nWhat do you mean when you say “ Learning metrics between candidate articles and transaction articles”? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1784907,
          "author_name": "kfujikawa",
          "author_url": "",
          "post_date": "05/11/2022 15:17:46",
          "content": "<p>The binary classifier takes as input the features for the transaction articles that has the top2~5 cosine similarity to the candidate articles.<br>\nThis cosine similarity is calculated by article embeddings that can be learned by BCELoss, so I understood that this is a kind of metric learning.<br>\n(Sorry if my understanding is wrong.)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1785649,
          "author_name": "hanejiyuto",
          "author_url": "",
          "post_date": "05/12/2022 09:55:15",
          "content": "<p><a href=\"https://www.kaggle.com/kfujikawa\" target=\"_blank\">@kfujikawa</a> <br>\nThank you! I understood your point.<br>\nI only know the concept of metric learning, but I think you're right.</p>\n<p>Could I ask you another question?<br>\nYou said \"Assign the following knowledge to the article's embedding vectors.\"<br>\nIf you don't mind, please show me some real code or dummy code to do so. I'm a beginner of PyTorch, I don't know how I could apply it in code.</p>\n<pre><code>class EmbeddingLayer(nn.Module):\n def __init__():\n    self.article_id_embedding = torch.nn.Embedding(\n        ARTICLE_ID_NUM, embedding_dim=EMBEDDING_DIM)\n    self.efficient_net_embedding = torch.nn.Linear(\n        EFFICIENT_NET_DIM, EMBEDDING_DIM)\n\n def forward(article_history):\n    efficient_net_history = get_efficient_net(article_history)\n    embedded_history = self.article_id_embedding(article_history) + self.efficient_net_embedding(efficient_net_history)\n\n    return embedded_history\n</code></pre>\n<p>I tried to write code following your figure in another comment. Is it what you meant?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1786447,
          "author_name": "kfujikawa",
          "author_url": "",
          "post_date": "05/12/2022 22:45:51",
          "content": "<p>You are right. Real code is here:</p>\n<pre><code>class AdditiveFeatureEmbedding(torch.nn.Module):\n    def __init__(\n        self,\n        categorical_embeddings: Optional[dict] = None,\n        numerical_embeddings: Optional[dict] = None,\n    ):\n        super().__init__()\n        self.categorical_embeddings = torch.nn.ModuleDict(\n            categorical_embeddings or {}\n        )\n        self.numerical_embeddings = torch.nn.ModuleDict(\n            numerical_embeddings or {}\n        )\n\n    @property\n    def columns(self):\n        return [\n            *self.categorical_embeddings.keys(),\n            *self.numerical_embeddings.keys(),\n        ]\n\n    def __call__(self, **inputs):\n        h1 = sum([\n            f(inputs[k])\n            for k, f in self.categorical_embeddings.items()\n        ])\n        h2 = sum([\n            f(inputs[k])\n            for k, f in self.numerical_embeddings.items()\n        ])\n        h = h1 + h2\n        return h\n</code></pre>\n<p>And then</p>\n<pre><code>AdditiveFeatureEmbedding(\n    categorical_embeddings={\n        \"article_embed_ids\": torch.nn.Embedding(\n            num_embeddings=NUM_EMBEDDINGS,\n            embedding_dim=HIDDEN_DIM,\n        ),\n        ...\n    },\n    numerical_embeddings={\n        \"efficientnet_pca\": torch.nn.LazyLinear(\n            out_features =HIDDEN_DIM\n        ),\n        ...\n    }\n)\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1786579,
          "author_name": "hanejiyuto",
          "author_url": "",
          "post_date": "05/13/2022 03:29:12",
          "content": "<p><a href=\"https://www.kaggle.com/kfujikawa\" target=\"_blank\">@kfujikawa</a> <br>\nThank you for sharing real code!!<br>\nI can learn from it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1785566,
      "author_name": "sanikamal",
      "author_url": "",
      "post_date": "05/12/2022 08:05:59",
      "content": "<p>This is a really good write-up of your solution. Congratulations on your gold medal👍👍</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1784169": "First of all, I would like to thank the competition hosts.\nHere I will try to summarize some of the main points of our solution.\n\n<img src=\"https://i.ibb.co/vxjNPwv/2022-05-10-21-39-35.png\" alt=\"2022-05-10-21-39-35\" border=\"0\">\n\n### **GBDT part ( @onodera )**\n\n@onodera tried to solve this recommendation task by splitting into 3 models with GBDT(XGB).\n1. Reorder\n    Predict whether the customer will purchase the article which was purchased by the customer before\n2. new item top1k\n    Predict whether the customer will purchase the most purchased article(top1k) in last 1 week\n3. Other Color\n    Predicting whether a customer will purchase a different color of a product they have previously purchased\n\nAnd then got **0.03242** on both LB by blending above outputs.\n\n### **NN part ( KF )**\n\nI used E2E NN approach based on the great [baseline notebook](https://www.kaggle.com/code/aerdem4/h-m-pure-pytorch-baseline) published by @aerdem4.\n\n1. Learning metrics between candidate articles and transaction articles\n  - Assign the following knowledge to the article's embedding vectors\n      - Article meta information ID (e.g. product code)\n      - EfficientNet image features compressed by PCA\n      - Graph embedding using Deep Walk and Word2Vec\n  - Training with the following classification losses.\n2. Binary classification of whether a candidate article will be purchased or not\n  - Top 2~5 transaction article features similar to each candidate articles\n      - Similarity scores computed by metric learning module\n      - Statistical features for time elapsed since last purchase\n  - Optimize BCELoss & DiceLoss\n\nAnd then got **0.298~0.303** on both LB for each single model.\n\n### **NN Stacking part ( @shimacos )**\n\n@shimacos built 2 Stacking model from KF's many NN models with LightGBM.\nBy including these models, the result of weighted blend improved from 0.03377 to 0.03394.\n\n### **Weighted blending ( KF )**\n\nI blended GBDT, NN, and stacked NN models using Optuna to optimize map12 directly.\nAnd then got CV: **0.04047**, Public: **0.03389**, and Private: **0.03394**.",
    "1784179": "Congratulations! In the NN part, I'm interested in your methodology for creating embeddings. I was surprised to see \"Efficientnet image features compressed by PCA\". This sounds very cool (purely for image tasks), and I would appreciate if you could either talk about or share some resources on how PCA could be applied to feature embeddings of images.",
    "1784370": "“Efficientnet image features compressed by PCA\" is just computed by timm and cuml.PCA, and then combined with other categorical / numerical features as shown in the figure.\n![figure](https://i.ibb.co/KsWSxBc/2022-05-11-15-21-05.png)",
    "1784540": "Great job on your 12th place finish! This is a really thorough write-up of your solution. It's great to see the different approaches you took and how you blended them together. Thanks for sharing!",
    "1784552": "kfujikawa \nCongratulations for your gold medal!!\nI have a question about NN part.\nWhat do you mean when you say “ Learning metrics between candidate articles and transaction articles”?",
    "1784907": "The binary classifier takes as input the features for the transaction articles that has the top2~5 cosine similarity to the candidate articles.\nThis cosine similarity is calculated by article embeddings that can be learned by BCELoss, so I understood that this is a kind of metric learning.\n(Sorry if my understanding is wrong.)",
    "1785566": "This is a really good write-up of your solution. Congratulations on your gold medal👍👍",
    "1785649": "kfujikawa \nThank you! I understood your point.\nI only know the concept of metric learning, but I think you're right.\n\nCould I ask you another question?\nYou said \"Assign the following knowledge to the article's embedding vectors.\"\nIf you don't mind, please show me some real code or dummy code to do so. I'm a beginner of PyTorch, I don't know how I could apply it in code.\n\n```\nclass EmbeddingLayer(nn.Module):\n def __init__():\n    self.article_id_embedding = torch.nn.Embedding(\n        ARTICLE_ID_NUM, embedding_dim=EMBEDDING_DIM)\n    self.efficient_net_embedding = torch.nn.Linear(\n        EFFICIENT_NET_DIM, EMBEDDING_DIM)\n\n def forward(article_history):\n    efficient_net_history = get_efficient_net(article_history)\n    embedded_history = self.article_id_embedding(article_history) + self.efficient_net_embedding(efficient_net_history)\n    \n    return embedded_history\n```\nI tried to write code following your figure in another comment. Is it what you meant?",
    "1786447": "You are right. Real code is here:\n```\nclass AdditiveFeatureEmbedding(torch.nn.Module):\n    def __init__(\n        self,\n        categorical_embeddings: Optional[dict] = None,\n        numerical_embeddings: Optional[dict] = None,\n    ):\n        super().__init__()\n        self.categorical_embeddings = torch.nn.ModuleDict(\n            categorical_embeddings or {}\n        )\n        self.numerical_embeddings = torch.nn.ModuleDict(\n            numerical_embeddings or {}\n        )\n\n    @property\n    def columns(self):\n        return [\n            *self.categorical_embeddings.keys(),\n            *self.numerical_embeddings.keys(),\n        ]\n\n    def __call__(self, **inputs):\n        h1 = sum([\n            f(inputs[k])\n            for k, f in self.categorical_embeddings.items()\n        ])\n        h2 = sum([\n            f(inputs[k])\n            for k, f in self.numerical_embeddings.items()\n        ])\n        h = h1 + h2\n        return h\n```\n\nAnd then\n\n```\nAdditiveFeatureEmbedding(\n    categorical_embeddings={\n        \"article_embed_ids\": torch.nn.Embedding(\n            num_embeddings=NUM_EMBEDDINGS,\n            embedding_dim=HIDDEN_DIM,\n        ),\n        ...\n    },\n    numerical_embeddings={\n        \"efficientnet_pca\": torch.nn.LazyLinear(\n            out_features =HIDDEN_DIM\n        ),\n        ...\n    }\n)\n```",
    "1786579": "kfujikawa \nThank you for sharing real code!!\nI can learn from it."
  },
  "source": "meta"
}