{
  "id": 324220,
  "title": "63th place solution",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/writeups/papapower-63th-place-solution",
  "author_name": "",
  "post_date": "2022-07-12T14:16:19.987Z",
  "votes": 12,
  "comment_count": 2,
  "views": 0,
  "content": "<h1>63th Place Solution</h1>\n<p>Thanks to my teammates <a href=\"https://www.kaggle.com/ootake\" target=\"_blank\">@ootake</a>, <a href=\"https://www.kaggle.com/ryoichi0917\" target=\"_blank\">@ryoichi0917</a>, <a href=\"https://www.kaggle.com/xia464\" target=\"_blank\">@xia464</a>, especially <a href=\"https://www.kaggle.com/tyonemoto\" target=\"_blank\">@tyonemoto</a> for inviting me.</p>\n<h2>Code (updated in 2022-07-12)</h2>\n<p>I publish the entire code.<br>\ncode: <a href=\"https://github.com/Y-Haneji/kaggle-handm\" target=\"_blank\">https://github.com/Y-Haneji/kaggle-handm</a></p>\n<h2>Last Submission</h2>\n<ul>\n<li>(lr=0.1) : Private LB: 0.02813, Public LB: 0.02826, Validation: 0.033623</li>\n<li>(lr=0.02): Private LB: 0.02772, Public LB: 0.0277, Validation: 0.033814</li>\n</ul>\n<h2>Summary</h2>\n<p>I think that the key points of this competition are these.</p>\n<ul>\n<li>Stable validation.</li>\n<li>Two stage: generating candidates by some strategies, then ranking them.</li>\n<li>Better strategies and scores from them.</li>\n<li>Using full information like not only transactions but also detail descriptions, images, and so on.</li>\n</ul>\n<h2>Ranking Model (me)</h2>\n<h3>Model</h3>\n<p>LightGBM lambdarank</p>\n<h3>Dataset</h3>\n<p>This competition gave no training datasets, only transactions datas and some metadatas. I needed to generate training datasets by myself. The dataset was like this.</p>\n<ul>\n<li>customer_id</li>\n<li>article_id (candidate)</li>\n<li>score from generative strategy (may not just one score)</li>\n<li>flags wheter the candidate included each strategy</li>\n<li>other features of customers</li>\n<li>other features of articles</li>\n<li>label (0 or 1)</li>\n</ul>\n<p>The features of customers and articles are purchase and sales of …</p>\n<ul>\n<li>each week</li>\n<li>agregation (sum, mean, max, min) of all weeks</li>\n<li>difference and division between each week and aggregation</li>\n</ul>\n<p>Label is whether the article was bought target_week.</p>\n<p>The details of strategies to generate candidates is written in below.</p>\n<h3>Training</h3>\n<ul>\n<li>5 target weeks, this means 5 folds training dataset.</li>\n<li>For training, I only generated candidates for the customers who had some purchases in the target week.</li>\n<li>Model Validation: NDCG@12. The dataset of next target week was validation dataset. The last target week's training dataset had no validation datasets.</li>\n<li>Submission Validation: MAP@12 for last week's transactions.</li>\n</ul>\n\n<h3>Parameter</h3>\n<ul>\n<li>objective: lambdarank</li>\n<li>lr: 0.1 or 0.02</li>\n<li>feature_fraction: 0.8</li>\n<li>baggin_fraction: 0.8</li>\n<li>early_stopping: 10 or 50 rounds (depend on lr)</li>\n</ul>\n<p>Other parameters is default. I tried Optuna but it didn't work well for lambdarank because it attempted to LOWER the MAP.</p>\n<h3>Prediction (Submission)</h3>\n<ul>\n<li>For prediction, I should generate candidates for all the customers, which causes a memory problem.</li>\n<li>Split the customers into batches.</li>\n<li>Ensemble 5 target week's models.</li>\n<li>Generally take longer time than training</li>\n<li>If validation was not good, I didn't compute submission.</li>\n</ul>\n<h2>Candidates Strategies</h2>\n<h3>Worked (included last submission)</h3>\n<ul>\n<li>trending (original: <a href=\"https://www.kaggle.com/code/byfone/h-m-trending-products-weekly\" target=\"_blank\">here</a>)</li>\n<li>recently purchased</li>\n<li>bin popular</li>\n<li>last week for each customer</li>\n<li>also bought this (original: <a href=\"https://www.kaggle.com/code/cdeotte/customers-who-bought-this-frequently-buy-this\" target=\"_blank\">here</a>)</li>\n<li>collabo</li>\n<li>t-SNE KNN</li>\n<li>Albert KNN</li>\n<li>sequence Word2Vec</li>\n</ul>\n<h3>Not Worked</h3>\n<ul>\n<li>Bert KNN</li>\n<li>VGG KNN</li>\n<li>Gru4Rec, Bert4Rec (LSTM) (original: <a href=\"https://www.kaggle.com/code/astrung/lstm-model-with-item-infor-fix-missing-last-item\" target=\"_blank\">here</a>)</li>\n<li>Implicit ALS</li>\n</ul>\n<h3>Details of Some Strategies</h3>\n<p>TO BE UPDATED: Some strategies will be written in some days.</p>\n<h4>bin popular (by me and by <a href=\"https://www.kaggle.com/tyonemoto\" target=\"_blank\">@tyonemoto</a>)</h4>\n<p>To recommend popular item was effecitve especially for cold start customers. Different kind of customers should have different preference. So we split customers into some groups by their features, like age, sex, and so on.</p>\n<p>In the experiment phase, we used age, price, and transaction channel for binning. The number of splitting is 10 for age, 20 for price, and 3 for channels. Recall of this strategy was 0.06750.</p>\n<p>Because of the deadline, we have no choice but to adopt only age bins.</p>\n<h4>collabo (by <a href=\"https://www.kaggle.com/tyonemoto\" target=\"_blank\">@tyonemoto</a>)</h4>\n<p>The collaboration articles with Kangol was launched a few weeks ago from the submission week. We recommended popular collaboration articles to customers who had bought some collaboration articles. Recall of this strategy was 0.00094.</p>\n<h4>t-SNE (by <a href=\"https://www.kaggle.com/tyonemoto\" target=\"_blank\">@tyonemoto</a>)</h4>\n<p>The metadatas of articles from the file of articles.csv and embedding vectors of detail descriptions by tf-idf and bert are compressed to two dimensions by t-SNE. We used count encoding method to deal with the category metadatas columns.</p>\n<h4>Albert (by <a href=\"https://www.kaggle.com/ryoichi0917\" target=\"_blank\">@ryoichi0917</a>)</h4>\n<h4>sequence Word2Vec (by <a href=\"https://www.kaggle.com/tyonemoto\" target=\"_blank\">@tyonemoto</a>)</h4>\n<h4>Bert (by <a href=\"https://www.kaggle.com/tyonemoto\" target=\"_blank\">@tyonemoto</a> and <a href=\"https://www.kaggle.com/ootake\" target=\"_blank\">@ootake</a>)</h4>\n<p>We used bert pretrained model named \"bert-uncased\", \"MRPC\", and \"SQuAD\" to extract features from articles' detail description.</p>\n<h4>VGG (by me)</h4>\n<p>I extracted features from images using VGG because their vectors were relatively low dimensions, which can avoid a memory problem. But even they reached 10GB files (when compressed to .npy).<br>\nVGG KNN doesn't worked, but I wonder if I could use other bigger models to extract features.</p>\n<h4>Implicit ALS (by <a href=\"https://www.kaggle.com/ootake\" target=\"_blank\">@ootake</a>)</h4>\n<h3>Comparision</h3>\n<p>The true purpose of this competition is not to explore strategies which have better Precision and Recall, but to obtain better MAP@12 for submission, so I compared each strategy based on MAP@12.</p>\n<table>\n<thead>\n<tr>\n<th>Strategy</th>\n<th>MAP@12 change</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>trending</td>\n<td>base lines</td>\n</tr>\n<tr>\n<td>recently</td>\n<td>base lines</td>\n</tr>\n<tr>\n<td>popular</td>\n<td>base lines</td>\n</tr>\n<tr>\n<td>last week</td>\n<td>-0.000131</td>\n</tr>\n<tr>\n<td>t-SNE</td>\n<td>+0.00053</td>\n</tr>\n<tr>\n<td>umap</td>\n<td>-0.0001</td>\n</tr>\n<tr>\n<td>t-SNE &amp; umap</td>\n<td>-0.000898</td>\n</tr>\n<tr>\n<td>also bought this</td>\n<td>+0.000533</td>\n</tr>\n<tr>\n<td>albert</td>\n<td>+0.000241</td>\n</tr>\n<tr>\n<td>collaboration</td>\n<td>+0.000134</td>\n</tr>\n<tr>\n<td>Bert</td>\n<td>+0.000051 (didn't correlate with LB)</td>\n</tr>\n<tr>\n<td>sequence Word2Vec w/ compression</td>\n<td>+0.000367</td>\n</tr>\n<tr>\n<td>bin popular</td>\n<td>+0.00058</td>\n</tr>\n<tr>\n<td>bert4rec</td>\n<td>-0.00037</td>\n</tr>\n<tr>\n<td>gru4rec</td>\n<td>-0.000347</td>\n</tr>\n<tr>\n<td>Implicit ALS</td>\n<td>+0.000142</td>\n</tr>\n<tr>\n<td>sequence Word2Vec</td>\n<td>-0.000563</td>\n</tr>\n<tr>\n<td>VGG</td>\n<td>-0.000375</td>\n</tr>\n</tbody>\n</table>\n<p>Note: This is the PROCESS of adding or subtracting the strategies rather than the result of comparision experiment.</p>\n<h2>Computing Resources</h2>\n<ul>\n<li>Untill 2 weeks from deadlines: MacBook Air (M1, 2020) (8 core CPU, 16 GM RAM, no GPU for deep learning).</li>\n<li>Last 2 weeks: GCP instances to deal with more candidates. (48 core CPU, 384 GB RAM)</li>\n</ul>\n<h2>Some Story</h2>\n<p>This is my first medal in Kaggle, so I'm very happy with my place! </p>\n<p>I am new to a recommend problem, but my teammates helped me a lot, and <a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a> and <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> was so kind to answer my questions in discussion.</p>",
  "messages": [
    {
      "id": "1783657",
      "postDate": "05/10/2022 15:00:05",
      "content": "<h1>63th Place Solution</h1>\n<p>Thanks to my teammates <a href=\"https://www.kaggle.com/ootake\" target=\"_blank\">@ootake</a>, <a href=\"https://www.kaggle.com/ryoichi0917\" target=\"_blank\">@ryoichi0917</a>, <a href=\"https://www.kaggle.com/xia464\" target=\"_blank\">@xia464</a>, especially <a href=\"https://www.kaggle.com/tyonemoto\" target=\"_blank\">@tyonemoto</a> for inviting me.</p>\n<h2>Code (updated in 2022-07-12)</h2>\n<p>I publish the entire code.<br>\ncode: <a href=\"https://github.com/Y-Haneji/kaggle-handm\" target=\"_blank\">https://github.com/Y-Haneji/kaggle-handm</a></p>\n<h2>Last Submission</h2>\n<ul>\n<li>(lr=0.1) : Private LB: 0.02813, Public LB: 0.02826, Validation: 0.033623</li>\n<li>(lr=0.02): Private LB: 0.02772, Public LB: 0.0277, Validation: 0.033814</li>\n</ul>\n<h2>Summary</h2>\n<p>I think that the key points of this competition are these.</p>\n<ul>\n<li>Stable validation.</li>\n<li>Two stage: generating candidates by some strategies, then ranking them.</li>\n<li>Better strategies and scores from them.</li>\n<li>Using full information like not only transactions but also detail descriptions, images, and so on.</li>\n</ul>\n<h2>Ranking Model (me)</h2>\n<h3>Model</h3>\n<p>LightGBM lambdarank</p>\n<h3>Dataset</h3>\n<p>This competition gave no training datasets, only transactions datas and some metadatas. I needed to generate training datasets by myself. The dataset was like this.</p>\n<ul>\n<li>customer_id</li>\n<li>article_id (candidate)</li>\n<li>score from generative strategy (may not just one score)</li>\n<li>flags wheter the candidate included each strategy</li>\n<li>other features of customers</li>\n<li>other features of articles</li>\n<li>label (0 or 1)</li>\n</ul>\n<p>The features of customers and articles are purchase and sales of …</p>\n<ul>\n<li>each week</li>\n<li>agregation (sum, mean, max, min) of all weeks</li>\n<li>difference and division between each week and aggregation</li>\n</ul>\n<p>Label is whether the article was bought target_week.</p>\n<p>The details of strategies to generate candidates is written in below.</p>\n<h3>Training</h3>\n<ul>\n<li>5 target weeks, this means 5 folds training dataset.</li>\n<li>For training, I only generated candidates for the customers who had some purchases in the target week.</li>\n<li>Model Validation: NDCG@12. The dataset of next target week was validation dataset. The last target week's training dataset had no validation datasets.</li>\n<li>Submission Validation: MAP@12 for last week's transactions.</li>\n</ul>\n\n<h3>Parameter</h3>\n<ul>\n<li>objective: lambdarank</li>\n<li>lr: 0.1 or 0.02</li>\n<li>feature_fraction: 0.8</li>\n<li>baggin_fraction: 0.8</li>\n<li>early_stopping: 10 or 50 rounds (depend on lr)</li>\n</ul>\n<p>Other parameters is default. I tried Optuna but it didn't work well for lambdarank because it attempted to LOWER the MAP.</p>\n<h3>Prediction (Submission)</h3>\n<ul>\n<li>For prediction, I should generate candidates for all the customers, which causes a memory problem.</li>\n<li>Split the customers into batches.</li>\n<li>Ensemble 5 target week's models.</li>\n<li>Generally take longer time than training</li>\n<li>If validation was not good, I didn't compute submission.</li>\n</ul>\n<h2>Candidates Strategies</h2>\n<h3>Worked (included last submission)</h3>\n<ul>\n<li>trending (original: <a href=\"https://www.kaggle.com/code/byfone/h-m-trending-products-weekly\" target=\"_blank\">here</a>)</li>\n<li>recently purchased</li>\n<li>bin popular</li>\n<li>last week for each customer</li>\n<li>also bought this (original: <a href=\"https://www.kaggle.com/code/cdeotte/customers-who-bought-this-frequently-buy-this\" target=\"_blank\">here</a>)</li>\n<li>collabo</li>\n<li>t-SNE KNN</li>\n<li>Albert KNN</li>\n<li>sequence Word2Vec</li>\n</ul>\n<h3>Not Worked</h3>\n<ul>\n<li>Bert KNN</li>\n<li>VGG KNN</li>\n<li>Gru4Rec, Bert4Rec (LSTM) (original: <a href=\"https://www.kaggle.com/code/astrung/lstm-model-with-item-infor-fix-missing-last-item\" target=\"_blank\">here</a>)</li>\n<li>Implicit ALS</li>\n</ul>\n<h3>Details of Some Strategies</h3>\n<p>TO BE UPDATED: Some strategies will be written in some days.</p>\n<h4>bin popular (by me and by <a href=\"https://www.kaggle.com/tyonemoto\" target=\"_blank\">@tyonemoto</a>)</h4>\n<p>To recommend popular item was effecitve especially for cold start customers. Different kind of customers should have different preference. So we split customers into some groups by their features, like age, sex, and so on.</p>\n<p>In the experiment phase, we used age, price, and transaction channel for binning. The number of splitting is 10 for age, 20 for price, and 3 for channels. Recall of this strategy was 0.06750.</p>\n<p>Because of the deadline, we have no choice but to adopt only age bins.</p>\n<h4>collabo (by <a href=\"https://www.kaggle.com/tyonemoto\" target=\"_blank\">@tyonemoto</a>)</h4>\n<p>The collaboration articles with Kangol was launched a few weeks ago from the submission week. We recommended popular collaboration articles to customers who had bought some collaboration articles. Recall of this strategy was 0.00094.</p>\n<h4>t-SNE (by <a href=\"https://www.kaggle.com/tyonemoto\" target=\"_blank\">@tyonemoto</a>)</h4>\n<p>The metadatas of articles from the file of articles.csv and embedding vectors of detail descriptions by tf-idf and bert are compressed to two dimensions by t-SNE. We used count encoding method to deal with the category metadatas columns.</p>\n<h4>Albert (by <a href=\"https://www.kaggle.com/ryoichi0917\" target=\"_blank\">@ryoichi0917</a>)</h4>\n<h4>sequence Word2Vec (by <a href=\"https://www.kaggle.com/tyonemoto\" target=\"_blank\">@tyonemoto</a>)</h4>\n<h4>Bert (by <a href=\"https://www.kaggle.com/tyonemoto\" target=\"_blank\">@tyonemoto</a> and <a href=\"https://www.kaggle.com/ootake\" target=\"_blank\">@ootake</a>)</h4>\n<p>We used bert pretrained model named \"bert-uncased\", \"MRPC\", and \"SQuAD\" to extract features from articles' detail description.</p>\n<h4>VGG (by me)</h4>\n<p>I extracted features from images using VGG because their vectors were relatively low dimensions, which can avoid a memory problem. But even they reached 10GB files (when compressed to .npy).<br>\nVGG KNN doesn't worked, but I wonder if I could use other bigger models to extract features.</p>\n<h4>Implicit ALS (by <a href=\"https://www.kaggle.com/ootake\" target=\"_blank\">@ootake</a>)</h4>\n<h3>Comparision</h3>\n<p>The true purpose of this competition is not to explore strategies which have better Precision and Recall, but to obtain better MAP@12 for submission, so I compared each strategy based on MAP@12.</p>\n<table>\n<thead>\n<tr>\n<th>Strategy</th>\n<th>MAP@12 change</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>trending</td>\n<td>base lines</td>\n</tr>\n<tr>\n<td>recently</td>\n<td>base lines</td>\n</tr>\n<tr>\n<td>popular</td>\n<td>base lines</td>\n</tr>\n<tr>\n<td>last week</td>\n<td>-0.000131</td>\n</tr>\n<tr>\n<td>t-SNE</td>\n<td>+0.00053</td>\n</tr>\n<tr>\n<td>umap</td>\n<td>-0.0001</td>\n</tr>\n<tr>\n<td>t-SNE &amp; umap</td>\n<td>-0.000898</td>\n</tr>\n<tr>\n<td>also bought this</td>\n<td>+0.000533</td>\n</tr>\n<tr>\n<td>albert</td>\n<td>+0.000241</td>\n</tr>\n<tr>\n<td>collaboration</td>\n<td>+0.000134</td>\n</tr>\n<tr>\n<td>Bert</td>\n<td>+0.000051 (didn't correlate with LB)</td>\n</tr>\n<tr>\n<td>sequence Word2Vec w/ compression</td>\n<td>+0.000367</td>\n</tr>\n<tr>\n<td>bin popular</td>\n<td>+0.00058</td>\n</tr>\n<tr>\n<td>bert4rec</td>\n<td>-0.00037</td>\n</tr>\n<tr>\n<td>gru4rec</td>\n<td>-0.000347</td>\n</tr>\n<tr>\n<td>Implicit ALS</td>\n<td>+0.000142</td>\n</tr>\n<tr>\n<td>sequence Word2Vec</td>\n<td>-0.000563</td>\n</tr>\n<tr>\n<td>VGG</td>\n<td>-0.000375</td>\n</tr>\n</tbody>\n</table>\n<p>Note: This is the PROCESS of adding or subtracting the strategies rather than the result of comparision experiment.</p>\n<h2>Computing Resources</h2>\n<ul>\n<li>Untill 2 weeks from deadlines: MacBook Air (M1, 2020) (8 core CPU, 16 GM RAM, no GPU for deep learning).</li>\n<li>Last 2 weeks: GCP instances to deal with more candidates. (48 core CPU, 384 GB RAM)</li>\n</ul>\n<h2>Some Story</h2>\n<p>This is my first medal in Kaggle, so I'm very happy with my place! </p>\n<p>I am new to a recommend problem, but my teammates helped me a lot, and <a href=\"https://www.kaggle.com/paweljankiewicz\" target=\"_blank\">@paweljankiewicz</a> and <a href=\"https://www.kaggle.com/lihaorocky\" target=\"_blank\">@lihaorocky</a> was so kind to answer my questions in discussion.</p>",
      "rawMarkdown": "# 63th Place Solution\n\nThanks to my teammates @ootake, @ryoichi0917, @xia464, especially @tyonemoto for inviting me.\n\n## Code (updated in 2022-07-12)\n\nI publish the entire code.\ncode: https://github.com/Y-Haneji/kaggle-handm\n\n## Last Submission\n\n- (lr=0.1) : Private LB: 0.02813, Public LB: 0.02826, Validation: 0.033623\n- (lr=0.02): Private LB: 0.02772, Public LB: 0.0277, Validation: 0.033814\n\n## Summary\n\nI think that the key points of this competition are these.\n\n- Stable validation.\n- Two stage: generating candidates by some strategies, then ranking them.\n- Better strategies and scores from them.\n- Using full information like not only transactions but also detail descriptions, images, and so on.\n\n## Ranking Model (me)\n\n### Model\n\nLightGBM lambdarank\n\n### Dataset\n\nThis competition gave no training datasets, only transactions datas and some metadatas. I needed to generate training datasets by myself. The dataset was like this.\n\n- customer_id\n- article_id (candidate)\n- score from generative strategy (may not just one score)\n- flags wheter the candidate included each strategy\n- other features of customers\n- other features of articles\n- label (0 or 1)\n\nThe features of customers and articles are purchase and sales of ...\n\n- each week\n- agregation (sum, mean, max, min) of all weeks\n- difference and division between each week and aggregation\n\nLabel is whether the article was bought target_week.\n\nThe details of strategies to generate candidates is written in below.\n\n### Training\n\n- 5 target weeks, this means 5 folds training dataset.\n- For training, I only generated candidates for the customers who had some purchases in the target week.\n- Model Validation: NDCG@12. The dataset of next target week was validation dataset. The last target week's training dataset had no validation datasets.\n- Submission Validation: MAP@12 for last week's transactions.\n\n<!-- TODO: CVの図を入れる -->\n\n### Parameter\n\n- objective: lambdarank\n- lr: 0.1 or 0.02\n- feature_fraction: 0.8\n- baggin_fraction: 0.8\n- early_stopping: 10 or 50 rounds (depend on lr)\n\nOther parameters is default. I tried Optuna but it didn't work well for lambdarank because it attempted to LOWER the MAP.\n\n### Prediction (Submission)\n\n- For prediction, I should generate candidates for all the customers, which causes a memory problem.\n- Split the customers into batches.\n- Ensemble 5 target week's models.\n- Generally take longer time than training\n- If validation was not good, I didn't compute submission.\n\n## Candidates Strategies\n\n### Worked (included last submission)\n\n- trending (original: [here](https://www.kaggle.com/code/byfone/h-m-trending-products-weekly))\n- recently purchased\n- bin popular\n- last week for each customer\n- also bought this (original: [here](https://www.kaggle.com/code/cdeotte/customers-who-bought-this-frequently-buy-this))\n- collabo\n- t-SNE KNN\n- Albert KNN\n- sequence Word2Vec\n\n### Not Worked\n\n- Bert KNN\n- VGG KNN\n- Gru4Rec, Bert4Rec (LSTM) (original: [here](https://www.kaggle.com/code/astrung/lstm-model-with-item-infor-fix-missing-last-item))\n- Implicit ALS\n\n### Details of Some Strategies\n\nTO BE UPDATED: Some strategies will be written in some days.\n\n#### bin popular (by me and by @tyonemoto)\n\nTo recommend popular item was effecitve especially for cold start customers. Different kind of customers should have different preference. So we split customers into some groups by their features, like age, sex, and so on.\n\nIn the experiment phase, we used age, price, and transaction channel for binning. The number of splitting is 10 for age, 20 for price, and 3 for channels. Recall of this strategy was 0.06750.\n\nBecause of the deadline, we have no choice but to adopt only age bins.\n\n#### collabo (by @tyonemoto)\n\nThe collaboration articles with Kangol was launched a few weeks ago from the submission week. We recommended popular collaboration articles to customers who had bought some collaboration articles. Recall of this strategy was 0.00094.\n\n#### t-SNE (by @tyonemoto)\n\nThe metadatas of articles from the file of articles.csv and embedding vectors of detail descriptions by tf-idf and bert are compressed to two dimensions by t-SNE. We used count encoding method to deal with the category metadatas columns.\n\n#### Albert (by @ryoichi0917)\n\n#### sequence Word2Vec (by @tyonemoto)\n\n#### Bert (by @tyonemoto and @ootake)\n\nWe used bert pretrained model named \"bert-uncased\", \"MRPC\", and \"SQuAD\" to extract features from articles' detail description.\n\n#### VGG (by me)\n\nI extracted features from images using VGG because their vectors were relatively low dimensions, which can avoid a memory problem. But even they reached 10GB files (when compressed to .npy).\nVGG KNN doesn't worked, but I wonder if I could use other bigger models to extract features.\n\n#### Implicit ALS (by @ootake)\n\n### Comparision\n\nThe true purpose of this competition is not to explore strategies which have better Precision and Recall, but to obtain better MAP@12 for submission, so I compared each strategy based on MAP@12.\n\n| Strategy | MAP@12 change |\n| --- | --- |\n| trending | base lines |\n| recently | base lines |\n| popular | base lines |\n| last week | -0.000131 |\n| t-SNE | +0.00053 |\n| umap | -0.0001 |\n| t-SNE & umap | -0.000898 |\n| also bought this | +0.000533 |\n| albert | +0.000241 |\n| collaboration | +0.000134 |\n| Bert | +0.000051 (didn't correlate with LB) |\n| sequence Word2Vec w/ compression | +0.000367 |\n| bin popular | +0.00058 |\n| bert4rec | -0.00037 |\n| gru4rec | -0.000347 |\n| Implicit ALS | +0.000142 |\n| sequence Word2Vec | -0.000563 |\n| VGG | -0.000375 |\n\nNote: This is the PROCESS of adding or subtracting the strategies rather than the result of comparision experiment.\n\n## Computing Resources\n\n- Untill 2 weeks from deadlines: MacBook Air (M1, 2020) (8 core CPU, 16 GM RAM, no GPU for deep learning).\n- Last 2 weeks: GCP instances to deal with more candidates. (48 core CPU, 384 GB RAM)\n\n## Some Story\n\nThis is my first medal in Kaggle, so I'm very happy with my place! \n\nI am new to a recommend problem, but my teammates helped me a lot, and @paweljankiewicz and @lihaorocky was so kind to answer my questions in discussion.",
      "votes": null
    },
    {
      "id": "1784101",
      "postDate": "05/10/2022 23:55:41",
      "content": "<p>Congratulations on your medal! This is an impressive and informative post. Your teammates seem to have been a great help, and it's great that you were able to learn from others in the community. Recommendation problems can be tough, but it sounds like you did a great job. Keep up the good work!</p>",
      "rawMarkdown": "Congratulations on your medal! This is an impressive and informative post. Your teammates seem to have been a great help, and it's great that you were able to learn from others in the community. Recommendation problems can be tough, but it sounds like you did a great job. Keep up the good work!",
      "votes": null
    },
    {
      "id": "1784510",
      "postDate": "05/11/2022 08:10:49",
      "content": "<p><a href=\"https://www.kaggle.com/satoshidatamoto\" target=\"_blank\">@satoshidatamoto</a> <br>\nThank you for your celebrating!</p>",
      "rawMarkdown": "satoshidatamoto \nThank you for your celebrating!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1784101,
      "author_name": "",
      "author_url": "",
      "post_date": "05/10/2022 23:55:41",
      "content": "<p>Congratulations on your medal! This is an impressive and informative post. Your teammates seem to have been a great help, and it's great that you were able to learn from others in the community. Recommendation problems can be tough, but it sounds like you did a great job. Keep up the good work!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1784510,
          "author_name": "hanejiyuto",
          "author_url": "",
          "post_date": "05/11/2022 08:10:49",
          "content": "<p><a href=\"https://www.kaggle.com/satoshidatamoto\" target=\"_blank\">@satoshidatamoto</a> <br>\nThank you for your celebrating!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1783657": "# 63th Place Solution\n\nThanks to my teammates @ootake, @ryoichi0917, @xia464, especially @tyonemoto for inviting me.\n\n## Code (updated in 2022-07-12)\n\nI publish the entire code.\ncode: https://github.com/Y-Haneji/kaggle-handm\n\n## Last Submission\n\n- (lr=0.1) : Private LB: 0.02813, Public LB: 0.02826, Validation: 0.033623\n- (lr=0.02): Private LB: 0.02772, Public LB: 0.0277, Validation: 0.033814\n\n## Summary\n\nI think that the key points of this competition are these.\n\n- Stable validation.\n- Two stage: generating candidates by some strategies, then ranking them.\n- Better strategies and scores from them.\n- Using full information like not only transactions but also detail descriptions, images, and so on.\n\n## Ranking Model (me)\n\n### Model\n\nLightGBM lambdarank\n\n### Dataset\n\nThis competition gave no training datasets, only transactions datas and some metadatas. I needed to generate training datasets by myself. The dataset was like this.\n\n- customer_id\n- article_id (candidate)\n- score from generative strategy (may not just one score)\n- flags wheter the candidate included each strategy\n- other features of customers\n- other features of articles\n- label (0 or 1)\n\nThe features of customers and articles are purchase and sales of ...\n\n- each week\n- agregation (sum, mean, max, min) of all weeks\n- difference and division between each week and aggregation\n\nLabel is whether the article was bought target_week.\n\nThe details of strategies to generate candidates is written in below.\n\n### Training\n\n- 5 target weeks, this means 5 folds training dataset.\n- For training, I only generated candidates for the customers who had some purchases in the target week.\n- Model Validation: NDCG@12. The dataset of next target week was validation dataset. The last target week's training dataset had no validation datasets.\n- Submission Validation: MAP@12 for last week's transactions.\n\n<!-- TODO: CVの図を入れる -->\n\n### Parameter\n\n- objective: lambdarank\n- lr: 0.1 or 0.02\n- feature_fraction: 0.8\n- baggin_fraction: 0.8\n- early_stopping: 10 or 50 rounds (depend on lr)\n\nOther parameters is default. I tried Optuna but it didn't work well for lambdarank because it attempted to LOWER the MAP.\n\n### Prediction (Submission)\n\n- For prediction, I should generate candidates for all the customers, which causes a memory problem.\n- Split the customers into batches.\n- Ensemble 5 target week's models.\n- Generally take longer time than training\n- If validation was not good, I didn't compute submission.\n\n## Candidates Strategies\n\n### Worked (included last submission)\n\n- trending (original: [here](https://www.kaggle.com/code/byfone/h-m-trending-products-weekly))\n- recently purchased\n- bin popular\n- last week for each customer\n- also bought this (original: [here](https://www.kaggle.com/code/cdeotte/customers-who-bought-this-frequently-buy-this))\n- collabo\n- t-SNE KNN\n- Albert KNN\n- sequence Word2Vec\n\n### Not Worked\n\n- Bert KNN\n- VGG KNN\n- Gru4Rec, Bert4Rec (LSTM) (original: [here](https://www.kaggle.com/code/astrung/lstm-model-with-item-infor-fix-missing-last-item))\n- Implicit ALS\n\n### Details of Some Strategies\n\nTO BE UPDATED: Some strategies will be written in some days.\n\n#### bin popular (by me and by @tyonemoto)\n\nTo recommend popular item was effecitve especially for cold start customers. Different kind of customers should have different preference. So we split customers into some groups by their features, like age, sex, and so on.\n\nIn the experiment phase, we used age, price, and transaction channel for binning. The number of splitting is 10 for age, 20 for price, and 3 for channels. Recall of this strategy was 0.06750.\n\nBecause of the deadline, we have no choice but to adopt only age bins.\n\n#### collabo (by @tyonemoto)\n\nThe collaboration articles with Kangol was launched a few weeks ago from the submission week. We recommended popular collaboration articles to customers who had bought some collaboration articles. Recall of this strategy was 0.00094.\n\n#### t-SNE (by @tyonemoto)\n\nThe metadatas of articles from the file of articles.csv and embedding vectors of detail descriptions by tf-idf and bert are compressed to two dimensions by t-SNE. We used count encoding method to deal with the category metadatas columns.\n\n#### Albert (by @ryoichi0917)\n\n#### sequence Word2Vec (by @tyonemoto)\n\n#### Bert (by @tyonemoto and @ootake)\n\nWe used bert pretrained model named \"bert-uncased\", \"MRPC\", and \"SQuAD\" to extract features from articles' detail description.\n\n#### VGG (by me)\n\nI extracted features from images using VGG because their vectors were relatively low dimensions, which can avoid a memory problem. But even they reached 10GB files (when compressed to .npy).\nVGG KNN doesn't worked, but I wonder if I could use other bigger models to extract features.\n\n#### Implicit ALS (by @ootake)\n\n### Comparision\n\nThe true purpose of this competition is not to explore strategies which have better Precision and Recall, but to obtain better MAP@12 for submission, so I compared each strategy based on MAP@12.\n\n| Strategy | MAP@12 change |\n| --- | --- |\n| trending | base lines |\n| recently | base lines |\n| popular | base lines |\n| last week | -0.000131 |\n| t-SNE | +0.00053 |\n| umap | -0.0001 |\n| t-SNE & umap | -0.000898 |\n| also bought this | +0.000533 |\n| albert | +0.000241 |\n| collaboration | +0.000134 |\n| Bert | +0.000051 (didn't correlate with LB) |\n| sequence Word2Vec w/ compression | +0.000367 |\n| bin popular | +0.00058 |\n| bert4rec | -0.00037 |\n| gru4rec | -0.000347 |\n| Implicit ALS | +0.000142 |\n| sequence Word2Vec | -0.000563 |\n| VGG | -0.000375 |\n\nNote: This is the PROCESS of adding or subtracting the strategies rather than the result of comparision experiment.\n\n## Computing Resources\n\n- Untill 2 weeks from deadlines: MacBook Air (M1, 2020) (8 core CPU, 16 GM RAM, no GPU for deep learning).\n- Last 2 weeks: GCP instances to deal with more candidates. (48 core CPU, 384 GB RAM)\n\n## Some Story\n\nThis is my first medal in Kaggle, so I'm very happy with my place! \n\nI am new to a recommend problem, but my teammates helped me a lot, and @paweljankiewicz and @lihaorocky was so kind to answer my questions in discussion.",
    "1784101": "Congratulations on your medal! This is an impressive and informative post. Your teammates seem to have been a great help, and it's great that you were able to learn from others in the community. Recommendation problems can be tough, but it sounds like you did a great job. Keep up the good work!",
    "1784510": "satoshidatamoto \nThank you for your celebrating!"
  },
  "source": "meta"
}