{
  "id": 382935,
  "title": "5th place (yet) solution (@EEHITer & @wj19971997)",
  "url": "/competitions/otto-recommender-system/discussion/382935",
  "author_name": "EEHITer",
  "post_date": "2023-02-01T15:43:36.421000",
  "votes": 16,
  "comment_count": 0,
  "views": 0,
  "content": "<p>First, thanks all my teammates, without their efforts we could not achieve the final results.<br>\nThanks <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, you provide a good baseline and easy to follow.<br>\nAs <a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">@carnozhao</a> mentioned, our final team consists of two teams, the solution I'd like to share is contributed by me and <a href=\"https://www.kaggle.com/wj19971997\" target=\"_blank\">@wj19971997</a>.<br>\nFor evaluation metrics, the types of orders and carts have a large weight, so our optimization is mainly focused on these two targets and we generate the same candidates for both types.</p>\n<h1>Cross Validation</h1>\n<p>Our local validation is obtained by <a href=\"https://www.kaggle.com/https\" target=\"_blank\">@https</a>://github.com/otto-de/recsys-dataset，the last week of data is adopted as the local test set and the first three weeks are used as the local training set. </p>\n<h1>Matching</h1>\n<ul>\n<li>Co-visitation matrix and re-interaction items provided <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> ：TOP-150<ul>\n<li>Buy2buy Co-visitation;</li>\n<li>Co-clicks Co-visitation;</li>\n<li>Carts-orders Co-visitation;</li></ul></li>\n<li>Item-CF: TOP-100;</li>\n<li>Embedding-based Retrieval U2I：TOP-50;</li>\n<li>Embedding-based Retrieval I2I：TOP-50;</li>\n<li>Bi-Graph: TOP-50;</li>\n</ul>\n<p>After merging the five recall methods, the average number of our candidate set is 261 and the max is 400. <br>\nThe score of top400@recall for each target is：</p>\n<ul>\n<li>types[orders], recall[0.739558]</li>\n<li>types[carts], recall[0.567205]</li>\n</ul>\n<h1>Feature Engineering</h1>\n<p>our features can be mainly divided into four categories：</p>\n<ul>\n<li>recall score and rank-index of each retrieval；</li>\n<li>Session features：<ul>\n<li>aid count/nunique group by session;</li>\n<li>aid ranks sorted by time in history interaction;</li>\n<li>the latest k interaction types;</li></ul></li>\n<li>Aid features:<ul>\n<li>generate from past 3 days/1 week/2 weeks/3 weeks</li>\n<li>session count/nunqiue;</li>\n<li>session count by each type;</li>\n<li>type statistics (mean/std/min/max)；</li>\n<li>the interaction of different types of statistics features;</li></ul></li>\n<li>Interaction features of session-aid pairs：<ul>\n<li>aid count/nunique group by ['session', 'aid'];</li>\n<li>type count/nunique/mean/sum group by ['session', 'aid'];</li>\n<li>similarity features based emebeddings：we adopt several models to generate different item embeddings and aggregated them to the session in different ways, then we calculate the cosine similarity via dot product;<ul>\n<li>Models：Word2vec; DeepWalk;</li>\n<li>Aggregation methods：<ul>\n<li>History items of all interactions;</li>\n<li>latest k interaction;</li>\n<li>History interaction by only orders/carts/clicks/orders+carts；</li></ul></li></ul></li>\n<li>similarity features generated by similarity matrix: <ul>\n<li>we query each similarity dictionary to obtain the similarity between the latest three items and the candidate item;</li>\n<li>For each session, we can obtain a similarity sequence between history aid and candidate aid via the similarity matrix, then we calculate the mean/std/max/min/sum of the sequence.<ul>\n<li>co-visitation matrix：buy2buy、carts2carts、carts2orders、carts2clicks、clicks2carts、clicks2orders、clicks2clicks、orders2orders、orders2carts、orders2clicks;</li>\n<li>Item-CF simlarity matrix;</li>\n<li>Bi-Graph similarity matrix;</li></ul></li></ul></li></ul></li>\n</ul>\n<h1>Predict</h1>\n<p>We adopt XGBoost and Catboost Ranker as our predict models, each model is 5-FOLD group by session;</p>\n<p>The loss function of XGBoost is log loss, and YetiRankPairwise is for Catboost Ranker.</p>\n<p>Our negative sampling rate is 0.1 for orders and 0.2 for carts.</p>\n<p>our single model in public is 0.601 without <a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">@carnozhao</a> 's features, by using the selected features they provided, our single model score in public achieved 0.603.</p>\n<h1>Some attempts but don't work</h1>\n<ul>\n<li>GNN:\nGNN can learn the multi-hop neighbor information on the graph. I think it can provide additional information for the predict model, but it doesn’t work. Maybe the graph-construction method, sampling method, and aggregation algorithm need to be fine-tuned. Following are our gnn details:<ul>\n<li>construct isomorphic graphs with historical interaction data;</li>\n<li>pairwise loss &amp; Inductive learning;</li></ul></li>\n<li>NN predict model；</li>\n</ul>\n<hr>\n<p>Finally, thanks again for all my teammates;</p>",
  "messages": [
    {
      "id": 2125339,
      "postDate": "2023-02-01T15:43:36.420Z",
      "content": "<p>First, thanks all my teammates, without their efforts we could not achieve the final results.<br>\nThanks <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, you provide a good baseline and easy to follow.<br>\nAs <a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">@carnozhao</a> mentioned, our final team consists of two teams, the solution I'd like to share is contributed by me and <a href=\"https://www.kaggle.com/wj19971997\" target=\"_blank\">@wj19971997</a>.<br>\nFor evaluation metrics, the types of orders and carts have a large weight, so our optimization is mainly focused on these two targets and we generate the same candidates for both types.</p>\n<h1>Cross Validation</h1>\n<p>Our local validation is obtained by <a href=\"https://www.kaggle.com/https\" target=\"_blank\">@https</a>://github.com/otto-de/recsys-dataset，the last week of data is adopted as the local test set and the first three weeks are used as the local training set. </p>\n<h1>Matching</h1>\n<ul>\n<li>Co-visitation matrix and re-interaction items provided <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> ：TOP-150<ul>\n<li>Buy2buy Co-visitation;</li>\n<li>Co-clicks Co-visitation;</li>\n<li>Carts-orders Co-visitation;</li></ul></li>\n<li>Item-CF: TOP-100;</li>\n<li>Embedding-based Retrieval U2I：TOP-50;</li>\n<li>Embedding-based Retrieval I2I：TOP-50;</li>\n<li>Bi-Graph: TOP-50;</li>\n</ul>\n<p>After merging the five recall methods, the average number of our candidate set is 261 and the max is 400. <br>\nThe score of top400@recall for each target is：</p>\n<ul>\n<li>types[orders], recall[0.739558]</li>\n<li>types[carts], recall[0.567205]</li>\n</ul>\n<h1>Feature Engineering</h1>\n<p>our features can be mainly divided into four categories：</p>\n<ul>\n<li>recall score and rank-index of each retrieval；</li>\n<li>Session features：<ul>\n<li>aid count/nunique group by session;</li>\n<li>aid ranks sorted by time in history interaction;</li>\n<li>the latest k interaction types;</li></ul></li>\n<li>Aid features:<ul>\n<li>generate from past 3 days/1 week/2 weeks/3 weeks</li>\n<li>session count/nunqiue;</li>\n<li>session count by each type;</li>\n<li>type statistics (mean/std/min/max)；</li>\n<li>the interaction of different types of statistics features;</li></ul></li>\n<li>Interaction features of session-aid pairs：<ul>\n<li>aid count/nunique group by ['session', 'aid'];</li>\n<li>type count/nunique/mean/sum group by ['session', 'aid'];</li>\n<li>similarity features based emebeddings：we adopt several models to generate different item embeddings and aggregated them to the session in different ways, then we calculate the cosine similarity via dot product;<ul>\n<li>Models：Word2vec; DeepWalk;</li>\n<li>Aggregation methods：<ul>\n<li>History items of all interactions;</li>\n<li>latest k interaction;</li>\n<li>History interaction by only orders/carts/clicks/orders+carts；</li></ul></li></ul></li>\n<li>similarity features generated by similarity matrix: <ul>\n<li>we query each similarity dictionary to obtain the similarity between the latest three items and the candidate item;</li>\n<li>For each session, we can obtain a similarity sequence between history aid and candidate aid via the similarity matrix, then we calculate the mean/std/max/min/sum of the sequence.<ul>\n<li>co-visitation matrix：buy2buy、carts2carts、carts2orders、carts2clicks、clicks2carts、clicks2orders、clicks2clicks、orders2orders、orders2carts、orders2clicks;</li>\n<li>Item-CF simlarity matrix;</li>\n<li>Bi-Graph similarity matrix;</li></ul></li></ul></li></ul></li>\n</ul>\n<h1>Predict</h1>\n<p>We adopt XGBoost and Catboost Ranker as our predict models, each model is 5-FOLD group by session;</p>\n<p>The loss function of XGBoost is log loss, and YetiRankPairwise is for Catboost Ranker.</p>\n<p>Our negative sampling rate is 0.1 for orders and 0.2 for carts.</p>\n<p>our single model in public is 0.601 without <a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">@carnozhao</a> 's features, by using the selected features they provided, our single model score in public achieved 0.603.</p>\n<h1>Some attempts but don't work</h1>\n<ul>\n<li>GNN:\nGNN can learn the multi-hop neighbor information on the graph. I think it can provide additional information for the predict model, but it doesn’t work. Maybe the graph-construction method, sampling method, and aggregation algorithm need to be fine-tuned. Following are our gnn details:<ul>\n<li>construct isomorphic graphs with historical interaction data;</li>\n<li>pairwise loss &amp; Inductive learning;</li></ul></li>\n<li>NN predict model；</li>\n</ul>\n<hr>\n<p>Finally, thanks again for all my teammates;</p>",
      "rawMarkdown": "First, thanks all my teammates, without their efforts we could not achieve the final results.\nThanks @cdeotte, you provide a good baseline and easy to follow.\nAs @carnozhao mentioned, our final team consists of two teams, the solution I'd like to share is contributed by me and @wj19971997.\nFor evaluation metrics, the types of orders and carts have a large weight, so our optimization is mainly focused on these two targets and we generate the same candidates for both types.\n\n# Cross Validation\nOur local validation is obtained by @https://github.com/otto-de/recsys-dataset，the last week of data is adopted as the local test set and the first three weeks are used as the local training set. \n# Matching\n- Co-visitation matrix and re-interaction items provided @cdeotte ：TOP-150\n - Buy2buy Co-visitation;\n  - Co-clicks Co-visitation;\n  - Carts-orders Co-visitation;\n- Item-CF: TOP-100;\n- Embedding-based Retrieval U2I：TOP-50;\n- Embedding-based Retrieval I2I：TOP-50;\n- Bi-Graph: TOP-50;\n\nAfter merging the five recall methods, the average number of our candidate set is 261 and the max is 400. \nThe score of top400@recall for each target is：\n- types[orders], recall[0.739558]\n- types[carts], recall[0.567205]\n# Feature Engineering\nour features can be mainly divided into four categories：\n- recall score and rank-index of each retrieval；\n- Session features：\n  - aid count/nunique group by session;\n  - aid ranks sorted by time in history interaction;\n  - the latest k interaction types;\n- Aid features:\n  - generate from past 3 days/1 week/2 weeks/3 weeks\n    - session count/nunqiue;\n    - session count by each type;\n    - type statistics (mean/std/min/max)；\n    - the interaction of different types of statistics features;\n- Interaction features of session-aid pairs：\n  - aid count/nunique group by ['session', 'aid'];\n  - type count/nunique/mean/sum group by ['session', 'aid'];\n  - similarity features based emebeddings：we adopt several models to generate different item embeddings and aggregated them to the session in different ways, then we calculate the cosine similarity via dot product;\n     - Models：Word2vec; DeepWalk;\n     - Aggregation methods：\n         - History items of all interactions;\n         - latest k interaction;\n         - History interaction by only orders/carts/clicks/orders+carts；\n  - similarity features generated by similarity matrix: \n      - we query each similarity dictionary to obtain the similarity between the latest three items and the candidate item;\n      - For each session, we can obtain a similarity sequence between history aid and candidate aid via the similarity matrix, then we calculate the mean/std/max/min/sum of the sequence.\n         - co-visitation matrix：buy2buy、carts2carts、carts2orders、carts2clicks、clicks2carts、clicks2orders、clicks2clicks、orders2orders、orders2carts、orders2clicks;\n         - Item-CF simlarity matrix;\n         - Bi-Graph similarity matrix;\n# Predict \n\nWe adopt XGBoost and Catboost Ranker as our predict models, each model is 5-FOLD group by session;\n\nThe loss function of XGBoost is log loss, and YetiRankPairwise is for Catboost Ranker.\n\nOur negative sampling rate is 0.1 for orders and 0.2 for carts.\n\nour single model in public is 0.601 without @carnozhao 's features, by using the selected features they provided, our single model score in public achieved 0.603.\n# Some attempts but don't work\n- GNN:\nGNN can learn the multi-hop neighbor information on the graph. I think it can provide additional information for the predict model, but it doesn’t work. Maybe the graph-construction method, sampling method, and aggregation algorithm need to be fine-tuned. Following are our gnn details:\n    - construct isomorphic graphs with historical interaction data;\n    - pairwise loss & Inductive learning;\n- NN predict model；\n\n\n----------------------------------------\n\nFinally, thanks again for all my teammates;",
      "votes": 16
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2125339": "First, thanks all my teammates, without their efforts we could not achieve the final results.\nThanks @cdeotte, you provide a good baseline and easy to follow.\nAs @carnozhao mentioned, our final team consists of two teams, the solution I'd like to share is contributed by me and @wj19971997.\nFor evaluation metrics, the types of orders and carts have a large weight, so our optimization is mainly focused on these two targets and we generate the same candidates for both types.\n\n# Cross Validation\nOur local validation is obtained by @https://github.com/otto-de/recsys-dataset，the last week of data is adopted as the local test set and the first three weeks are used as the local training set. \n# Matching\n- Co-visitation matrix and re-interaction items provided @cdeotte ：TOP-150\n - Buy2buy Co-visitation;\n  - Co-clicks Co-visitation;\n  - Carts-orders Co-visitation;\n- Item-CF: TOP-100;\n- Embedding-based Retrieval U2I：TOP-50;\n- Embedding-based Retrieval I2I：TOP-50;\n- Bi-Graph: TOP-50;\n\nAfter merging the five recall methods, the average number of our candidate set is 261 and the max is 400. \nThe score of top400@recall for each target is：\n- types[orders], recall[0.739558]\n- types[carts], recall[0.567205]\n# Feature Engineering\nour features can be mainly divided into four categories：\n- recall score and rank-index of each retrieval；\n- Session features：\n  - aid count/nunique group by session;\n  - aid ranks sorted by time in history interaction;\n  - the latest k interaction types;\n- Aid features:\n  - generate from past 3 days/1 week/2 weeks/3 weeks\n    - session count/nunqiue;\n    - session count by each type;\n    - type statistics (mean/std/min/max)；\n    - the interaction of different types of statistics features;\n- Interaction features of session-aid pairs：\n  - aid count/nunique group by ['session', 'aid'];\n  - type count/nunique/mean/sum group by ['session', 'aid'];\n  - similarity features based emebeddings：we adopt several models to generate different item embeddings and aggregated them to the session in different ways, then we calculate the cosine similarity via dot product;\n     - Models：Word2vec; DeepWalk;\n     - Aggregation methods：\n         - History items of all interactions;\n         - latest k interaction;\n         - History interaction by only orders/carts/clicks/orders+carts；\n  - similarity features generated by similarity matrix: \n      - we query each similarity dictionary to obtain the similarity between the latest three items and the candidate item;\n      - For each session, we can obtain a similarity sequence between history aid and candidate aid via the similarity matrix, then we calculate the mean/std/max/min/sum of the sequence.\n         - co-visitation matrix：buy2buy、carts2carts、carts2orders、carts2clicks、clicks2carts、clicks2orders、clicks2clicks、orders2orders、orders2carts、orders2clicks;\n         - Item-CF simlarity matrix;\n         - Bi-Graph similarity matrix;\n# Predict \n\nWe adopt XGBoost and Catboost Ranker as our predict models, each model is 5-FOLD group by session;\n\nThe loss function of XGBoost is log loss, and YetiRankPairwise is for Catboost Ranker.\n\nOur negative sampling rate is 0.1 for orders and 0.2 for carts.\n\nour single model in public is 0.601 without @carnozhao 's features, by using the selected features they provided, our single model score in public achieved 0.603.\n# Some attempts but don't work\n- GNN:\nGNN can learn the multi-hop neighbor information on the graph. I think it can provide additional information for the predict model, but it doesn’t work. Maybe the graph-construction method, sampling method, and aggregation algorithm need to be fine-tuned. Following are our gnn details:\n    - construct isomorphic graphs with historical interaction data;\n    - pairwise loss & Inductive learning;\n- NN predict model；\n\n\n----------------------------------------\n\nFinally, thanks again for all my teammates;"
  }
}