{
  "id": 382834,
  "title": "12th place solution",
  "url": "/competitions/otto-recommender-system/writeups/buumoo-12th-place-solution",
  "author_name": "",
  "post_date": "2023-02-10T04:14:57.180Z",
  "votes": 37,
  "comment_count": 17,
  "views": 0,
  "content": "<h3>1. Retrieval Stage</h3>\n<p>• Re_visitition:<br>\nAll past events<br>\n   • Co_visitition:<br>\nTop_150 of [ Top_140(type_weighted) + Top_140(buy2buy) + Top_50(time_weighted) ]</p>\n<p>• SRGNN:<br>\nTop_172<br>\nDNN model can boost the retriever recall significantly in my solution by using even slightly smaller average candidates number.<br>\nThe recall here indicates recall@full_candidates</p>\n<table>\n<thead>\n<tr>\n<th>retrieval_strategy</th>\n<th>recall_clicks</th>\n<th>recall_carts</th>\n<th>recall_orders</th>\n<th>recall_weighted_sum</th>\n<th>average candidates number per session</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>all_past_events | top_258_co_visitition</td>\n<td>0.697</td>\n<td>0.555</td>\n<td>0.732</td>\n<td>0.675</td>\n<td>236.47</td>\n</tr>\n<tr>\n<td>all_past_events | top_150_co_visitition | top_169_CORE_trm</td>\n<td>0.723</td>\n<td>0.575</td>\n<td>0.742</td>\n<td>0.690</td>\n<td>234.26</td>\n</tr>\n<tr>\n<td>all_past_events | top_150_co_visitition | top_172_SRGNN</td>\n<td>0.734</td>\n<td>0.581</td>\n<td>0.746</td>\n<td>0.695</td>\n<td>234.07</td>\n</tr>\n<tr>\n<td>Combine multiple DNN models does not work for me, single SRGNN performs best under dozens of experiments.</td>\n<td></td>\n<td></td>\n<td></td>\n<td></td>\n<td></td>\n</tr>\n</tbody>\n</table>\n<h4>Data preparation &amp; augmentation for SRGNN:</h4>\n<p>Delete items appearing less than 3 times.<br>\ntrain_set: 3rd_week + truncated_4th_week_former_part<br>\nvalid_set: truncated_4th_week_latter_part (with former part as basis)</p>\n<h5>Augmentattion:</h5>\n<p>Train_set: For each sequence, iteratively seperate last item as target, truncate or pad the rest to max_length of 50 as training sequence, until only 1 item left.<br>\nValid_set: For each sequence, use truncated_4th_week_former_part as basis, then iteratively concatenate truncated_4th_week_latter_part as target, until last item in the latter part, as illustrated below:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4993592%2F081097ff55a958dc4326dc27ed76cbbb%2FMind%20map%20(1).png?generation=1675234390597141&amp;alt=media\" alt=\"\"><br>\nAs in test case, we have former part as basis for prediction, so I used this method to best mimic this scenario. It increased the validation samples as well compared with using latter part only, in which case the first item of latter part can not be used as target.<br>\nThe validation metric is recall@20, which correlates very well with my retrieval strategy's final recall across all experiments &amp; DNN models.<br>\nI did not try other augmentations, such as Swapping, Shuffling, etc…, simply because the data scale is already considered huge for my machine.</p>\n<h3>2. Ranking Stage</h3>\n<h4>CV strategy for ranking:</h4>\n<p>Using only the 4th week as training data for forming the session-item pairs.<br>\nFor each type: clicks/carts/orders, drop all sessions without positve targets, then do GroupKfold for each, respectively.<br>\nNegative sample downsampling: positive:negative = 1:20</p>\n<h4>Feature Engineering:</h4>\n<ol>\n<li><p>Session features: statistics about session: <br>\nsession_clicks/carts/orders_count, </p>\n<p>session_length</p>\n<p>session_items_nunique</p>\n<p>session_duration </p>\n<p>etc…</p></li>\n<li><p>Item features: statistics about item:<br>\nsimple statistics</p>\n<p>item_total_count</p>\n<p>item_count as clicks/carts/orders</p>\n<p>clicks/carts/orders_counts in each week</p>\n<p>clicks/carts/orders_counts changing trends in temporal order, weekly basis</p>\n<p>clicks/carts/orders_unique_session_counts in each week</p>\n<p>clicks/carts/orders_unique_session_counts changing trends in temporal order, weekly basis</p>\n<p>etc…</p></li>\n<li><p>Session-item interation features:<br>\nitem appearing count in session</p>\n<p>item appearing as clicks/carts/orders count in session</p>\n<p>item recent 1 hour/day/week clicks/carts/orders count from prediction timestamp</p>\n<p>relative time gap between the item clicked/carted/ordered last time and prediction timestamp</p>\n<p>co_visition weights(time weight, type weight, just count) for items in each session</p>\n<p>dnn model cosine similaries between session and item embeddings</p>\n<p>etc…</p></li>\n</ol>\n<h4>Models used:</h4>\n<p>Lgbm, Catboost, Xgboost, MLP<br>\nClassification and ranking algorithm perform similarly in my solution.<br>\nAnd ensembling classification and ranking does not provide any boost.<br>\nSo only binary classification is used at the end.</p>\n<h4>Ensemble:</h4>\n<p>Blending of 4 models, each with 5 folds.<br>\nStacking with meta model leads to a silimar performance, thus only blending is used.</p>\n<h4>Post processing:</h4>\n<p>There are 3327 sessions with less than 20 candidates.<br>\nSo I re-calculated the co_visitition items for these sessions without constraining the time range, namely with full data.<br>\nAfter this supplement, the average candidates number is above 19.<br>\nUsing hottest items within half day around the prediction timestamp to supplement the last drop.</p>\n<h3>Acknowledgement</h3>\n<p>Thank <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>  for introducing polars to the community, which saves me in this competition with limited ram.<br>\nIts lazy evalution and native multithreading support make it such a nice Lib on CPU platform.<br>\nThank <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> for sharing the knowledge and valuable notebooks as always.<br>\nThank all guys who shared their knowledge and insights in notebooks and discussions. Learned a lot from you guys.</p>",
  "messages": [
    {
      "id": "2124709",
      "postDate": "02/01/2023 06:57:44",
      "content": "<h3>1. Retrieval Stage</h3>\n<p>• Re_visitition:<br>\nAll past events<br>\n   • Co_visitition:<br>\nTop_150 of [ Top_140(type_weighted) + Top_140(buy2buy) + Top_50(time_weighted) ]</p>\n<p>• SRGNN:<br>\nTop_172<br>\nDNN model can boost the retriever recall significantly in my solution by using even slightly smaller average candidates number.<br>\nThe recall here indicates recall@full_candidates</p>\n<table>\n<thead>\n<tr>\n<th>retrieval_strategy</th>\n<th>recall_clicks</th>\n<th>recall_carts</th>\n<th>recall_orders</th>\n<th>recall_weighted_sum</th>\n<th>average candidates number per session</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>all_past_events | top_258_co_visitition</td>\n<td>0.697</td>\n<td>0.555</td>\n<td>0.732</td>\n<td>0.675</td>\n<td>236.47</td>\n</tr>\n<tr>\n<td>all_past_events | top_150_co_visitition | top_169_CORE_trm</td>\n<td>0.723</td>\n<td>0.575</td>\n<td>0.742</td>\n<td>0.690</td>\n<td>234.26</td>\n</tr>\n<tr>\n<td>all_past_events | top_150_co_visitition | top_172_SRGNN</td>\n<td>0.734</td>\n<td>0.581</td>\n<td>0.746</td>\n<td>0.695</td>\n<td>234.07</td>\n</tr>\n<tr>\n<td>Combine multiple DNN models does not work for me, single SRGNN performs best under dozens of experiments.</td>\n<td></td>\n<td></td>\n<td></td>\n<td></td>\n<td></td>\n</tr>\n</tbody>\n</table>\n<h4>Data preparation &amp; augmentation for SRGNN:</h4>\n<p>Delete items appearing less than 3 times.<br>\ntrain_set: 3rd_week + truncated_4th_week_former_part<br>\nvalid_set: truncated_4th_week_latter_part (with former part as basis)</p>\n<h5>Augmentattion:</h5>\n<p>Train_set: For each sequence, iteratively seperate last item as target, truncate or pad the rest to max_length of 50 as training sequence, until only 1 item left.<br>\nValid_set: For each sequence, use truncated_4th_week_former_part as basis, then iteratively concatenate truncated_4th_week_latter_part as target, until last item in the latter part, as illustrated below:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4993592%2F081097ff55a958dc4326dc27ed76cbbb%2FMind%20map%20(1).png?generation=1675234390597141&amp;alt=media\" alt=\"\"><br>\nAs in test case, we have former part as basis for prediction, so I used this method to best mimic this scenario. It increased the validation samples as well compared with using latter part only, in which case the first item of latter part can not be used as target.<br>\nThe validation metric is recall@20, which correlates very well with my retrieval strategy's final recall across all experiments &amp; DNN models.<br>\nI did not try other augmentations, such as Swapping, Shuffling, etc…, simply because the data scale is already considered huge for my machine.</p>\n<h3>2. Ranking Stage</h3>\n<h4>CV strategy for ranking:</h4>\n<p>Using only the 4th week as training data for forming the session-item pairs.<br>\nFor each type: clicks/carts/orders, drop all sessions without positve targets, then do GroupKfold for each, respectively.<br>\nNegative sample downsampling: positive:negative = 1:20</p>\n<h4>Feature Engineering:</h4>\n<ol>\n<li><p>Session features: statistics about session: <br>\nsession_clicks/carts/orders_count, </p>\n<p>session_length</p>\n<p>session_items_nunique</p>\n<p>session_duration </p>\n<p>etc…</p></li>\n<li><p>Item features: statistics about item:<br>\nsimple statistics</p>\n<p>item_total_count</p>\n<p>item_count as clicks/carts/orders</p>\n<p>clicks/carts/orders_counts in each week</p>\n<p>clicks/carts/orders_counts changing trends in temporal order, weekly basis</p>\n<p>clicks/carts/orders_unique_session_counts in each week</p>\n<p>clicks/carts/orders_unique_session_counts changing trends in temporal order, weekly basis</p>\n<p>etc…</p></li>\n<li><p>Session-item interation features:<br>\nitem appearing count in session</p>\n<p>item appearing as clicks/carts/orders count in session</p>\n<p>item recent 1 hour/day/week clicks/carts/orders count from prediction timestamp</p>\n<p>relative time gap between the item clicked/carted/ordered last time and prediction timestamp</p>\n<p>co_visition weights(time weight, type weight, just count) for items in each session</p>\n<p>dnn model cosine similaries between session and item embeddings</p>\n<p>etc…</p></li>\n</ol>\n<h4>Models used:</h4>\n<p>Lgbm, Catboost, Xgboost, MLP<br>\nClassification and ranking algorithm perform similarly in my solution.<br>\nAnd ensembling classification and ranking does not provide any boost.<br>\nSo only binary classification is used at the end.</p>\n<h4>Ensemble:</h4>\n<p>Blending of 4 models, each with 5 folds.<br>\nStacking with meta model leads to a silimar performance, thus only blending is used.</p>\n<h4>Post processing:</h4>\n<p>There are 3327 sessions with less than 20 candidates.<br>\nSo I re-calculated the co_visitition items for these sessions without constraining the time range, namely with full data.<br>\nAfter this supplement, the average candidates number is above 19.<br>\nUsing hottest items within half day around the prediction timestamp to supplement the last drop.</p>\n<h3>Acknowledgement</h3>\n<p>Thank <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>  for introducing polars to the community, which saves me in this competition with limited ram.<br>\nIts lazy evalution and native multithreading support make it such a nice Lib on CPU platform.<br>\nThank <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> for sharing the knowledge and valuable notebooks as always.<br>\nThank all guys who shared their knowledge and insights in notebooks and discussions. Learned a lot from you guys.</p>",
      "rawMarkdown": "### 1. Retrieval Stage\n\n    • Re_visitition:\n\tAll past events\n\n    • Co_visitition:\n\tTop_150 of [ Top_140(type_weighted) + Top_140(buy2buy) + Top_50(time_weighted) ]\n\t\n    • SRGNN:\n\tTop_172\n\n\tDNN model can boost the retriever recall significantly in my solution by using even slightly smaller average candidates number.\n\tThe recall here indicates recall@full_candidates\n\n| retrieval_strategy        | recall_clicks           | recall_carts  | recall_orders | recall_weighted_sum | average candidates number per session\n| ------------- |:-------------:| -----:|  -----:| -----:| -----:|\n| all_past_events \\| top_258_co_visitition  | 0.697| 0.555 | 0.732 | 0.675 | 236.47 |\n| all_past_events \\| top_150_co_visitition \\| top_169_CORE_trm     | 0.723   |   0.575| 0.742| 0.690 | 234.26 |\n| all_past_events \\| top_150_co_visitition \\| top_172_SRGNN     | 0.734  |   0.581 | 0.746 | 0.695| 234.07 |\n\nCombine multiple DNN models does not work for me, single SRGNN performs best under dozens of experiments.\n\n\n#### Data preparation & augmentation for SRGNN:\n\nDelete items appearing less than 3 times.\n\ntrain_set: 3rd_week + truncated_4th_week_former_part\n\nvalid_set: truncated_4th_week_latter_part (with former part as basis)\n\n##### Augmentattion: \nTrain_set: For each sequence, iteratively seperate last item as target, truncate or pad the rest to max_length of 50 as training sequence, until only 1 item left.\n\nValid_set: For each sequence, use truncated_4th_week_former_part as basis, then iteratively concatenate truncated_4th_week_latter_part as target, until last item in the latter part, as illustrated below:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4993592%2F081097ff55a958dc4326dc27ed76cbbb%2FMind%20map%20(1).png?generation=1675234390597141&alt=media)\n\nAs in test case, we have former part as basis for prediction, so I used this method to best mimic this scenario. It increased the validation samples as well compared with using latter part only, in which case the first item of latter part can not be used as target.\n\nThe validation metric is recall@20, which correlates very well with my retrieval strategy's final recall across all experiments & DNN models.\n\nI did not try other augmentations, such as Swapping, Shuffling, etc..., simply because the data scale is already considered huge for my machine.\n\n### 2. Ranking Stage\n\n#### CV strategy for ranking:\n\nUsing only the 4th week as training data for forming the session-item pairs.\nFor each type: clicks/carts/orders, drop all sessions without positve targets, then do GroupKfold for each, respectively.\n\nNegative sample downsampling: positive:negative = 1:20\n\n#### Feature Engineering:\n1. Session features: statistics about session: \n\n    session_clicks/carts/orders_count, \n    \n    session_length\n    \n    session_items_nunique\n    \n    session_duration \n    \n    etc...\n\n2. Item features: statistics about item:\n\n    simple statistics\n    \n    item_total_count\n    \n    item_count as clicks/carts/orders\n    \n    clicks/carts/orders_counts in each week\n    \n    clicks/carts/orders_counts changing trends in temporal order, weekly basis\n    \n    clicks/carts/orders_unique_session_counts in each week\n    \n    clicks/carts/orders_unique_session_counts changing trends in temporal order, weekly basis\n    \n    etc...\n\n3. Session-item interation features:\n\n    item appearing count in session\n    \n    item appearing as clicks/carts/orders count in session\n    \n    item recent 1 hour/day/week clicks/carts/orders count from prediction timestamp\n    \n    relative time gap between the item clicked/carted/ordered last time and prediction timestamp\n    \n    co_visition weights(time weight, type weight, just count) for items in each session\n    \n    dnn model cosine similaries between session and item embeddings\n    \n    etc...\n\n\n#### Models used:\nLgbm, Catboost, Xgboost, MLP\n\nClassification and ranking algorithm perform similarly in my solution.\n\nAnd ensembling classification and ranking does not provide any boost.\n\nSo only binary classification is used at the end.\n\n\n#### Ensemble:\nBlending of 4 models, each with 5 folds.\n\nStacking with meta model leads to a silimar performance, thus only blending is used.\n\n#### Post processing:\nThere are 3327 sessions with less than 20 candidates.\n\nSo I re-calculated the co_visitition items for these sessions without constraining the time range, namely with full data.\n\nAfter this supplement, the average candidates number is above 19.\n\nUsing hottest items within half day around the prediction timestamp to supplement the last drop.\n\n\n### Acknowledgement\n\nThank @radek1  for introducing polars to the community, which saves me in this competition with limited ram.\n\nIts lazy evalution and native multithreading support make it such a nice Lib on CPU platform.\n\nThank @cdeotte for sharing the knowledge and valuable notebooks as always.\n\nThank all guys who shared their knowledge and insights in notebooks and discussions. Learned a lot from you guys.",
      "votes": null
    },
    {
      "id": "2124746",
      "postDate": "02/01/2023 07:41:18",
      "content": "<p>Thanks a lot for your sharing. Hope you can share the config and source code for the SRGNN model in the near future. Graph is really efficient for single model.</p>",
      "rawMarkdown": "Thanks a lot for your sharing. Hope you can share the config and source code for the SRGNN model in the near future. Graph is really efficient for single model.",
      "votes": null
    },
    {
      "id": "2124882",
      "postDate": "02/01/2023 09:20:43",
      "content": "<p>Thanks. So u just do the plain cv and OOF prediction in Radek's dataset. I still wonder why it's seems you cv is more robust than others? </p>",
      "rawMarkdown": "Thanks. So u just do the plain cv and OOF prediction in Radek's dataset. I still wonder why it's seems you cv is more robust than others?",
      "votes": null
    },
    {
      "id": "2124899",
      "postDate": "02/01/2023 09:33:24",
      "content": "<p>No, I created my own dataset using host version 1 script, where there is unseen items in valid set. And it turns out to be not that robust. Only correlate with public LB well, the private LB severely shakes down.</p>",
      "rawMarkdown": "No, I created my own dataset using host version 1 script, where there is unseen items in valid set. And it turns out to be not that robust. Only correlate with public LB well, the private LB severely shakes down.",
      "votes": null
    },
    {
      "id": "2124975",
      "postDate": "02/01/2023 10:39:41",
      "content": "<p>SRGNN ! Great I really like Graph based solution. Personnally,  I used GraphSAGE and Node2Vec to generate embeddings from USERs and AIDs in order to compute the similarity score between them,I can say that using different approaches can help ! ( those features belong to the top 10 of my LGBM ) Unfortunately I didn't work much on candidates generation.</p>\n<p>I also noticed than using those graph features was enough. From my public and private scores, I can say that using 20 features ( 15 of them are graph based features ), gave me 0.579, same as using 120 features !</p>\n<p>Thanks for sharing <a href=\"https://www.kaggle.com/buumoo\" target=\"_blank\">@buumoo</a> </p>",
      "rawMarkdown": "SRGNN ! Great I really like Graph based solution. Personnally,  I used GraphSAGE and Node2Vec to generate embeddings from USERs and AIDs in order to compute the similarity score between them,I can say that using different approaches can help ! ( those features belong to the top 10 of my LGBM ) Unfortunately I didn't work much on candidates generation.\n\nI also noticed than using those graph features was enough. From my public and private scores, I can say that using 20 features ( 15 of them are graph based features ), gave me 0.579, same as using 120 features !\n\nThanks for sharing @buumoo",
      "votes": null
    },
    {
      "id": "2124983",
      "postDate": "02/01/2023 10:53:12",
      "content": "<p>Yes, your feeling is right. SRGNN similarity score is the top 1 feature in some of my models.</p>",
      "rawMarkdown": "Yes, your feeling is right. SRGNN similarity score is the top 1 feature in some of my models.",
      "votes": null
    },
    {
      "id": "2124995",
      "postDate": "02/01/2023 11:09:12",
      "content": "<p>can you further explain this :  \"Re_visitition: All past events\", with a little exemple ?</p>",
      "rawMarkdown": "can you further explain this :  \"Re_visitition: All past events\", with a little exemple ?",
      "votes": null
    },
    {
      "id": "2125001",
      "postDate": "02/01/2023 11:12:03",
      "content": "<p>Just use all items in one session as its candidates.</p>",
      "rawMarkdown": "Just use all items in one session as its candidates.",
      "votes": null
    },
    {
      "id": "2125022",
      "postDate": "02/01/2023 11:21:51",
      "content": "<p>Ok I got it ! I thought about another strategy. Yes you mean the unique aids of each session ! thanks for the insight</p>",
      "rawMarkdown": "Ok I got it ! I thought about another strategy. Yes you mean the unique aids of each session ! thanks for the insight",
      "votes": null
    },
    {
      "id": "2125087",
      "postDate": "02/01/2023 12:11:40",
      "content": "<p>Thanks for your sharing. I have not tried NN based solution, as think the item/aid is high cardinality, may I know if any special handling for your SRGNN?</p>",
      "rawMarkdown": "Thanks for your sharing. I have not tried NN based solution, as think the item/aid is high cardinality, may I know if any special handling for your SRGNN?",
      "votes": null
    },
    {
      "id": "2125094",
      "postDate": "02/01/2023 12:17:32",
      "content": "<p>Just modified the data processing part to accelerate training speed.<br>\nFilter out less frequent item to reduce cardinality.</p>",
      "rawMarkdown": "Just modified the data processing part to accelerate training speed.\nFilter out less frequent item to reduce cardinality.",
      "votes": null
    },
    {
      "id": "2126116",
      "postDate": "02/02/2023 05:24:31",
      "content": "<p>Congrats on your strong finish!</p>\n<p>The Devastator.</p>",
      "rawMarkdown": "Congrats on your strong finish!\n\n\nThe Devastator.",
      "votes": null
    },
    {
      "id": "2127538",
      "postDate": "02/03/2023 00:41:58",
      "content": "<p>solo gold medal! Congras.</p>",
      "rawMarkdown": "solo gold medal! Congras.",
      "votes": null
    },
    {
      "id": "2127642",
      "postDate": "02/03/2023 02:59:52",
      "content": "<p>Congratulations for u up to golden now. It's seems someone was dropped in top15.</p>",
      "rawMarkdown": "Congratulations for u up to golden now. It's seems someone was dropped in top15.",
      "votes": null
    },
    {
      "id": "2127813",
      "postDate": "02/03/2023 07:49:54",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/buumoo\" target=\"_blank\">@buumoo</a> ! A really well deserved solo gold - love your approach to candidate generation with SRGNN and data augmentation too (I definitely missed that trick!), and great features as well :)</p>",
      "rawMarkdown": "Congratulations @buumoo ! A really well deserved solo gold - love your approach to candidate generation with SRGNN and data augmentation too (I definitely missed that trick!), and great features as well :)",
      "votes": null
    },
    {
      "id": "2129334",
      "postDate": "02/04/2023 14:24:54",
      "content": "<p>Congratulations!!!! <a href=\"https://www.kaggle.com/buumoo\" target=\"_blank\">@buumoo</a> 👍 Hope someday I can get my solo gold just like you</p>",
      "rawMarkdown": "Congratulations!!!! @buumoo 👍 Hope someday I can get my solo gold just like you",
      "votes": null
    },
    {
      "id": "2129336",
      "postDate": "02/04/2023 14:28:35",
      "content": "<p>Also, could you plz share the feature importance?(if you ever did the <code>plot_importance()</code>)</p>",
      "rawMarkdown": "Also, could you plz share the feature importance?(if you ever did the `plot_importance()`)",
      "votes": null
    },
    {
      "id": "2134446",
      "postDate": "02/08/2023 01:48:51",
      "content": "<p>Thanks mate</p>",
      "rawMarkdown": "Thanks mate",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2124746,
      "author_name": "bibanh",
      "author_url": "",
      "post_date": "02/01/2023 07:41:18",
      "content": "<p>Thanks a lot for your sharing. Hope you can share the config and source code for the SRGNN model in the near future. Graph is really efficient for single model.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2124882,
      "author_name": "earlee0412",
      "author_url": "",
      "post_date": "02/01/2023 09:20:43",
      "content": "<p>Thanks. So u just do the plain cv and OOF prediction in Radek's dataset. I still wonder why it's seems you cv is more robust than others? </p>",
      "votes": null,
      "replies": [
        {
          "id": 2124899,
          "author_name": "buumoo",
          "author_url": "",
          "post_date": "02/01/2023 09:33:24",
          "content": "<p>No, I created my own dataset using host version 1 script, where there is unseen items in valid set. And it turns out to be not that robust. Only correlate with public LB well, the private LB severely shakes down.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2127642,
              "author_name": "earlee0412",
              "author_url": "",
              "post_date": "02/03/2023 02:59:52",
              "content": "<p>Congratulations for u up to golden now. It's seems someone was dropped in top15.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2124975,
      "author_name": "rayanaay",
      "author_url": "",
      "post_date": "02/01/2023 10:39:41",
      "content": "<p>SRGNN ! Great I really like Graph based solution. Personnally,  I used GraphSAGE and Node2Vec to generate embeddings from USERs and AIDs in order to compute the similarity score between them,I can say that using different approaches can help ! ( those features belong to the top 10 of my LGBM ) Unfortunately I didn't work much on candidates generation.</p>\n<p>I also noticed than using those graph features was enough. From my public and private scores, I can say that using 20 features ( 15 of them are graph based features ), gave me 0.579, same as using 120 features !</p>\n<p>Thanks for sharing <a href=\"https://www.kaggle.com/buumoo\" target=\"_blank\">@buumoo</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 2124983,
          "author_name": "buumoo",
          "author_url": "",
          "post_date": "02/01/2023 10:53:12",
          "content": "<p>Yes, your feeling is right. SRGNN similarity score is the top 1 feature in some of my models.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2124995,
              "author_name": "rayanaay",
              "author_url": "",
              "post_date": "02/01/2023 11:09:12",
              "content": "<p>can you further explain this :  \"Re_visitition: All past events\", with a little exemple ?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2125001,
                  "author_name": "buumoo",
                  "author_url": "",
                  "post_date": "02/01/2023 11:12:03",
                  "content": "<p>Just use all items in one session as its candidates.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2125022,
                      "author_name": "rayanaay",
                      "author_url": "",
                      "post_date": "02/01/2023 11:21:51",
                      "content": "<p>Ok I got it ! I thought about another strategy. Yes you mean the unique aids of each session ! thanks for the insight</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2125087,
      "author_name": "gongbi",
      "author_url": "",
      "post_date": "02/01/2023 12:11:40",
      "content": "<p>Thanks for your sharing. I have not tried NN based solution, as think the item/aid is high cardinality, may I know if any special handling for your SRGNN?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2125094,
          "author_name": "buumoo",
          "author_url": "",
          "post_date": "02/01/2023 12:17:32",
          "content": "<p>Just modified the data processing part to accelerate training speed.<br>\nFilter out less frequent item to reduce cardinality.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2126116,
      "author_name": "thedevastator",
      "author_url": "",
      "post_date": "02/02/2023 05:24:31",
      "content": "<p>Congrats on your strong finish!</p>\n<p>The Devastator.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2127538,
      "author_name": "hookman",
      "author_url": "",
      "post_date": "02/03/2023 00:41:58",
      "content": "<p>solo gold medal! Congras.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2127813,
      "author_name": "judehunt23",
      "author_url": "",
      "post_date": "02/03/2023 07:49:54",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/buumoo\" target=\"_blank\">@buumoo</a> ! A really well deserved solo gold - love your approach to candidate generation with SRGNN and data augmentation too (I definitely missed that trick!), and great features as well :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 2134446,
          "author_name": "buumoo",
          "author_url": "",
          "post_date": "02/08/2023 01:48:51",
          "content": "<p>Thanks mate</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2129334,
      "author_name": "yuzhang0422",
      "author_url": "",
      "post_date": "02/04/2023 14:24:54",
      "content": "<p>Congratulations!!!! <a href=\"https://www.kaggle.com/buumoo\" target=\"_blank\">@buumoo</a> 👍 Hope someday I can get my solo gold just like you</p>",
      "votes": null,
      "replies": [
        {
          "id": 2129336,
          "author_name": "yuzhang0422",
          "author_url": "",
          "post_date": "02/04/2023 14:28:35",
          "content": "<p>Also, could you plz share the feature importance?(if you ever did the <code>plot_importance()</code>)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2124709": "### 1. Retrieval Stage\n\n    • Re_visitition:\n\tAll past events\n\n    • Co_visitition:\n\tTop_150 of [ Top_140(type_weighted) + Top_140(buy2buy) + Top_50(time_weighted) ]\n\t\n    • SRGNN:\n\tTop_172\n\n\tDNN model can boost the retriever recall significantly in my solution by using even slightly smaller average candidates number.\n\tThe recall here indicates recall@full_candidates\n\n| retrieval_strategy        | recall_clicks           | recall_carts  | recall_orders | recall_weighted_sum | average candidates number per session\n| ------------- |:-------------:| -----:|  -----:| -----:| -----:|\n| all_past_events \\| top_258_co_visitition  | 0.697| 0.555 | 0.732 | 0.675 | 236.47 |\n| all_past_events \\| top_150_co_visitition \\| top_169_CORE_trm     | 0.723   |   0.575| 0.742| 0.690 | 234.26 |\n| all_past_events \\| top_150_co_visitition \\| top_172_SRGNN     | 0.734  |   0.581 | 0.746 | 0.695| 234.07 |\n\nCombine multiple DNN models does not work for me, single SRGNN performs best under dozens of experiments.\n\n\n#### Data preparation & augmentation for SRGNN:\n\nDelete items appearing less than 3 times.\n\ntrain_set: 3rd_week + truncated_4th_week_former_part\n\nvalid_set: truncated_4th_week_latter_part (with former part as basis)\n\n##### Augmentattion: \nTrain_set: For each sequence, iteratively seperate last item as target, truncate or pad the rest to max_length of 50 as training sequence, until only 1 item left.\n\nValid_set: For each sequence, use truncated_4th_week_former_part as basis, then iteratively concatenate truncated_4th_week_latter_part as target, until last item in the latter part, as illustrated below:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4993592%2F081097ff55a958dc4326dc27ed76cbbb%2FMind%20map%20(1).png?generation=1675234390597141&alt=media)\n\nAs in test case, we have former part as basis for prediction, so I used this method to best mimic this scenario. It increased the validation samples as well compared with using latter part only, in which case the first item of latter part can not be used as target.\n\nThe validation metric is recall@20, which correlates very well with my retrieval strategy's final recall across all experiments & DNN models.\n\nI did not try other augmentations, such as Swapping, Shuffling, etc..., simply because the data scale is already considered huge for my machine.\n\n### 2. Ranking Stage\n\n#### CV strategy for ranking:\n\nUsing only the 4th week as training data for forming the session-item pairs.\nFor each type: clicks/carts/orders, drop all sessions without positve targets, then do GroupKfold for each, respectively.\n\nNegative sample downsampling: positive:negative = 1:20\n\n#### Feature Engineering:\n1. Session features: statistics about session: \n\n    session_clicks/carts/orders_count, \n    \n    session_length\n    \n    session_items_nunique\n    \n    session_duration \n    \n    etc...\n\n2. Item features: statistics about item:\n\n    simple statistics\n    \n    item_total_count\n    \n    item_count as clicks/carts/orders\n    \n    clicks/carts/orders_counts in each week\n    \n    clicks/carts/orders_counts changing trends in temporal order, weekly basis\n    \n    clicks/carts/orders_unique_session_counts in each week\n    \n    clicks/carts/orders_unique_session_counts changing trends in temporal order, weekly basis\n    \n    etc...\n\n3. Session-item interation features:\n\n    item appearing count in session\n    \n    item appearing as clicks/carts/orders count in session\n    \n    item recent 1 hour/day/week clicks/carts/orders count from prediction timestamp\n    \n    relative time gap between the item clicked/carted/ordered last time and prediction timestamp\n    \n    co_visition weights(time weight, type weight, just count) for items in each session\n    \n    dnn model cosine similaries between session and item embeddings\n    \n    etc...\n\n\n#### Models used:\nLgbm, Catboost, Xgboost, MLP\n\nClassification and ranking algorithm perform similarly in my solution.\n\nAnd ensembling classification and ranking does not provide any boost.\n\nSo only binary classification is used at the end.\n\n\n#### Ensemble:\nBlending of 4 models, each with 5 folds.\n\nStacking with meta model leads to a silimar performance, thus only blending is used.\n\n#### Post processing:\nThere are 3327 sessions with less than 20 candidates.\n\nSo I re-calculated the co_visitition items for these sessions without constraining the time range, namely with full data.\n\nAfter this supplement, the average candidates number is above 19.\n\nUsing hottest items within half day around the prediction timestamp to supplement the last drop.\n\n\n### Acknowledgement\n\nThank @radek1  for introducing polars to the community, which saves me in this competition with limited ram.\n\nIts lazy evalution and native multithreading support make it such a nice Lib on CPU platform.\n\nThank @cdeotte for sharing the knowledge and valuable notebooks as always.\n\nThank all guys who shared their knowledge and insights in notebooks and discussions. Learned a lot from you guys.",
    "2124746": "Thanks a lot for your sharing. Hope you can share the config and source code for the SRGNN model in the near future. Graph is really efficient for single model.",
    "2124882": "Thanks. So u just do the plain cv and OOF prediction in Radek's dataset. I still wonder why it's seems you cv is more robust than others?",
    "2124899": "No, I created my own dataset using host version 1 script, where there is unseen items in valid set. And it turns out to be not that robust. Only correlate with public LB well, the private LB severely shakes down.",
    "2124975": "SRGNN ! Great I really like Graph based solution. Personnally,  I used GraphSAGE and Node2Vec to generate embeddings from USERs and AIDs in order to compute the similarity score between them,I can say that using different approaches can help ! ( those features belong to the top 10 of my LGBM ) Unfortunately I didn't work much on candidates generation.\n\nI also noticed than using those graph features was enough. From my public and private scores, I can say that using 20 features ( 15 of them are graph based features ), gave me 0.579, same as using 120 features !\n\nThanks for sharing @buumoo",
    "2124983": "Yes, your feeling is right. SRGNN similarity score is the top 1 feature in some of my models.",
    "2124995": "can you further explain this :  \"Re_visitition: All past events\", with a little exemple ?",
    "2125001": "Just use all items in one session as its candidates.",
    "2125022": "Ok I got it ! I thought about another strategy. Yes you mean the unique aids of each session ! thanks for the insight",
    "2125087": "Thanks for your sharing. I have not tried NN based solution, as think the item/aid is high cardinality, may I know if any special handling for your SRGNN?",
    "2125094": "Just modified the data processing part to accelerate training speed.\nFilter out less frequent item to reduce cardinality.",
    "2126116": "Congrats on your strong finish!\n\n\nThe Devastator.",
    "2127538": "solo gold medal! Congras.",
    "2127642": "Congratulations for u up to golden now. It's seems someone was dropped in top15.",
    "2127813": "Congratulations @buumoo ! A really well deserved solo gold - love your approach to candidate generation with SRGNN and data augmentation too (I definitely missed that trick!), and great features as well :)",
    "2129334": "Congratulations!!!! @buumoo 👍 Hope someday I can get my solo gold just like you",
    "2129336": "Also, could you plz share the feature importance?(if you ever did the `plot_importance()`)",
    "2134446": "Thanks mate"
  },
  "source": "meta"
}