{
  "id": 383382,
  "title": "16th Place Solution (Association x UserCF x NN-based x Matrix Factorization x Covisit x LightGBM)",
  "url": "/competitions/otto-recommender-system/writeups/luftschloss-16th-place-solution-association-x-user",
  "author_name": "",
  "post_date": "2023-02-07T23:12:34.113Z",
  "votes": 28,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Thanks to Otto and Kaggle team for hosting such a challenging competition. I tried almost all methods I could come up with and seeing the score improve step by step was thrilling, especially at the end of the competition.</p>\n<p>The sad part is the Grandmaster/Master are allegedly involved in cheating, and many participants have been suffering/discouraged from such an unforgivable activity that hurts Kaggle's integrity and reputation.</p>\n<p>I would like to share my solution, <strong>which is 0.01% behind the Gold medal zone at this moment lol.</strong><br>\nHope this helps, and I'm happy to answer questions.</p>\n<h1><strong>Candidate Selection</strong></h1>\n<ul>\n<li>revisit</li>\n<li>top20 covisitation matrix (two different matrix)</li>\n<li>top3 3-hop graph from covisitation matrix<ul>\n<li>lets say sessions's last aid is X, and X's top 3 covisit aids are A,B,C. I call them '1-hop' from X. A's top3 covisit aids are D,E,F. I call them '2-hop from X'. I did this for three hops so you have 3^3 cand to 1 aid)</li></ul></li>\n<li>top10 from 3D-covisitation matrix</li>\n<li>top20 from simple itemcf/usercf (similaripy)</li>\n<li>top20 from gru4rec (recbole)</li>\n</ul>\n<p><strong>The average cand per session is almost 100, and recall for valid is 0.6396</strong></p>\n<h1><strong>Feature Engineering</strong></h1>\n<ul>\n<li>simple interaction between aid and session (ex: time lapse from session's last action to candidate aid)</li>\n<li>session-only based (ex: num of clicks of that session)</li>\n<li>aid-only based (ex: num of clicks of that aid)</li>\n<li>covisitation matrix score from four different covisitation matrices</li>\n<li>association rules score (jaccard/dice/lift---)</li>\n<li>simple user/item cf score (using <strong>'similaripy'</strong>) </li>\n<li>item2vec cosine similarity (using <strong>'gensim'</strong>)</li>\n<li>matrix factorization type similarity (using <strong>'implicit'</strong>)<ul>\n<li>BPR</li>\n<li>ALS</li>\n<li>LightFM</li>\n<li>SVD </li></ul></li>\n<li>neural net based scores (using <strong>'recbole'</strong>)<ul>\n<li>lightGCN</li>\n<li>GRU4Rec</li>\n<li>BERT4Rec</li>\n<li>RecVAE</li>\n<li>SasREC</li>\n<li>SRGNN</li></ul></li>\n</ul>\n<p>The total num of features is 291.<br>\nFrom my observation, there was no silver bullet feature(s). All types of features slightly contributed to the model. <br>\nNN features are relatively costly because you need time and computing resources to get them but still get marginal boost.<br>\nAssociation rules and simple userCF are easy to handle because they are not so computationally intensive.</p>\n<p><strong>Single model best CV is .5808 and public .60206/private .6017</strong></p>\n<h1><strong>Modeling</strong></h1>\n<p>LightGBM Ranker with custom metrics(Recall at 20)<br>\nLightGBM binary/Catboost binary/Catboost Yetirank are all worse than LightGBM Ranker.</p>\n<p>The final submission is the ensemble from five slightly different candidates/model specifications. Ensemble helped me a little(+0.03%).</p>\n<h1><strong>What I should have done to get solo gold</strong></h1>\n<ul>\n<li>Catboost Pairwise (some report it works well)</li>\n<li>More Candidate (maybe 150-200 per candidates)</li>\n</ul>",
  "messages": [
    {
      "id": "2128097",
      "postDate": "02/03/2023 13:15:26",
      "content": "<p>Thanks to Otto and Kaggle team for hosting such a challenging competition. I tried almost all methods I could come up with and seeing the score improve step by step was thrilling, especially at the end of the competition.</p>\n<p>The sad part is the Grandmaster/Master are allegedly involved in cheating, and many participants have been suffering/discouraged from such an unforgivable activity that hurts Kaggle's integrity and reputation.</p>\n<p>I would like to share my solution, <strong>which is 0.01% behind the Gold medal zone at this moment lol.</strong><br>\nHope this helps, and I'm happy to answer questions.</p>\n<h1><strong>Candidate Selection</strong></h1>\n<ul>\n<li>revisit</li>\n<li>top20 covisitation matrix (two different matrix)</li>\n<li>top3 3-hop graph from covisitation matrix<ul>\n<li>lets say sessions's last aid is X, and X's top 3 covisit aids are A,B,C. I call them '1-hop' from X. A's top3 covisit aids are D,E,F. I call them '2-hop from X'. I did this for three hops so you have 3^3 cand to 1 aid)</li></ul></li>\n<li>top10 from 3D-covisitation matrix</li>\n<li>top20 from simple itemcf/usercf (similaripy)</li>\n<li>top20 from gru4rec (recbole)</li>\n</ul>\n<p><strong>The average cand per session is almost 100, and recall for valid is 0.6396</strong></p>\n<h1><strong>Feature Engineering</strong></h1>\n<ul>\n<li>simple interaction between aid and session (ex: time lapse from session's last action to candidate aid)</li>\n<li>session-only based (ex: num of clicks of that session)</li>\n<li>aid-only based (ex: num of clicks of that aid)</li>\n<li>covisitation matrix score from four different covisitation matrices</li>\n<li>association rules score (jaccard/dice/lift---)</li>\n<li>simple user/item cf score (using <strong>'similaripy'</strong>) </li>\n<li>item2vec cosine similarity (using <strong>'gensim'</strong>)</li>\n<li>matrix factorization type similarity (using <strong>'implicit'</strong>)<ul>\n<li>BPR</li>\n<li>ALS</li>\n<li>LightFM</li>\n<li>SVD </li></ul></li>\n<li>neural net based scores (using <strong>'recbole'</strong>)<ul>\n<li>lightGCN</li>\n<li>GRU4Rec</li>\n<li>BERT4Rec</li>\n<li>RecVAE</li>\n<li>SasREC</li>\n<li>SRGNN</li></ul></li>\n</ul>\n<p>The total num of features is 291.<br>\nFrom my observation, there was no silver bullet feature(s). All types of features slightly contributed to the model. <br>\nNN features are relatively costly because you need time and computing resources to get them but still get marginal boost.<br>\nAssociation rules and simple userCF are easy to handle because they are not so computationally intensive.</p>\n<p><strong>Single model best CV is .5808 and public .60206/private .6017</strong></p>\n<h1><strong>Modeling</strong></h1>\n<p>LightGBM Ranker with custom metrics(Recall at 20)<br>\nLightGBM binary/Catboost binary/Catboost Yetirank are all worse than LightGBM Ranker.</p>\n<p>The final submission is the ensemble from five slightly different candidates/model specifications. Ensemble helped me a little(+0.03%).</p>\n<h1><strong>What I should have done to get solo gold</strong></h1>\n<ul>\n<li>Catboost Pairwise (some report it works well)</li>\n<li>More Candidate (maybe 150-200 per candidates)</li>\n</ul>",
      "rawMarkdown": "Thanks to Otto and Kaggle team for hosting such a challenging competition. I tried almost all methods I could come up with and seeing the score improve step by step was thrilling, especially at the end of the competition.\n\nThe sad part is the Grandmaster/Master are allegedly involved in cheating, and many participants have been suffering/discouraged from such an unforgivable activity that hurts Kaggle's integrity and reputation.\n\nI would like to share my solution, **which is 0.01% behind the Gold medal zone at this moment lol.**\nHope this helps, and I'm happy to answer questions.\n\n# **Candidate Selection**\n- revisit\n- top20 covisitation matrix (two different matrix)\n- top3 3-hop graph from covisitation matrix\n  - lets say sessions's last aid is X, and X's top 3 covisit aids are A,B,C. I call them '1-hop' from X. A's top3 covisit aids are D,E,F. I call them '2-hop from X'. I did this for three hops so you have 3^3 cand to 1 aid)\n- top10 from 3D-covisitation matrix\n- top20 from simple itemcf/usercf (similaripy)\n- top20 from gru4rec (recbole)\n\n**The average cand per session is almost 100, and recall for valid is 0.6396**\n\n# **Feature Engineering**\n- simple interaction between aid and session (ex: time lapse from session's last action to candidate aid)\n- session-only based (ex: num of clicks of that session)\n- aid-only based (ex: num of clicks of that aid)\n- covisitation matrix score from four different covisitation matrices\n- association rules score (jaccard/dice/lift---)\n- simple user/item cf score (using **'similaripy'**) \n- item2vec cosine similarity (using **'gensim'**)\n- matrix factorization type similarity (using **'implicit'**)\n  - BPR\n  - ALS\n  - LightFM\n  - SVD \n- neural net based scores (using **'recbole'**)\n - lightGCN\n - GRU4Rec\n - BERT4Rec\n - RecVAE\n - SasREC\n - SRGNN\n\nThe total num of features is 291.\nFrom my observation, there was no silver bullet feature(s). All types of features slightly contributed to the model. \nNN features are relatively costly because you need time and computing resources to get them but still get marginal boost.\nAssociation rules and simple userCF are easy to handle because they are not so computationally intensive.\n\n**Single model best CV is .5808 and public .60206/private .6017**\n\n# **Modeling**\nLightGBM Ranker with custom metrics(Recall at 20)\nLightGBM binary/Catboost binary/Catboost Yetirank are all worse than LightGBM Ranker.\n\nThe final submission is the ensemble from five slightly different candidates/model specifications. Ensemble helped me a little(+0.03%).\n\n# **What I should have done to get solo gold**\n - Catboost Pairwise (some report it works well)\n - More Candidate (maybe 150-200 per candidates)",
      "votes": null
    },
    {
      "id": "2128308",
      "postDate": "02/03/2023 16:01:09",
      "content": "<p>Good job. The leaderboard is not final, so you may still get a solo gold.</p>",
      "rawMarkdown": "Good job. The leaderboard is not final, so you may still get a solo gold.",
      "votes": null
    },
    {
      "id": "2131707",
      "postDate": "02/06/2023 09:58:46",
      "content": "<p>Thank you for your kind reply! As two teams seemed to be involved in cheating, I may get a solo gold. <br>\nHope Kaggle team thinks I deserve that, lol.</p>",
      "rawMarkdown": "Thank you for your kind reply! As two teams seemed to be involved in cheating, I may get a solo gold. \nHope Kaggle team thinks I deserve that, lol.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2128308,
      "author_name": "tilii7",
      "author_url": "",
      "post_date": "02/03/2023 16:01:09",
      "content": "<p>Good job. The leaderboard is not final, so you may still get a solo gold.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2131707,
          "author_name": "tommy1028",
          "author_url": "",
          "post_date": "02/06/2023 09:58:46",
          "content": "<p>Thank you for your kind reply! As two teams seemed to be involved in cheating, I may get a solo gold. <br>\nHope Kaggle team thinks I deserve that, lol.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2128097": "Thanks to Otto and Kaggle team for hosting such a challenging competition. I tried almost all methods I could come up with and seeing the score improve step by step was thrilling, especially at the end of the competition.\n\nThe sad part is the Grandmaster/Master are allegedly involved in cheating, and many participants have been suffering/discouraged from such an unforgivable activity that hurts Kaggle's integrity and reputation.\n\nI would like to share my solution, **which is 0.01% behind the Gold medal zone at this moment lol.**\nHope this helps, and I'm happy to answer questions.\n\n# **Candidate Selection**\n- revisit\n- top20 covisitation matrix (two different matrix)\n- top3 3-hop graph from covisitation matrix\n  - lets say sessions's last aid is X, and X's top 3 covisit aids are A,B,C. I call them '1-hop' from X. A's top3 covisit aids are D,E,F. I call them '2-hop from X'. I did this for three hops so you have 3^3 cand to 1 aid)\n- top10 from 3D-covisitation matrix\n- top20 from simple itemcf/usercf (similaripy)\n- top20 from gru4rec (recbole)\n\n**The average cand per session is almost 100, and recall for valid is 0.6396**\n\n# **Feature Engineering**\n- simple interaction between aid and session (ex: time lapse from session's last action to candidate aid)\n- session-only based (ex: num of clicks of that session)\n- aid-only based (ex: num of clicks of that aid)\n- covisitation matrix score from four different covisitation matrices\n- association rules score (jaccard/dice/lift---)\n- simple user/item cf score (using **'similaripy'**) \n- item2vec cosine similarity (using **'gensim'**)\n- matrix factorization type similarity (using **'implicit'**)\n  - BPR\n  - ALS\n  - LightFM\n  - SVD \n- neural net based scores (using **'recbole'**)\n - lightGCN\n - GRU4Rec\n - BERT4Rec\n - RecVAE\n - SasREC\n - SRGNN\n\nThe total num of features is 291.\nFrom my observation, there was no silver bullet feature(s). All types of features slightly contributed to the model. \nNN features are relatively costly because you need time and computing resources to get them but still get marginal boost.\nAssociation rules and simple userCF are easy to handle because they are not so computationally intensive.\n\n**Single model best CV is .5808 and public .60206/private .6017**\n\n# **Modeling**\nLightGBM Ranker with custom metrics(Recall at 20)\nLightGBM binary/Catboost binary/Catboost Yetirank are all worse than LightGBM Ranker.\n\nThe final submission is the ensemble from five slightly different candidates/model specifications. Ensemble helped me a little(+0.03%).\n\n# **What I should have done to get solo gold**\n - Catboost Pairwise (some report it works well)\n - More Candidate (maybe 150-200 per candidates)",
    "2128308": "Good job. The leaderboard is not final, so you may still get a solo gold.",
    "2131707": "Thank you for your kind reply! As two teams seemed to be involved in cheating, I may get a solo gold. \nHope Kaggle team thinks I deserve that, lol."
  },
  "source": "meta"
}