{
  "id": 383229,
  "title": "13th Place Solution",
  "url": "/competitions/otto-recommender-system/writeups/newteam-13th-place-solution",
  "author_name": "",
  "post_date": "2023-02-12T17:49:21.797Z",
  "votes": 27,
  "comment_count": 5,
  "views": 0,
  "content": "<p>First of all, thank you for organizing a great competition. I would also like to thank everyone who participated in the competition with us and the four of us who worked together as a team.<br>\nSepcially thanks to my teammates: <a href=\"https://www.kaggle.com/trasibulo\" target=\"_blank\">@trasibulo</a> <a href=\"https://www.kaggle.com/virilo\" target=\"_blank\">@virilo</a> and <a href=\"https://www.kaggle.com/albert2017\" target=\"_blank\">@albert2017</a> </p>\n<h2>Candidates Selection</h2>\n<p>We merge several types of models for gettint the final candidates. Each model contribute with different number of candidates and then we ensemble them over a ranking based weights to get the final 150 candidates. We tested 100 and 150, we noticed a slightly improvement with 150.</p>\n<h3>Scores:</h3>\n<p>Our final top150 recall was: <br><br>\n   clicks recall=0.7241 <br><br>\n   carts recall=0.5636 <br><br>\n   orders recall=0.7399 <br><br>\n   <strong>Overall Recall(top 150)</strong> = 0.6854 <br></p>\n<h3>Base models</h3>\n<ul>\n<li>historical items weighted by time and type.</li>\n<li>fnoa candidates (50-150 candidates depending on the historic).</li>\n<li>transformers4rec based model in order to get candidates. (100)</li>\n<li>short term coocurrence (100)</li>\n<li>covisitation matrix (100)</li>\n<li>w2vec knn (100)</li>\n</ul>\n<h2>Ranking</h2>\n<h3>Fnoa</h3>\n<p>I started with pandas but finally I changed all my pipeline to polars.<br>\nI decided to do one model per item type: clicks, carts, orders</p>\n<h4>Features used:</h4>\n<p>item_features<br>\nsession_features<br>\nuser_items_features<br>\nfeatures from candidates generation (probs, similarity, etc..)</p>\n<h4>Models:</h4>\n<p>We tried lgb and catboost;<br>\nAlthough lgb was slightly better on LB, catboost was faster so at the end I decided to go only with catboost.</p>\n<h4>Things that didnt work:</h4>\n<p>I tried many things for improving validation score at the end but everything i tried didn't seem to help<br>\nUse carts predictions as feature for orders<br>\nStacking<br>\nDifferent models<br>\nTrain with all data<br>\nw2vec similarity features</p>\n<h3>Xiao</h3>\n<p>Xgboost model and lgbm model. (comming soon)</p>\n<h3>Virilo</h3>\n<p>BST Transformer.<br>\n<a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/384441\" target=\"_blank\">https://www.kaggle.com/competitions/otto-recommender-system/discussion/384441</a></p>\n<h2>Ensemble &amp; Stacking</h2>\n<p>We tried both ensemble and stacking.<br>\n<strong>Stacking:</strong> we didnt managed to make it work on LB (while we were getting a good validation score, it was not reflected on LB)<br>\n<strong>Ensemble:</strong> it seemed to be more stable but also didnt give us the expected boost.<br>\nAt the end our final submission was an ensemble of probabilities from 3 models: xgb,lgb,catboost<br>\nLB results-&gt; single model (0.602 low); ensemble (0.602 high)</p>",
  "messages": [
    {
      "id": "2127187",
      "postDate": "02/02/2023 18:29:39",
      "content": "<p>First of all, thank you for organizing a great competition. I would also like to thank everyone who participated in the competition with us and the four of us who worked together as a team.<br>\nSepcially thanks to my teammates: <a href=\"https://www.kaggle.com/trasibulo\" target=\"_blank\">@trasibulo</a> <a href=\"https://www.kaggle.com/virilo\" target=\"_blank\">@virilo</a> and <a href=\"https://www.kaggle.com/albert2017\" target=\"_blank\">@albert2017</a> </p>\n<h2>Candidates Selection</h2>\n<p>We merge several types of models for gettint the final candidates. Each model contribute with different number of candidates and then we ensemble them over a ranking based weights to get the final 150 candidates. We tested 100 and 150, we noticed a slightly improvement with 150.</p>\n<h3>Scores:</h3>\n<p>Our final top150 recall was: <br><br>\n   clicks recall=0.7241 <br><br>\n   carts recall=0.5636 <br><br>\n   orders recall=0.7399 <br><br>\n   <strong>Overall Recall(top 150)</strong> = 0.6854 <br></p>\n<h3>Base models</h3>\n<ul>\n<li>historical items weighted by time and type.</li>\n<li>fnoa candidates (50-150 candidates depending on the historic).</li>\n<li>transformers4rec based model in order to get candidates. (100)</li>\n<li>short term coocurrence (100)</li>\n<li>covisitation matrix (100)</li>\n<li>w2vec knn (100)</li>\n</ul>\n<h2>Ranking</h2>\n<h3>Fnoa</h3>\n<p>I started with pandas but finally I changed all my pipeline to polars.<br>\nI decided to do one model per item type: clicks, carts, orders</p>\n<h4>Features used:</h4>\n<p>item_features<br>\nsession_features<br>\nuser_items_features<br>\nfeatures from candidates generation (probs, similarity, etc..)</p>\n<h4>Models:</h4>\n<p>We tried lgb and catboost;<br>\nAlthough lgb was slightly better on LB, catboost was faster so at the end I decided to go only with catboost.</p>\n<h4>Things that didnt work:</h4>\n<p>I tried many things for improving validation score at the end but everything i tried didn't seem to help<br>\nUse carts predictions as feature for orders<br>\nStacking<br>\nDifferent models<br>\nTrain with all data<br>\nw2vec similarity features</p>\n<h3>Xiao</h3>\n<p>Xgboost model and lgbm model. (comming soon)</p>\n<h3>Virilo</h3>\n<p>BST Transformer.<br>\n<a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/384441\" target=\"_blank\">https://www.kaggle.com/competitions/otto-recommender-system/discussion/384441</a></p>\n<h2>Ensemble &amp; Stacking</h2>\n<p>We tried both ensemble and stacking.<br>\n<strong>Stacking:</strong> we didnt managed to make it work on LB (while we were getting a good validation score, it was not reflected on LB)<br>\n<strong>Ensemble:</strong> it seemed to be more stable but also didnt give us the expected boost.<br>\nAt the end our final submission was an ensemble of probabilities from 3 models: xgb,lgb,catboost<br>\nLB results-&gt; single model (0.602 low); ensemble (0.602 high)</p>",
      "rawMarkdown": "First of all, thank you for organizing a great competition. I would also like to thank everyone who participated in the competition with us and the four of us who worked together as a team.\n\nSepcially thanks to my teammates: @trasibulo @virilo and @albert2017 \n\n## Candidates Selection\n\nWe merge several types of models for gettint the final candidates. Each model contribute with different number of candidates and then we ensemble them over a ranking based weights to get the final 150 candidates. We tested 100 and 150, we noticed a slightly improvement with 150.\n\n### Scores: \n\nOur final top150 recall was: <br>\n    clicks recall=0.7241 <br>\n    carts recall=0.5636 <br>\n    orders recall=0.7399 <br>\n    **Overall Recall(top 150)** = 0.6854 <br>\n\n### Base models\n\n- historical items weighted by time and type.\n- fnoa candidates (50-150 candidates depending on the historic).\n- transformers4rec based model in order to get candidates. (100)\n- short term coocurrence (100)\n- covisitation matrix (100)\n- w2vec knn (100)\n\n## Ranking\n\n### Fnoa \n\nI started with pandas but finally I changed all my pipeline to polars.\n\nI decided to do one model per item type: clicks, carts, orders\n\n#### Features used:\nitem_features\nsession_features\nuser_items_features\nfeatures from candidates generation (probs, similarity, etc..)\n\n#### Models:\nWe tried lgb and catboost;\nAlthough lgb was slightly better on LB, catboost was faster so at the end I decided to go only with catboost.\n\n#### Things that didnt work:\nI tried many things for improving validation score at the end but everything i tried didn't seem to help\nUse carts predictions as feature for orders\nStacking\nDifferent models\nTrain with all data\nw2vec similarity features\n\n### Xiao\nXgboost model and lgbm model. (comming soon)\n\n### Virilo\nBST Transformer.\nhttps://www.kaggle.com/competitions/otto-recommender-system/discussion/384441\n\n## Ensemble & Stacking\nWe tried both ensemble and stacking.\n\n**Stacking:** we didnt managed to make it work on LB (while we were getting a good validation score, it was not reflected on LB)\n\n**Ensemble:** it seemed to be more stable but also didnt give us the expected boost.\n\nAt the end our final submission was an ensemble of probabilities from 3 models: xgb,lgb,catboost\n\nLB results-> single model (0.602 low); ensemble (0.602 high)",
      "votes": null
    },
    {
      "id": "2127249",
      "postDate": "02/02/2023 19:18:11",
      "content": "<p>Polars is new to me, thanks for the mention <a href=\"https://www.kaggle.com/enric1296\" target=\"_blank\">@enric1296</a>! I am sure this should be good as I can see some users using it frequently now-a-days. <br>\nCongrats for the great result! Best regards!</p>",
      "rawMarkdown": "Polars is new to me, thanks for the mention @enric1296! I am sure this should be good as I can see some users using it frequently now-a-days. \nCongrats for the great result! Best regards!",
      "votes": null
    },
    {
      "id": "2127259",
      "postDate": "02/02/2023 19:23:07",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/enric1296\" target=\"_blank\">@enric1296</a> </p>",
      "rawMarkdown": "Congrats @enric1296",
      "votes": null
    },
    {
      "id": "2127745",
      "postDate": "02/03/2023 05:44:36",
      "content": "<p>Well Done mate 👍</p>",
      "rawMarkdown": "Well Done mate 👍",
      "votes": null
    },
    {
      "id": "2128137",
      "postDate": "02/03/2023 13:46:04",
      "content": "<p>Congrats to all of you!</p>",
      "rawMarkdown": "Congrats to all of you!",
      "votes": null
    },
    {
      "id": "2134442",
      "postDate": "02/08/2023 01:47:12",
      "content": "<p>Congrats on the gold, well done.</p>",
      "rawMarkdown": "Congrats on the gold, well done.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2127249,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "02/02/2023 19:18:11",
      "content": "<p>Polars is new to me, thanks for the mention <a href=\"https://www.kaggle.com/enric1296\" target=\"_blank\">@enric1296</a>! I am sure this should be good as I can see some users using it frequently now-a-days. <br>\nCongrats for the great result! Best regards!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2127259,
      "author_name": "abhi011097",
      "author_url": "",
      "post_date": "02/02/2023 19:23:07",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/enric1296\" target=\"_blank\">@enric1296</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2127745,
      "author_name": "ayushnitb",
      "author_url": "",
      "post_date": "02/03/2023 05:44:36",
      "content": "<p>Well Done mate 👍</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2128137,
      "author_name": "indiratsirikhova",
      "author_url": "",
      "post_date": "02/03/2023 13:46:04",
      "content": "<p>Congrats to all of you!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2134442,
      "author_name": "buumoo",
      "author_url": "",
      "post_date": "02/08/2023 01:47:12",
      "content": "<p>Congrats on the gold, well done.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2127187": "First of all, thank you for organizing a great competition. I would also like to thank everyone who participated in the competition with us and the four of us who worked together as a team.\n\nSepcially thanks to my teammates: @trasibulo @virilo and @albert2017 \n\n## Candidates Selection\n\nWe merge several types of models for gettint the final candidates. Each model contribute with different number of candidates and then we ensemble them over a ranking based weights to get the final 150 candidates. We tested 100 and 150, we noticed a slightly improvement with 150.\n\n### Scores: \n\nOur final top150 recall was: <br>\n    clicks recall=0.7241 <br>\n    carts recall=0.5636 <br>\n    orders recall=0.7399 <br>\n    **Overall Recall(top 150)** = 0.6854 <br>\n\n### Base models\n\n- historical items weighted by time and type.\n- fnoa candidates (50-150 candidates depending on the historic).\n- transformers4rec based model in order to get candidates. (100)\n- short term coocurrence (100)\n- covisitation matrix (100)\n- w2vec knn (100)\n\n## Ranking\n\n### Fnoa \n\nI started with pandas but finally I changed all my pipeline to polars.\n\nI decided to do one model per item type: clicks, carts, orders\n\n#### Features used:\nitem_features\nsession_features\nuser_items_features\nfeatures from candidates generation (probs, similarity, etc..)\n\n#### Models:\nWe tried lgb and catboost;\nAlthough lgb was slightly better on LB, catboost was faster so at the end I decided to go only with catboost.\n\n#### Things that didnt work:\nI tried many things for improving validation score at the end but everything i tried didn't seem to help\nUse carts predictions as feature for orders\nStacking\nDifferent models\nTrain with all data\nw2vec similarity features\n\n### Xiao\nXgboost model and lgbm model. (comming soon)\n\n### Virilo\nBST Transformer.\nhttps://www.kaggle.com/competitions/otto-recommender-system/discussion/384441\n\n## Ensemble & Stacking\nWe tried both ensemble and stacking.\n\n**Stacking:** we didnt managed to make it work on LB (while we were getting a good validation score, it was not reflected on LB)\n\n**Ensemble:** it seemed to be more stable but also didnt give us the expected boost.\n\nAt the end our final submission was an ensemble of probabilities from 3 models: xgb,lgb,catboost\n\nLB results-> single model (0.602 low); ensemble (0.602 high)",
    "2127249": "Polars is new to me, thanks for the mention @enric1296! I am sure this should be good as I can see some users using it frequently now-a-days. \nCongrats for the great result! Best regards!",
    "2127259": "Congrats @enric1296",
    "2127745": "Well Done mate 👍",
    "2128137": "Congrats to all of you!",
    "2134442": "Congrats on the gold, well done."
  },
  "source": "meta"
}