{
  "id": 382879,
  "title": "3rd place(imaginary) solution",
  "url": "/competitions/otto-recommender-system/writeups/m-s-a-s-3rd-place-imaginary-solution",
  "author_name": "",
  "post_date": "2023-02-09T22:43:12.337Z",
  "votes": 98,
  "comment_count": 20,
  "views": 0,
  "content": "<p><strong>PS: This solution is actually equivalent to 3rd place, but it was ban because a cheater was on the team. So please enjoy it as an \"imaginary\" 3rd place solution.</strong></p>\n<p>First of all, thank you for organizing a great competition. I would also like to thank everyone who participated in the competition with us and the four of us who worked together as a team.<br>\nAbout one week before the end of the competition, we formed a team. However, just on the next day, we were warned one of our teammate might have some problem(<a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/381321\" target=\"_blank\">link</a>). And since then, we other four were working hard together to achieve the best possible result. <br>\nWe can proudly tell you that this 3rd place score is definitely a result with no outside influence. (Of course, as we have already said, we leave it up to the kaggle team to decide how our results will ultimately be handled.)<br>\nIn this discussion, I would be happy to share with you the solutions of the four of us.<br>\nOf course we can share the code as well, but we hope we can open it after the rankings are finalized.<br>\n(The code is not yet public at this time, so it would be difficult for cheaters to imitate our accuracy.)</p>\n<p>We write the solution each one below.</p>\n<hr>\n<h2>Alvor part</h2>\n<h3>Candidates selection:</h3>\n<ul>\n<li>all items from the session's history</li>\n<li>items from co-visitation matrices. I did many experiments (window size, weights, action types etc.) and finally end up with 5 best matrices (in terms of maximum possible recall when using each matrix). Each of these 5 matrices I use both for candidates selection and for feature engineering (sum/avg/max candidate score, etc.)</li>\n<li>I did not have fixed number of candidates per session</li>\n<li>median number of candidates per session: 67</li>\n<li>but there were sessions with very high number of candidates, so the average candidates number per session is 121.46</li>\n<li>my recalls per target:<br>\nclicks recall: 0.631531<br>\ncarts recall: 0.540274<br>\norders recall: 0.729626<br>\nTotal: 0.6630<br>\nWhat did not work for me:</li>\n<li>add candidates from Word2Vec embeddings (K nearest neighbours)</li>\n<li>add candidates from Matrix Factorization embeddings (K nearest neighbours)</li>\n<li>simply increase number of candidates in my methods</li>\n<li>all these attempts increased the maximum possible recall, but decreased the validation score</li>\n</ul>\n<h3>Feature engineering:</h3>\n<ul>\n<li>Sessions features</li>\n<li>Items features</li>\n<li>Items-2-Sessions interactions features</li>\n<li>Co-Visitation matrices features (sum/avg/max of candidate score, etc.)</li>\n<li>Word2Vec features: euclidean and angular distance between the candidate's embedding and session's last item's embedding</li>\n<li>Matrix Factorization features: euclidean and angular distance between the candidate's embedding and last session's item's embedding</li>\n<li>Features selection based on Feature importance of my experimental models<br>\nSome of my best candidates:</li>\n<li>index of last interaction with the candidate in the session</li>\n<li>relative number of interactions with given candidate among all session's interactions</li>\n<li>timestamp differences</li>\n<li>angular distance between candidate's w2v embedding and last session's item w2v embedding</li>\n<li>number of co-visitation matrices containing given item as a candidate<br>\nWhat did not work for me:</li>\n<li>Word2Vec and Matrix Factorization features: euclidean and angular distance between the candidate's embedding and session's penultimate item's embedding</li>\n<li>Word2Vec and Matrix Factorization features: average/min euclidean and angular distance between the candidate's embedding and all session's items embeddings</li>\n<li>Word2Vec and Matrix Factorization features: euclidean and angular distance between the candidate's embedding and session's last cart/order item's embedding (there was a minor improvement, but I abandoned these features)</li>\n</ul>\n<h3>Model</h3>\n<ul>\n<li>Group 3-Fold (by sessions) LightGBM Classifier (binary_logloss)</li>\n<li>One model per one target (clicks/carts/orders)</li>\n<li>No negative downsample (during almost the entire competition) so it required a lot of memory and time</li>\n<li>1/5 negative downsample for carts/orders and 1/15 for clicks (last week of competition)</li>\n<li>But I use another trick to reduce memory/time usage: When I train a model for carts/orders, I use only sessions for which I have at least one ground-truth candidate among my candidates. Because the main goal of model is to distinguish between good candidates and bad candidates. So I think that sessions with only bad candidates don't help my model.<br>\n(It could be by 2 reasons:<br>\n1) the session has not ground true cart/order at all<br>\n2) the session has ground-true cart/order, but I didn't select it among my candidates)</li>\n<li>'dart' boosting for carts nodel, 'gbdt' boosting for clicks and orders model.</li>\n<li>add \"carts\" prediction as a feature for \"orders\" model (the best feature)<br>\nWhat did not work for me:</li>\n<li>Add second click to the positive samples as well as first click</li>\n<li>add \"clicks\" prediction as a feature for \"orders\" model</li>\n<li>multiple attempts to build a second model to re-arrange only top-X candidates from the first model</li>\n</ul>\n<hr>\n<h2>Makotu part</h2>\n<p>My pipeline consists mainly of preprocess/make co-matrix/candidate / make feature / modeling parts.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F397595%2Fcff84b84bab47c015ef64ab84e20de5d%2Fmakotu_solution_v2.png?generation=1675314772195222&amp;alt=media\" alt=\"\"></p>\n<h3>Data</h3>\n<ul>\n<li>I used Radek's train/validation/test data</li>\n</ul>\n<h3>Preprocess</h3>\n<ul>\n<li>make aid vector using w2v</li>\n<li>make BPR(Bayesian Personalized Ranking) feature (recommended Sirius)</li>\n</ul>\n<h3>Make-co-matrix and candidate</h3>\n<ul>\n<li>I made many pattern co-matrix. Basically, co-matrix is created in the  form of probabilities. For example, the following pattern is used to quantify the relationship between aidA and aidB.<ul>\n<li>aid B count within same session / all aid A count</li>\n<li>aid B count within same session (after aid A click) / all aid A count</li>\n<li>aid B count within same session / all aid A count. (But if the same user has the same pair of idA and idB, make them unique and then aggregate)</li>\n<li>aid B count within same session / all aid A count. (with time weighted)</li>\n<li>aid B count within same session and within 1 hour / all aid A count.<br>\nlike this.</li></ul></li>\n<li>Select a candidate based on session last action aid / top action aid / with in 1 hour action aid / within 1 day action aid</li>\n<li>candidate recall<ul>\n<li>order: 0.732 (@190 candidates)</li>\n<li>cart: 0.546 (@190 candidates)</li>\n<li>click: 0.682 (@140 candidates)</li></ul></li>\n</ul>\n<h3>Make feature</h3>\n<ul>\n<li>User / aid / user&amp;aid interaction features.</li>\n<li>features that worked<ul>\n<li>last action aid</li>\n<li>distance from last aid w2v to candidate aid w2v</li>\n<li>BPR feature (recommened sirius)</li>\n<li>co-matrix features (ex. probability of clicking on candidate aid after the last aid)</li>\n<li>etc</li></ul></li>\n</ul>\n<h3>Modeling</h3>\n<ul>\n<li>Catboost (loss: PairLogitPairwise)</li>\n<li>Group Kfold(group: session)</li>\n<li>Add oof features in addition to the above features<ul>\n<li>click model does not use oof.<br>\nHowever, two patterns of click models were created. in addition to predicting the next clicked id, it also predicts all subsequent clicks of the id, and then uses as oof.</li>\n<li>cart model add above 2 model oof.</li>\n<li>order model add click 2 model oofs, cart model's oof</li></ul></li>\n<li>CV click: 0.5632 cart: 0.4429 order: 0.6699</li>\n<li>LB: 0.602</li>\n</ul>\n<h3>Stacking</h3>\n<p>I was not involved in the stacking of results with other members, as all the stacking was carried out by shimacos with great skill. I hope you will wait for his addition. </p>\n<h3>(add) code of my part</h3>\n<p><a href=\"https://github.com/makotu1208/Otto-kaggle-3rd-solution-makotupart\" target=\"_blank\">https://github.com/makotu1208/Otto-kaggle-3rd-solution-makotupart</a></p>\n<hr>\n<h2>Shimacos part</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F397595%2F9602e046e48f6d2028d0a05ccf898720%2Fsolution_shimacos.png?generation=1675249326142169&amp;alt=media\" alt=\"\"></p>\n<h3>data</h3>\n<ul>\n<li>First, I used my own original truncate dataset.<ul>\n<li>train: 2022-08-14 ~ 2022-08-21 valid 2022-08-21 ~ 2022-08-28</li>\n<li>train: 2022-08-21 ~ 2022-08-28 test: 2022-08-28 ~ 2022-09-04 (Training iteration is the calculated one in the previous step)</li></ul></li>\n<li>And Roughly 10 or so models were made and stacking. (Single 0.592 -&gt; 0.595)</li>\n<li>However since everyone except me was using the radek dataset, I shifted to using that one after team merge. This was to maximize ensemble results.</li>\n</ul>\n<h3>preprocess</h3>\n<ul>\n<li>made aid vector<ul>\n<li>w2v with gensim</li>\n<li>n2v with pytorch_geometric</li></ul></li>\n<li>made session, aid vector<ul>\n<li>bpr with implicit</li></ul></li>\n</ul>\n<h3>Candidates</h3>\n<ul>\n<li>As shown in the overview figure.</li>\n<li>I created candidates and features by BigQuery. For example, the covisit candidates can be calculated in less than one minute for both train and test.</li>\n<li>Recall<ul>\n<li>average candidate count: 188.69</li>\n<li>click: 0.6799</li>\n<li>cart: 0.5555</li>\n<li>order: 0.7359</li>\n<li>Overall: 0.6762</li></ul></li>\n</ul>\n<h3>Feature</h3>\n<ul>\n<li>As shown in the overview figure.</li>\n<li>I used the overall average for smoothing when calculating cvr.</li>\n<li>High importance feature (CatBoost LossFunctionChange importance)<ul>\n<li>difference between the last time the aid was actioned across the overall data and the last time of the session.</li>\n<li>candidate feature by next visit covisit</li>\n<li>I think it was because in the end it was important to hit the next aid after the last action.</li>\n<li>difference feature of w2v and bpr</li>\n<li>covisit cvr feature</li></ul></li>\n</ul>\n<h3>Model</h3>\n<ul>\n<li>LightGBM (lambdarank) and CatBoost (PairLogit)</li>\n<li>I used CatBoost the day before the last day and found an improvement of about 0.002-0.003 over LightGBM. I should have used it earlier…<ul>\n<li>I got 0.596 and 0.598 at private LB finally.</li></ul></li>\n</ul>\n<h3>Stacking</h3>\n<ul>\n<li>At the team merge, the scores were only 0.600 (makotu), 0.600 (sirius),0.598 (Alvor and Shimacos), but By rank blending each score as follows, we were able to produce a 0.603, which was the second at the time.<ul>\n<li>Use the top50 of each prediction.</li>\n<li>And the each prediction was simply outer joined to create a pair of user items.</li>\n<li>fillna with 1 / 50</li>\n<li>rank blending with formula <code>1/rank_a + 1/rank_b + 1/rank_c</code> and sort descending.</li></ul></li>\n<li>As the number of models increased, it became difficult to find the optimal weights, so stacking was performed.<ul>\n<li>Dataset is the same as above.</li>\n<li>Feature: prediction score and rank</li></ul></li>\n<li>This resulted in a score of 0.604 by LightGBM. (CV: 0.59624)</li>\n<li>The same was created in CatBoost (CV: 0.59592) and averaged.</li>\n<li>Finally, I interpolated predictions by aid with the highest number of actions for each type in the test period.</li>\n<li>Random Thought<ul>\n<li>Too many more models could lower the score. This is probably due to too many negative examples.</li>\n<li>Stacking models created by different candidates tended to increase the overall score. </li></ul></li>\n</ul>\n<hr>\n<h2>Sirius part</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F397595%2Ffa7ab3c2005f4af8929e6ed5e7b6f42b%2Fsolution_sirius.png?generation=1675249468500843&amp;alt=media\" alt=\"\"></p>\n<h3>Recalling</h3>\n<ul>\n<li>All the recalling methods are listed in the figure.</li>\n<li>Avg candidates cnt for each user: 224.74</li>\n<li>Overall recall: 0.677232,<ul>\n<li>Clicks recall:0.688493</li>\n<li>Carts recall:0.557334</li>\n<li>Orders recall:0.735304</li></ul></li>\n</ul>\n<h3>Ranking</h3>\n<h4>Samples</h4>\n<ul>\n<li>Click model: the generated candidates which the user clicked are labeled with 1</li>\n<li>Cart/Order model: the generated candidates which the user carted or ordered are labeled with 1. I merged cart and order labels because of their sharing of similar semantic</li>\n<li>Downsampling of negtive samples is abopted. I just keep the negative samples with amount of <code>20*len(pos_samples)</code></li>\n</ul>\n<h4>Features</h4>\n<ul>\n<li>All the features are listed in the figure.</li>\n<li>The feature with highest importance in my model is BPR①, implemented with Implicit package.  (<a href=\"https://www.kaggle.com/code/sirius81/bpr-feature/notebook\" target=\"_blank\">Notebook</a> will be public after everything is verified.)</li>\n<li>Bi-gram② is a terminology of NLP, meaning the successive visiting of aid1, aid2 here. It is similar to co-visit matrix, but it only counts the next one action. And I normlize the count by the hotness of item then. This group of features gave the biggest boost among all my experiments, ~0.006.  (<a href=\"https://www.kaggle.com/code/sirius81/otto-bigram-feature/notebook\" target=\"_blank\">Notebook</a> will be public after everything is verified.)</li>\n<li>OOF of other action③ help improve my score ~0.001.</li>\n<li>Features ①+③+ shimacos’s w2v improved Makotu's CV ~0.001 and features ①+② improved Alvor's CV ~0.003.</li>\n</ul>\n<h4>Models</h4>\n<ul>\n<li>Catboost is trained with Logloss. When training Cart/Order model, I set different weights to positive sample, 5 and 10 for is_carts and is_orders respectively. Along with the strategy of mergeing Cart labels and Order labels, it boosts CV ~0.002, compared with training Cart model and Order model separately.</li>\n<li>MLP is trained with the same features as catboost, plus an standard scaler to normlize the features. But it is not accurate and diverse enough to improve the blending score. So in our final submission, it is not included. Beside MLP, I have tried to introduce GRU and pretrained item embedding into the NN model, but all failed.</li>\n</ul>\n<h3>Ensemble</h3>\n<p>That’s <a href=\"https://www.kaggle.com/shimacos\" target=\"_blank\">@shimacos</a> ’s credit. Very impressive work! </p>",
  "messages": [
    {
      "id": "2124992",
      "postDate": "02/01/2023 11:07:27",
      "content": "<p><strong>PS: This solution is actually equivalent to 3rd place, but it was ban because a cheater was on the team. So please enjoy it as an \"imaginary\" 3rd place solution.</strong></p>\n<p>First of all, thank you for organizing a great competition. I would also like to thank everyone who participated in the competition with us and the four of us who worked together as a team.<br>\nAbout one week before the end of the competition, we formed a team. However, just on the next day, we were warned one of our teammate might have some problem(<a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/381321\" target=\"_blank\">link</a>). And since then, we other four were working hard together to achieve the best possible result. <br>\nWe can proudly tell you that this 3rd place score is definitely a result with no outside influence. (Of course, as we have already said, we leave it up to the kaggle team to decide how our results will ultimately be handled.)<br>\nIn this discussion, I would be happy to share with you the solutions of the four of us.<br>\nOf course we can share the code as well, but we hope we can open it after the rankings are finalized.<br>\n(The code is not yet public at this time, so it would be difficult for cheaters to imitate our accuracy.)</p>\n<p>We write the solution each one below.</p>\n<hr>\n<h2>Alvor part</h2>\n<h3>Candidates selection:</h3>\n<ul>\n<li>all items from the session's history</li>\n<li>items from co-visitation matrices. I did many experiments (window size, weights, action types etc.) and finally end up with 5 best matrices (in terms of maximum possible recall when using each matrix). Each of these 5 matrices I use both for candidates selection and for feature engineering (sum/avg/max candidate score, etc.)</li>\n<li>I did not have fixed number of candidates per session</li>\n<li>median number of candidates per session: 67</li>\n<li>but there were sessions with very high number of candidates, so the average candidates number per session is 121.46</li>\n<li>my recalls per target:<br>\nclicks recall: 0.631531<br>\ncarts recall: 0.540274<br>\norders recall: 0.729626<br>\nTotal: 0.6630<br>\nWhat did not work for me:</li>\n<li>add candidates from Word2Vec embeddings (K nearest neighbours)</li>\n<li>add candidates from Matrix Factorization embeddings (K nearest neighbours)</li>\n<li>simply increase number of candidates in my methods</li>\n<li>all these attempts increased the maximum possible recall, but decreased the validation score</li>\n</ul>\n<h3>Feature engineering:</h3>\n<ul>\n<li>Sessions features</li>\n<li>Items features</li>\n<li>Items-2-Sessions interactions features</li>\n<li>Co-Visitation matrices features (sum/avg/max of candidate score, etc.)</li>\n<li>Word2Vec features: euclidean and angular distance between the candidate's embedding and session's last item's embedding</li>\n<li>Matrix Factorization features: euclidean and angular distance between the candidate's embedding and last session's item's embedding</li>\n<li>Features selection based on Feature importance of my experimental models<br>\nSome of my best candidates:</li>\n<li>index of last interaction with the candidate in the session</li>\n<li>relative number of interactions with given candidate among all session's interactions</li>\n<li>timestamp differences</li>\n<li>angular distance between candidate's w2v embedding and last session's item w2v embedding</li>\n<li>number of co-visitation matrices containing given item as a candidate<br>\nWhat did not work for me:</li>\n<li>Word2Vec and Matrix Factorization features: euclidean and angular distance between the candidate's embedding and session's penultimate item's embedding</li>\n<li>Word2Vec and Matrix Factorization features: average/min euclidean and angular distance between the candidate's embedding and all session's items embeddings</li>\n<li>Word2Vec and Matrix Factorization features: euclidean and angular distance between the candidate's embedding and session's last cart/order item's embedding (there was a minor improvement, but I abandoned these features)</li>\n</ul>\n<h3>Model</h3>\n<ul>\n<li>Group 3-Fold (by sessions) LightGBM Classifier (binary_logloss)</li>\n<li>One model per one target (clicks/carts/orders)</li>\n<li>No negative downsample (during almost the entire competition) so it required a lot of memory and time</li>\n<li>1/5 negative downsample for carts/orders and 1/15 for clicks (last week of competition)</li>\n<li>But I use another trick to reduce memory/time usage: When I train a model for carts/orders, I use only sessions for which I have at least one ground-truth candidate among my candidates. Because the main goal of model is to distinguish between good candidates and bad candidates. So I think that sessions with only bad candidates don't help my model.<br>\n(It could be by 2 reasons:<br>\n1) the session has not ground true cart/order at all<br>\n2) the session has ground-true cart/order, but I didn't select it among my candidates)</li>\n<li>'dart' boosting for carts nodel, 'gbdt' boosting for clicks and orders model.</li>\n<li>add \"carts\" prediction as a feature for \"orders\" model (the best feature)<br>\nWhat did not work for me:</li>\n<li>Add second click to the positive samples as well as first click</li>\n<li>add \"clicks\" prediction as a feature for \"orders\" model</li>\n<li>multiple attempts to build a second model to re-arrange only top-X candidates from the first model</li>\n</ul>\n<hr>\n<h2>Makotu part</h2>\n<p>My pipeline consists mainly of preprocess/make co-matrix/candidate / make feature / modeling parts.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F397595%2Fcff84b84bab47c015ef64ab84e20de5d%2Fmakotu_solution_v2.png?generation=1675314772195222&amp;alt=media\" alt=\"\"></p>\n<h3>Data</h3>\n<ul>\n<li>I used Radek's train/validation/test data</li>\n</ul>\n<h3>Preprocess</h3>\n<ul>\n<li>make aid vector using w2v</li>\n<li>make BPR(Bayesian Personalized Ranking) feature (recommended Sirius)</li>\n</ul>\n<h3>Make-co-matrix and candidate</h3>\n<ul>\n<li>I made many pattern co-matrix. Basically, co-matrix is created in the  form of probabilities. For example, the following pattern is used to quantify the relationship between aidA and aidB.<ul>\n<li>aid B count within same session / all aid A count</li>\n<li>aid B count within same session (after aid A click) / all aid A count</li>\n<li>aid B count within same session / all aid A count. (But if the same user has the same pair of idA and idB, make them unique and then aggregate)</li>\n<li>aid B count within same session / all aid A count. (with time weighted)</li>\n<li>aid B count within same session and within 1 hour / all aid A count.<br>\nlike this.</li></ul></li>\n<li>Select a candidate based on session last action aid / top action aid / with in 1 hour action aid / within 1 day action aid</li>\n<li>candidate recall<ul>\n<li>order: 0.732 (@190 candidates)</li>\n<li>cart: 0.546 (@190 candidates)</li>\n<li>click: 0.682 (@140 candidates)</li></ul></li>\n</ul>\n<h3>Make feature</h3>\n<ul>\n<li>User / aid / user&amp;aid interaction features.</li>\n<li>features that worked<ul>\n<li>last action aid</li>\n<li>distance from last aid w2v to candidate aid w2v</li>\n<li>BPR feature (recommened sirius)</li>\n<li>co-matrix features (ex. probability of clicking on candidate aid after the last aid)</li>\n<li>etc</li></ul></li>\n</ul>\n<h3>Modeling</h3>\n<ul>\n<li>Catboost (loss: PairLogitPairwise)</li>\n<li>Group Kfold(group: session)</li>\n<li>Add oof features in addition to the above features<ul>\n<li>click model does not use oof.<br>\nHowever, two patterns of click models were created. in addition to predicting the next clicked id, it also predicts all subsequent clicks of the id, and then uses as oof.</li>\n<li>cart model add above 2 model oof.</li>\n<li>order model add click 2 model oofs, cart model's oof</li></ul></li>\n<li>CV click: 0.5632 cart: 0.4429 order: 0.6699</li>\n<li>LB: 0.602</li>\n</ul>\n<h3>Stacking</h3>\n<p>I was not involved in the stacking of results with other members, as all the stacking was carried out by shimacos with great skill. I hope you will wait for his addition. </p>\n<h3>(add) code of my part</h3>\n<p><a href=\"https://github.com/makotu1208/Otto-kaggle-3rd-solution-makotupart\" target=\"_blank\">https://github.com/makotu1208/Otto-kaggle-3rd-solution-makotupart</a></p>\n<hr>\n<h2>Shimacos part</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F397595%2F9602e046e48f6d2028d0a05ccf898720%2Fsolution_shimacos.png?generation=1675249326142169&amp;alt=media\" alt=\"\"></p>\n<h3>data</h3>\n<ul>\n<li>First, I used my own original truncate dataset.<ul>\n<li>train: 2022-08-14 ~ 2022-08-21 valid 2022-08-21 ~ 2022-08-28</li>\n<li>train: 2022-08-21 ~ 2022-08-28 test: 2022-08-28 ~ 2022-09-04 (Training iteration is the calculated one in the previous step)</li></ul></li>\n<li>And Roughly 10 or so models were made and stacking. (Single 0.592 -&gt; 0.595)</li>\n<li>However since everyone except me was using the radek dataset, I shifted to using that one after team merge. This was to maximize ensemble results.</li>\n</ul>\n<h3>preprocess</h3>\n<ul>\n<li>made aid vector<ul>\n<li>w2v with gensim</li>\n<li>n2v with pytorch_geometric</li></ul></li>\n<li>made session, aid vector<ul>\n<li>bpr with implicit</li></ul></li>\n</ul>\n<h3>Candidates</h3>\n<ul>\n<li>As shown in the overview figure.</li>\n<li>I created candidates and features by BigQuery. For example, the covisit candidates can be calculated in less than one minute for both train and test.</li>\n<li>Recall<ul>\n<li>average candidate count: 188.69</li>\n<li>click: 0.6799</li>\n<li>cart: 0.5555</li>\n<li>order: 0.7359</li>\n<li>Overall: 0.6762</li></ul></li>\n</ul>\n<h3>Feature</h3>\n<ul>\n<li>As shown in the overview figure.</li>\n<li>I used the overall average for smoothing when calculating cvr.</li>\n<li>High importance feature (CatBoost LossFunctionChange importance)<ul>\n<li>difference between the last time the aid was actioned across the overall data and the last time of the session.</li>\n<li>candidate feature by next visit covisit</li>\n<li>I think it was because in the end it was important to hit the next aid after the last action.</li>\n<li>difference feature of w2v and bpr</li>\n<li>covisit cvr feature</li></ul></li>\n</ul>\n<h3>Model</h3>\n<ul>\n<li>LightGBM (lambdarank) and CatBoost (PairLogit)</li>\n<li>I used CatBoost the day before the last day and found an improvement of about 0.002-0.003 over LightGBM. I should have used it earlier…<ul>\n<li>I got 0.596 and 0.598 at private LB finally.</li></ul></li>\n</ul>\n<h3>Stacking</h3>\n<ul>\n<li>At the team merge, the scores were only 0.600 (makotu), 0.600 (sirius),0.598 (Alvor and Shimacos), but By rank blending each score as follows, we were able to produce a 0.603, which was the second at the time.<ul>\n<li>Use the top50 of each prediction.</li>\n<li>And the each prediction was simply outer joined to create a pair of user items.</li>\n<li>fillna with 1 / 50</li>\n<li>rank blending with formula <code>1/rank_a + 1/rank_b + 1/rank_c</code> and sort descending.</li></ul></li>\n<li>As the number of models increased, it became difficult to find the optimal weights, so stacking was performed.<ul>\n<li>Dataset is the same as above.</li>\n<li>Feature: prediction score and rank</li></ul></li>\n<li>This resulted in a score of 0.604 by LightGBM. (CV: 0.59624)</li>\n<li>The same was created in CatBoost (CV: 0.59592) and averaged.</li>\n<li>Finally, I interpolated predictions by aid with the highest number of actions for each type in the test period.</li>\n<li>Random Thought<ul>\n<li>Too many more models could lower the score. This is probably due to too many negative examples.</li>\n<li>Stacking models created by different candidates tended to increase the overall score. </li></ul></li>\n</ul>\n<hr>\n<h2>Sirius part</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F397595%2Ffa7ab3c2005f4af8929e6ed5e7b6f42b%2Fsolution_sirius.png?generation=1675249468500843&amp;alt=media\" alt=\"\"></p>\n<h3>Recalling</h3>\n<ul>\n<li>All the recalling methods are listed in the figure.</li>\n<li>Avg candidates cnt for each user: 224.74</li>\n<li>Overall recall: 0.677232,<ul>\n<li>Clicks recall:0.688493</li>\n<li>Carts recall:0.557334</li>\n<li>Orders recall:0.735304</li></ul></li>\n</ul>\n<h3>Ranking</h3>\n<h4>Samples</h4>\n<ul>\n<li>Click model: the generated candidates which the user clicked are labeled with 1</li>\n<li>Cart/Order model: the generated candidates which the user carted or ordered are labeled with 1. I merged cart and order labels because of their sharing of similar semantic</li>\n<li>Downsampling of negtive samples is abopted. I just keep the negative samples with amount of <code>20*len(pos_samples)</code></li>\n</ul>\n<h4>Features</h4>\n<ul>\n<li>All the features are listed in the figure.</li>\n<li>The feature with highest importance in my model is BPR①, implemented with Implicit package.  (<a href=\"https://www.kaggle.com/code/sirius81/bpr-feature/notebook\" target=\"_blank\">Notebook</a> will be public after everything is verified.)</li>\n<li>Bi-gram② is a terminology of NLP, meaning the successive visiting of aid1, aid2 here. It is similar to co-visit matrix, but it only counts the next one action. And I normlize the count by the hotness of item then. This group of features gave the biggest boost among all my experiments, ~0.006.  (<a href=\"https://www.kaggle.com/code/sirius81/otto-bigram-feature/notebook\" target=\"_blank\">Notebook</a> will be public after everything is verified.)</li>\n<li>OOF of other action③ help improve my score ~0.001.</li>\n<li>Features ①+③+ shimacos’s w2v improved Makotu's CV ~0.001 and features ①+② improved Alvor's CV ~0.003.</li>\n</ul>\n<h4>Models</h4>\n<ul>\n<li>Catboost is trained with Logloss. When training Cart/Order model, I set different weights to positive sample, 5 and 10 for is_carts and is_orders respectively. Along with the strategy of mergeing Cart labels and Order labels, it boosts CV ~0.002, compared with training Cart model and Order model separately.</li>\n<li>MLP is trained with the same features as catboost, plus an standard scaler to normlize the features. But it is not accurate and diverse enough to improve the blending score. So in our final submission, it is not included. Beside MLP, I have tried to introduce GRU and pretrained item embedding into the NN model, but all failed.</li>\n</ul>\n<h3>Ensemble</h3>\n<p>That’s <a href=\"https://www.kaggle.com/shimacos\" target=\"_blank\">@shimacos</a> ’s credit. Very impressive work! </p>",
      "rawMarkdown": "**PS: This solution is actually equivalent to 3rd place, but it was ban because a cheater was on the team. So please enjoy it as an \"imaginary\" 3rd place solution.**\n\nFirst of all, thank you for organizing a great competition. I would also like to thank everyone who participated in the competition with us and the four of us who worked together as a team.\nAbout one week before the end of the competition, we formed a team. However, just on the next day, we were warned one of our teammate might have some problem([link](https://www.kaggle.com/competitions/otto-recommender-system/discussion/381321)). And since then, we other four were working hard together to achieve the best possible result. \nWe can proudly tell you that this 3rd place score is definitely a result with no outside influence. (Of course, as we have already said, we leave it up to the kaggle team to decide how our results will ultimately be handled.)\nIn this discussion, I would be happy to share with you the solutions of the four of us.\nOf course we can share the code as well, but we hope we can open it after the rankings are finalized.\n(The code is not yet public at this time, so it would be difficult for cheaters to imitate our accuracy.)\n\nWe write the solution each one below.\n\n-----------------------------------------------------------------------\n\n## Alvor part\n### Candidates selection:\n- all items from the session's history\n- items from co-visitation matrices. I did many experiments (window size, weights, action types etc.) and finally end up with 5 best matrices (in terms of maximum possible recall when using each matrix). Each of these 5 matrices I use both for candidates selection and for feature engineering (sum/avg/max candidate score, etc.)\n- I did not have fixed number of candidates per session\n- median number of candidates per session: 67\n- but there were sessions with very high number of candidates, so the average candidates number per session is 121.46\n- my recalls per target:\nclicks recall: 0.631531\ncarts recall: 0.540274\norders recall: 0.729626\nTotal: 0.6630\nWhat did not work for me:\n- add candidates from Word2Vec embeddings (K nearest neighbours)\n- add candidates from Matrix Factorization embeddings (K nearest neighbours)\n- simply increase number of candidates in my methods\n- all these attempts increased the maximum possible recall, but decreased the validation score\n### Feature engineering:\n- Sessions features\n- Items features\n- Items-2-Sessions interactions features\n- Co-Visitation matrices features (sum/avg/max of candidate score, etc.)\n- Word2Vec features: euclidean and angular distance between the candidate's embedding and session's last item's embedding\n- Matrix Factorization features: euclidean and angular distance between the candidate's embedding and last session's item's embedding\n- Features selection based on Feature importance of my experimental models\nSome of my best candidates:\n- index of last interaction with the candidate in the session\n- relative number of interactions with given candidate among all session's interactions\n- timestamp differences\n- angular distance between candidate's w2v embedding and last session's item w2v embedding\n- number of co-visitation matrices containing given item as a candidate\nWhat did not work for me:\n- Word2Vec and Matrix Factorization features: euclidean and angular distance between the candidate's embedding and session's penultimate item's embedding\n- Word2Vec and Matrix Factorization features: average/min euclidean and angular distance between the candidate's embedding and all session's items embeddings\n- Word2Vec and Matrix Factorization features: euclidean and angular distance between the candidate's embedding and session's last cart/order item's embedding (there was a minor improvement, but I abandoned these features)\n### Model\n- Group 3-Fold (by sessions) LightGBM Classifier (binary_logloss)\n- One model per one target (clicks/carts/orders)\n- No negative downsample (during almost the entire competition) so it required a lot of memory and time\n- 1/5 negative downsample for carts/orders and 1/15 for clicks (last week of competition)\n- But I use another trick to reduce memory/time usage: When I train a model for carts/orders, I use only sessions for which I have at least one ground-truth candidate among my candidates. Because the main goal of model is to distinguish between good candidates and bad candidates. So I think that sessions with only bad candidates don't help my model.\n  (It could be by 2 reasons:\n  1) the session has not ground true cart/order at all\n  2) the session has ground-true cart/order, but I didn't select it among my candidates)\n- 'dart' boosting for carts nodel, 'gbdt' boosting for clicks and orders model.\n- add \"carts\" prediction as a feature for \"orders\" model (the best feature)\nWhat did not work for me:\n- Add second click to the positive samples as well as first click\n- add \"clicks\" prediction as a feature for \"orders\" model\n- multiple attempts to build a second model to re-arrange only top-X candidates from the first model\n\n-----------------------------------------------------------------------\n\n## Makotu part\nMy pipeline consists mainly of preprocess/make co-matrix/candidate / make feature / modeling parts.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F397595%2Fcff84b84bab47c015ef64ab84e20de5d%2Fmakotu_solution_v2.png?generation=1675314772195222&alt=media)\n\n### Data\n- I used Radek's train/validation/test data\n### Preprocess\n- make aid vector using w2v\n- make BPR(Bayesian Personalized Ranking) feature (recommended Sirius)\n### Make-co-matrix and candidate\n- I made many pattern co-matrix. Basically, co-matrix is created in the  form of probabilities. For example, the following pattern is used to quantify the relationship between aidA and aidB.\n    - aid B count within same session / all aid A count\n    - aid B count within same session (after aid A click) / all aid A count\n    - aid B count within same session / all aid A count. (But if the same user has the same pair of idA and idB, make them unique and then aggregate)\n    - aid B count within same session / all aid A count. (with time weighted)\n    - aid B count within same session and within 1 hour / all aid A count.\n    like this.\n- Select a candidate based on session last action aid / top action aid / with in 1 hour action aid / within 1 day action aid\n- candidate recall\n    - order: 0.732 (@190 candidates)\n    - cart: 0.546 (@190 candidates)\n    - click: 0.682 (@140 candidates)\n### Make feature\n- User / aid / user&aid interaction features.\n- features that worked\n    - last action aid\n    - distance from last aid w2v to candidate aid w2v\n    - BPR feature (recommened sirius)\n    - co-matrix features (ex. probability of clicking on candidate aid after the last aid)\n    - etc\n### Modeling\n- Catboost (loss: PairLogitPairwise)\n- Group Kfold(group: session)\n- Add oof features in addition to the above features\n    - click model does not use oof.\n    However, two patterns of click models were created. in addition to predicting the next clicked id, it also predicts all subsequent clicks of the id, and then uses as oof.\n    - cart model add above 2 model oof.\n    - order model add click 2 model oofs, cart model's oof\n- CV click: 0.5632 cart: 0.4429 order: 0.6699\n- LB: 0.602\n### Stacking\nI was not involved in the stacking of results with other members, as all the stacking was carried out by shimacos with great skill. I hope you will wait for his addition. \n### (add) code of my part\nhttps://github.com/makotu1208/Otto-kaggle-3rd-solution-makotupart\n\n-----------------------------------------------------------------------\n\n## Shimacos part\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F397595%2F9602e046e48f6d2028d0a05ccf898720%2Fsolution_shimacos.png?generation=1675249326142169&alt=media)\n\n### data\n- First, I used my own original truncate dataset.\n  - train: 2022-08-14 ~ 2022-08-21 valid 2022-08-21 ~ 2022-08-28\n  - train: 2022-08-21 ~ 2022-08-28 test: 2022-08-28 ~ 2022-09-04 (Training iteration is the calculated one in the previous step)\n- And Roughly 10 or so models were made and stacking. (Single 0.592 -> 0.595)\n- However since everyone except me was using the radek dataset, I shifted to using that one after team merge. This was to maximize ensemble results.\n### preprocess\n- made aid vector\n  - w2v with gensim\n  - n2v with pytorch_geometric\n- made session, aid vector\n  - bpr with implicit\n### Candidates\n- As shown in the overview figure.\n- I created candidates and features by BigQuery. For example, the covisit candidates can be calculated in less than one minute for both train and test.\n- Recall\n  - average candidate count: 188.69\n  - click: 0.6799\n  - cart: 0.5555\n  - order: 0.7359\n  - Overall: 0.6762\n### Feature\n- As shown in the overview figure.\n- I used the overall average for smoothing when calculating cvr.\n- High importance feature (CatBoost LossFunctionChange importance)\n  - difference between the last time the aid was actioned across the overall data and the last time of the session.\n  - candidate feature by next visit covisit\n    - I think it was because in the end it was important to hit the next aid after the last action.\n  - difference feature of w2v and bpr\n  - covisit cvr feature\n### Model\n- LightGBM (lambdarank) and CatBoost (PairLogit)\n- I used CatBoost the day before the last day and found an improvement of about 0.002-0.003 over LightGBM. I should have used it earlier...\n  - I got 0.596 and 0.598 at private LB finally.\n### Stacking\n- At the team merge, the scores were only 0.600 (makotu), 0.600 (sirius),0.598 (Alvor and Shimacos), but By rank blending each score as follows, we were able to produce a 0.603, which was the second at the time.\n  - Use the top50 of each prediction.\n  - And the each prediction was simply outer joined to create a pair of user items.\n  - fillna with 1 / 50\n  - rank blending with formula `1/rank_a + 1/rank_b + 1/rank_c ` and sort descending.\n- As the number of models increased, it became difficult to find the optimal weights, so stacking was performed.\n  - Dataset is the same as above.\n  - Feature: prediction score and rank\n- This resulted in a score of 0.604 by LightGBM. (CV: 0.59624)\n- The same was created in CatBoost (CV: 0.59592) and averaged.\n- Finally, I interpolated predictions by aid with the highest number of actions for each type in the test period.\n- Random Thought\n  - Too many more models could lower the score. This is probably due to too many negative examples.\n  - Stacking models created by different candidates tended to increase the overall score. \n\n----------------------------------------------------------------------- \n\n## Sirius part\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F397595%2Ffa7ab3c2005f4af8929e6ed5e7b6f42b%2Fsolution_sirius.png?generation=1675249468500843&alt=media)\n\n### Recalling\n- All the recalling methods are listed in the figure.\n- Avg candidates cnt for each user: 224.74\n- Overall recall: 0.677232,\n    - Clicks recall:0.688493\n    - Carts recall:0.557334\n    - Orders recall:0.735304\n### Ranking\n#### Samples\n- Click model: the generated candidates which the user clicked are labeled with 1\n- Cart/Order model: the generated candidates which the user carted or ordered are labeled with 1. I merged cart and order labels because of their sharing of similar semantic\n- Downsampling of negtive samples is abopted. I just keep the negative samples with amount of `20*len(pos_samples)`\n\n#### Features\n- All the features are listed in the figure.\n- The feature with highest importance in my model is BPR①, implemented with Implicit package.  ([Notebook](https://www.kaggle.com/code/sirius81/bpr-feature/notebook) will be public after everything is verified.)\n- Bi-gram② is a terminology of NLP, meaning the successive visiting of aid1, aid2 here. It is similar to co-visit matrix, but it only counts the next one action. And I normlize the count by the hotness of item then. This group of features gave the biggest boost among all my experiments, ~0.006.  ([Notebook](https://www.kaggle.com/code/sirius81/otto-bigram-feature/notebook) will be public after everything is verified.)\n- OOF of other action③ help improve my score ~0.001.\n- Features ①+③+ shimacos’s w2v improved Makotu's CV ~0.001 and features ①+② improved Alvor's CV ~0.003.\n\n#### Models\n- Catboost is trained with Logloss. When training Cart/Order model, I set different weights to positive sample, 5 and 10 for is_carts and is_orders respectively. Along with the strategy of mergeing Cart labels and Order labels, it boosts CV ~0.002, compared with training Cart model and Order model separately.\n- MLP is trained with the same features as catboost, plus an standard scaler to normlize the features. But it is not accurate and diverse enough to improve the blending score. So in our final submission, it is not included. Beside MLP, I have tried to introduce GRU and pretrained item embedding into the NN model, but all failed.\n### Ensemble\nThat’s @shimacos ’s credit. Very impressive work!",
      "votes": null
    },
    {
      "id": "2125148",
      "postDate": "02/01/2023 13:15:47",
      "content": "<p>Thanks for the write up</p>\n<p><code>As the number of models increased, it became difficult to find the optimal weights, so stacking was performed.\nDataset is the same as above.\nFeature: prediction score and rank</code></p>\n<p>Which dataset was used for stacking? Was there a OOF set?</p>",
      "rawMarkdown": "Thanks for the write up\n\n`As the number of models increased, it became difficult to find the optimal weights, so stacking was performed.\nDataset is the same as above.\nFeature: prediction score and rank`\n\nWhich dataset was used for stacking? Was there a OOF set?",
      "votes": null
    },
    {
      "id": "2125159",
      "postDate": "02/01/2023 13:24:32",
      "content": "<p>Yes, I used oof of radek dataset.<br>\nThis was possible because our team all used Redek's dataset.</p>",
      "rawMarkdown": "Yes, I used oof of radek dataset.\nThis was possible because our team all used Redek's dataset.",
      "votes": null
    },
    {
      "id": "2125165",
      "postDate": "02/01/2023 13:44:21",
      "content": "<p>Are you not using Radek's 4th week truncated dataset to train your individual reranker models? Makotu, Alvor and Sirius Model? On which dataset do you train these models and on which is the ensembled trained?</p>",
      "rawMarkdown": "Are you not using Radek's 4th week truncated dataset to train your individual reranker models? Makotu, Alvor and Sirius Model? On which dataset do you train these models and on which is the ensembled trained?",
      "votes": null
    },
    {
      "id": "2125172",
      "postDate": "02/01/2023 13:52:59",
      "content": "<p>We used Radek's 4th week truncated dataset to train by GroupKFold.<br>\nThen built stacking model by the oof prediction.<br>\nThe following is a section on how to create the dataset.</p>\n<blockquote>\n  <p>Use the top50 of each prediction.<br>\n  And the each prediction was simply outer joined to create a pair of user items.</p>\n</blockquote>",
      "rawMarkdown": "We used Radek's 4th week truncated dataset to train by GroupKFold.\nThen built stacking model by the oof prediction.\nThe following is a section on how to create the dataset.\n> Use the top50 of each prediction.\nAnd the each prediction was simply outer joined to create a pair of user items.",
      "votes": null
    },
    {
      "id": "2125188",
      "postDate": "02/01/2023 14:05:04",
      "content": "<p>Wow, so much content here. Great solution, congrats guys!</p>",
      "rawMarkdown": "Wow, so much content here. Great solution, congrats guys!",
      "votes": null
    },
    {
      "id": "2125266",
      "postDate": "02/01/2023 14:56:05",
      "content": "<p>No matter what the final result will be, cooperating with you guys is super pleasant. Thank you very much!</p>",
      "rawMarkdown": "No matter what the final result will be, cooperating with you guys is super pleasant. Thank you very much!",
      "votes": null
    },
    {
      "id": "2125526",
      "postDate": "02/01/2023 17:34:52",
      "content": "<p>Great explanation a lot to learn here .. specially use of angular distance (which we havent used and weren't aware of). Lessons to take in future comps . Also thanks <a href=\"https://www.kaggle.com/sirius81\" target=\"_blank\">@sirius81</a> for your tips from your previous H&amp;M post helped us a lot for this competition :) </p>",
      "rawMarkdown": "Great explanation a lot to learn here .. specially use of angular distance (which we havent used and weren't aware of). Lessons to take in future comps . Also thanks @sirius81 for your tips from your previous H&M post helped us a lot for this competition :)",
      "votes": null
    },
    {
      "id": "2125861",
      "postDate": "02/02/2023 00:13:53",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/gauravbrills\" target=\"_blank\">@gauravbrills</a> , thanks for your comments. Very glad to hear that!</p>",
      "rawMarkdown": "Hi @gauravbrills , thanks for your comments. Very glad to hear that!",
      "votes": null
    },
    {
      "id": "2125944",
      "postDate": "02/02/2023 02:42:55",
      "content": "<p>I admire your team! Thank you for your hard work and nice sharing.</p>",
      "rawMarkdown": "I admire your team! Thank you for your hard work and nice sharing.",
      "votes": null
    },
    {
      "id": "2125996",
      "postDate": "02/02/2023 03:39:59",
      "content": "<p>Will you provide the code later? I would like to learn the details of the processing inside</p>",
      "rawMarkdown": "Will you provide the code later? I would like to learn the details of the processing inside",
      "votes": null
    },
    {
      "id": "2126073",
      "postDate": "02/02/2023 04:42:24",
      "content": "<p>Thanks for sharing and great team from 4 of you!! I hope Kaggle makes fair decision. </p>\n<blockquote>\n  <p>Bi-gram② is a terminology of NLP, meaning the successive visiting of here. It is similar to co-visit matrix, but it only counts the next one action.</p>\n</blockquote>\n<p>Is it similar(or equal) to co-occurence i.e. count of (item, next_item)?</p>",
      "rawMarkdown": "Thanks for sharing and great team from 4 of you!! I hope Kaggle makes fair decision. \n>Bi-gram② is a terminology of NLP, meaning the successive visiting of here. It is similar to co-visit matrix, but it only counts the next one action.\n\n\nIs it similar(or equal) to co-occurence i.e. count of (item, next_item)?",
      "votes": null
    },
    {
      "id": "2126081",
      "postDate": "02/02/2023 04:52:30",
      "content": "<p>Yes, Exactly.</p>",
      "rawMarkdown": "Yes, Exactly.",
      "votes": null
    },
    {
      "id": "2127822",
      "postDate": "02/03/2023 08:10:00",
      "content": "<p>Thank you for sharing! Really impressed by stacking, bi-gram, and the speed of BigQuery. <br>\nBesides BigQuery, what computational resources did you use, and what were the main considerations?</p>",
      "rawMarkdown": "Thank you for sharing! Really impressed by stacking, bi-gram, and the speed of BigQuery. \nBesides BigQuery, what computational resources did you use, and what were the main considerations?",
      "votes": null
    },
    {
      "id": "2127855",
      "postDate": "02/03/2023 08:52:03",
      "content": "<p>Thanks for sharing. Quite Remarkable.</p>",
      "rawMarkdown": "Thanks for sharing. Quite Remarkable.",
      "votes": null
    },
    {
      "id": "2127864",
      "postDate": "02/03/2023 09:26:03",
      "content": "<p>Impressive !</p>",
      "rawMarkdown": "Impressive !",
      "votes": null
    },
    {
      "id": "2129644",
      "postDate": "02/04/2023 19:12:40",
      "content": "<p>congratulations <a href=\"https://www.kaggle.com/mhyodo\" target=\"_blank\">@mhyodo</a> !!</p>",
      "rawMarkdown": "congratulations @mhyodo !!",
      "votes": null
    },
    {
      "id": "2131791",
      "postDate": "02/06/2023 11:24:09",
      "content": "<p><a href=\"https://www.kaggle.com/code/sirius81/otto-bigram-feature/notebook\" target=\"_blank\">https://www.kaggle.com/code/sirius81/otto-bigram-feature/notebook</a>  404😂</p>",
      "rawMarkdown": "https://www.kaggle.com/code/sirius81/otto-bigram-feature/notebook  404😂",
      "votes": null
    },
    {
      "id": "2131801",
      "postDate": "02/06/2023 11:42:52",
      "content": "<p>\"<a href=\"https://www.kaggle.com/code/sirius81/otto-bigram-feature/notebook\" target=\"_blank\">Notebook</a> will be public after everything is verified.\"  But I have added you in the sharing list now.</p>",
      "rawMarkdown": "\"[Notebook](https://www.kaggle.com/code/sirius81/otto-bigram-feature/notebook) will be public after everything is verified.\"  But I have added you in the sharing list now.",
      "votes": null
    },
    {
      "id": "2132694",
      "postDate": "02/07/2023 02:13:18",
      "content": "<p>Thanks, I learned a lot of new things from your code</p>",
      "rawMarkdown": "Thanks, I learned a lot of new things from your code",
      "votes": null
    },
    {
      "id": "2143194",
      "postDate": "02/14/2023 05:55:12",
      "content": "<p>excuse me, may i ask how to determine the w2v model trained well, i train the w2v model for recall and cv top 200 get round 0.600, and i use the w2v embedding for rerank but my rerank doesn't work, i have not find the reason,and i think my w2v embedding is not good, so how to consider it</p>",
      "rawMarkdown": "excuse me, may i ask how to determine the w2v model trained well, i train the w2v model for recall and cv top 200 get round 0.600, and i use the w2v embedding for rerank but my rerank doesn't work, i have not find the reason,and i think my w2v embedding is not good, so how to consider it",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2125148,
      "author_name": "benediktschifferer",
      "author_url": "",
      "post_date": "02/01/2023 13:15:47",
      "content": "<p>Thanks for the write up</p>\n<p><code>As the number of models increased, it became difficult to find the optimal weights, so stacking was performed.\nDataset is the same as above.\nFeature: prediction score and rank</code></p>\n<p>Which dataset was used for stacking? Was there a OOF set?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2125159,
          "author_name": "shimacos",
          "author_url": "",
          "post_date": "02/01/2023 13:24:32",
          "content": "<p>Yes, I used oof of radek dataset.<br>\nThis was possible because our team all used Redek's dataset.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2125165,
              "author_name": "benediktschifferer",
              "author_url": "",
              "post_date": "02/01/2023 13:44:21",
              "content": "<p>Are you not using Radek's 4th week truncated dataset to train your individual reranker models? Makotu, Alvor and Sirius Model? On which dataset do you train these models and on which is the ensembled trained?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2125172,
                  "author_name": "shimacos",
                  "author_url": "",
                  "post_date": "02/01/2023 13:52:59",
                  "content": "<p>We used Radek's 4th week truncated dataset to train by GroupKFold.<br>\nThen built stacking model by the oof prediction.<br>\nThe following is a section on how to create the dataset.</p>\n<blockquote>\n  <p>Use the top50 of each prediction.<br>\n  And the each prediction was simply outer joined to create a pair of user items.</p>\n</blockquote>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2125188,
      "author_name": "michau96",
      "author_url": "",
      "post_date": "02/01/2023 14:05:04",
      "content": "<p>Wow, so much content here. Great solution, congrats guys!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2125266,
      "author_name": "sirius81",
      "author_url": "",
      "post_date": "02/01/2023 14:56:05",
      "content": "<p>No matter what the final result will be, cooperating with you guys is super pleasant. Thank you very much!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2125526,
      "author_name": "gauravbrills",
      "author_url": "",
      "post_date": "02/01/2023 17:34:52",
      "content": "<p>Great explanation a lot to learn here .. specially use of angular distance (which we havent used and weren't aware of). Lessons to take in future comps . Also thanks <a href=\"https://www.kaggle.com/sirius81\" target=\"_blank\">@sirius81</a> for your tips from your previous H&amp;M post helped us a lot for this competition :) </p>",
      "votes": null,
      "replies": [
        {
          "id": 2125861,
          "author_name": "sirius81",
          "author_url": "",
          "post_date": "02/02/2023 00:13:53",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/gauravbrills\" target=\"_blank\">@gauravbrills</a> , thanks for your comments. Very glad to hear that!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2125944,
      "author_name": "songwonho",
      "author_url": "",
      "post_date": "02/02/2023 02:42:55",
      "content": "<p>I admire your team! Thank you for your hard work and nice sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2125996,
      "author_name": "yasso1",
      "author_url": "",
      "post_date": "02/02/2023 03:39:59",
      "content": "<p>Will you provide the code later? I would like to learn the details of the processing inside</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2126073,
      "author_name": "bibek777",
      "author_url": "",
      "post_date": "02/02/2023 04:42:24",
      "content": "<p>Thanks for sharing and great team from 4 of you!! I hope Kaggle makes fair decision. </p>\n<blockquote>\n  <p>Bi-gram② is a terminology of NLP, meaning the successive visiting of here. It is similar to co-visit matrix, but it only counts the next one action.</p>\n</blockquote>\n<p>Is it similar(or equal) to co-occurence i.e. count of (item, next_item)?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2126081,
          "author_name": "sirius81",
          "author_url": "",
          "post_date": "02/02/2023 04:52:30",
          "content": "<p>Yes, Exactly.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2127822,
      "author_name": "sunoonlee",
      "author_url": "",
      "post_date": "02/03/2023 08:10:00",
      "content": "<p>Thank you for sharing! Really impressed by stacking, bi-gram, and the speed of BigQuery. <br>\nBesides BigQuery, what computational resources did you use, and what were the main considerations?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2127855,
      "author_name": "zhenwangryan",
      "author_url": "",
      "post_date": "02/03/2023 08:52:03",
      "content": "<p>Thanks for sharing. Quite Remarkable.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2127864,
      "author_name": "alexpierron",
      "author_url": "",
      "post_date": "02/03/2023 09:26:03",
      "content": "<p>Impressive !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2129644,
      "author_name": "sohilsharma1996",
      "author_url": "",
      "post_date": "02/04/2023 19:12:40",
      "content": "<p>congratulations <a href=\"https://www.kaggle.com/mhyodo\" target=\"_blank\">@mhyodo</a> !!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2131791,
      "author_name": "yasso1",
      "author_url": "",
      "post_date": "02/06/2023 11:24:09",
      "content": "<p><a href=\"https://www.kaggle.com/code/sirius81/otto-bigram-feature/notebook\" target=\"_blank\">https://www.kaggle.com/code/sirius81/otto-bigram-feature/notebook</a>  404😂</p>",
      "votes": null,
      "replies": [
        {
          "id": 2131801,
          "author_name": "sirius81",
          "author_url": "",
          "post_date": "02/06/2023 11:42:52",
          "content": "<p>\"<a href=\"https://www.kaggle.com/code/sirius81/otto-bigram-feature/notebook\" target=\"_blank\">Notebook</a> will be public after everything is verified.\"  But I have added you in the sharing list now.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2132694,
              "author_name": "yasso1",
              "author_url": "",
              "post_date": "02/07/2023 02:13:18",
              "content": "<p>Thanks, I learned a lot of new things from your code</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2143194,
      "author_name": "wcqglhf",
      "author_url": "",
      "post_date": "02/14/2023 05:55:12",
      "content": "<p>excuse me, may i ask how to determine the w2v model trained well, i train the w2v model for recall and cv top 200 get round 0.600, and i use the w2v embedding for rerank but my rerank doesn't work, i have not find the reason,and i think my w2v embedding is not good, so how to consider it</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2124992": "**PS: This solution is actually equivalent to 3rd place, but it was ban because a cheater was on the team. So please enjoy it as an \"imaginary\" 3rd place solution.**\n\nFirst of all, thank you for organizing a great competition. I would also like to thank everyone who participated in the competition with us and the four of us who worked together as a team.\nAbout one week before the end of the competition, we formed a team. However, just on the next day, we were warned one of our teammate might have some problem([link](https://www.kaggle.com/competitions/otto-recommender-system/discussion/381321)). And since then, we other four were working hard together to achieve the best possible result. \nWe can proudly tell you that this 3rd place score is definitely a result with no outside influence. (Of course, as we have already said, we leave it up to the kaggle team to decide how our results will ultimately be handled.)\nIn this discussion, I would be happy to share with you the solutions of the four of us.\nOf course we can share the code as well, but we hope we can open it after the rankings are finalized.\n(The code is not yet public at this time, so it would be difficult for cheaters to imitate our accuracy.)\n\nWe write the solution each one below.\n\n-----------------------------------------------------------------------\n\n## Alvor part\n### Candidates selection:\n- all items from the session's history\n- items from co-visitation matrices. I did many experiments (window size, weights, action types etc.) and finally end up with 5 best matrices (in terms of maximum possible recall when using each matrix). Each of these 5 matrices I use both for candidates selection and for feature engineering (sum/avg/max candidate score, etc.)\n- I did not have fixed number of candidates per session\n- median number of candidates per session: 67\n- but there were sessions with very high number of candidates, so the average candidates number per session is 121.46\n- my recalls per target:\nclicks recall: 0.631531\ncarts recall: 0.540274\norders recall: 0.729626\nTotal: 0.6630\nWhat did not work for me:\n- add candidates from Word2Vec embeddings (K nearest neighbours)\n- add candidates from Matrix Factorization embeddings (K nearest neighbours)\n- simply increase number of candidates in my methods\n- all these attempts increased the maximum possible recall, but decreased the validation score\n### Feature engineering:\n- Sessions features\n- Items features\n- Items-2-Sessions interactions features\n- Co-Visitation matrices features (sum/avg/max of candidate score, etc.)\n- Word2Vec features: euclidean and angular distance between the candidate's embedding and session's last item's embedding\n- Matrix Factorization features: euclidean and angular distance between the candidate's embedding and last session's item's embedding\n- Features selection based on Feature importance of my experimental models\nSome of my best candidates:\n- index of last interaction with the candidate in the session\n- relative number of interactions with given candidate among all session's interactions\n- timestamp differences\n- angular distance between candidate's w2v embedding and last session's item w2v embedding\n- number of co-visitation matrices containing given item as a candidate\nWhat did not work for me:\n- Word2Vec and Matrix Factorization features: euclidean and angular distance between the candidate's embedding and session's penultimate item's embedding\n- Word2Vec and Matrix Factorization features: average/min euclidean and angular distance between the candidate's embedding and all session's items embeddings\n- Word2Vec and Matrix Factorization features: euclidean and angular distance between the candidate's embedding and session's last cart/order item's embedding (there was a minor improvement, but I abandoned these features)\n### Model\n- Group 3-Fold (by sessions) LightGBM Classifier (binary_logloss)\n- One model per one target (clicks/carts/orders)\n- No negative downsample (during almost the entire competition) so it required a lot of memory and time\n- 1/5 negative downsample for carts/orders and 1/15 for clicks (last week of competition)\n- But I use another trick to reduce memory/time usage: When I train a model for carts/orders, I use only sessions for which I have at least one ground-truth candidate among my candidates. Because the main goal of model is to distinguish between good candidates and bad candidates. So I think that sessions with only bad candidates don't help my model.\n  (It could be by 2 reasons:\n  1) the session has not ground true cart/order at all\n  2) the session has ground-true cart/order, but I didn't select it among my candidates)\n- 'dart' boosting for carts nodel, 'gbdt' boosting for clicks and orders model.\n- add \"carts\" prediction as a feature for \"orders\" model (the best feature)\nWhat did not work for me:\n- Add second click to the positive samples as well as first click\n- add \"clicks\" prediction as a feature for \"orders\" model\n- multiple attempts to build a second model to re-arrange only top-X candidates from the first model\n\n-----------------------------------------------------------------------\n\n## Makotu part\nMy pipeline consists mainly of preprocess/make co-matrix/candidate / make feature / modeling parts.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F397595%2Fcff84b84bab47c015ef64ab84e20de5d%2Fmakotu_solution_v2.png?generation=1675314772195222&alt=media)\n\n### Data\n- I used Radek's train/validation/test data\n### Preprocess\n- make aid vector using w2v\n- make BPR(Bayesian Personalized Ranking) feature (recommended Sirius)\n### Make-co-matrix and candidate\n- I made many pattern co-matrix. Basically, co-matrix is created in the  form of probabilities. For example, the following pattern is used to quantify the relationship between aidA and aidB.\n    - aid B count within same session / all aid A count\n    - aid B count within same session (after aid A click) / all aid A count\n    - aid B count within same session / all aid A count. (But if the same user has the same pair of idA and idB, make them unique and then aggregate)\n    - aid B count within same session / all aid A count. (with time weighted)\n    - aid B count within same session and within 1 hour / all aid A count.\n    like this.\n- Select a candidate based on session last action aid / top action aid / with in 1 hour action aid / within 1 day action aid\n- candidate recall\n    - order: 0.732 (@190 candidates)\n    - cart: 0.546 (@190 candidates)\n    - click: 0.682 (@140 candidates)\n### Make feature\n- User / aid / user&aid interaction features.\n- features that worked\n    - last action aid\n    - distance from last aid w2v to candidate aid w2v\n    - BPR feature (recommened sirius)\n    - co-matrix features (ex. probability of clicking on candidate aid after the last aid)\n    - etc\n### Modeling\n- Catboost (loss: PairLogitPairwise)\n- Group Kfold(group: session)\n- Add oof features in addition to the above features\n    - click model does not use oof.\n    However, two patterns of click models were created. in addition to predicting the next clicked id, it also predicts all subsequent clicks of the id, and then uses as oof.\n    - cart model add above 2 model oof.\n    - order model add click 2 model oofs, cart model's oof\n- CV click: 0.5632 cart: 0.4429 order: 0.6699\n- LB: 0.602\n### Stacking\nI was not involved in the stacking of results with other members, as all the stacking was carried out by shimacos with great skill. I hope you will wait for his addition. \n### (add) code of my part\nhttps://github.com/makotu1208/Otto-kaggle-3rd-solution-makotupart\n\n-----------------------------------------------------------------------\n\n## Shimacos part\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F397595%2F9602e046e48f6d2028d0a05ccf898720%2Fsolution_shimacos.png?generation=1675249326142169&alt=media)\n\n### data\n- First, I used my own original truncate dataset.\n  - train: 2022-08-14 ~ 2022-08-21 valid 2022-08-21 ~ 2022-08-28\n  - train: 2022-08-21 ~ 2022-08-28 test: 2022-08-28 ~ 2022-09-04 (Training iteration is the calculated one in the previous step)\n- And Roughly 10 or so models were made and stacking. (Single 0.592 -> 0.595)\n- However since everyone except me was using the radek dataset, I shifted to using that one after team merge. This was to maximize ensemble results.\n### preprocess\n- made aid vector\n  - w2v with gensim\n  - n2v with pytorch_geometric\n- made session, aid vector\n  - bpr with implicit\n### Candidates\n- As shown in the overview figure.\n- I created candidates and features by BigQuery. For example, the covisit candidates can be calculated in less than one minute for both train and test.\n- Recall\n  - average candidate count: 188.69\n  - click: 0.6799\n  - cart: 0.5555\n  - order: 0.7359\n  - Overall: 0.6762\n### Feature\n- As shown in the overview figure.\n- I used the overall average for smoothing when calculating cvr.\n- High importance feature (CatBoost LossFunctionChange importance)\n  - difference between the last time the aid was actioned across the overall data and the last time of the session.\n  - candidate feature by next visit covisit\n    - I think it was because in the end it was important to hit the next aid after the last action.\n  - difference feature of w2v and bpr\n  - covisit cvr feature\n### Model\n- LightGBM (lambdarank) and CatBoost (PairLogit)\n- I used CatBoost the day before the last day and found an improvement of about 0.002-0.003 over LightGBM. I should have used it earlier...\n  - I got 0.596 and 0.598 at private LB finally.\n### Stacking\n- At the team merge, the scores were only 0.600 (makotu), 0.600 (sirius),0.598 (Alvor and Shimacos), but By rank blending each score as follows, we were able to produce a 0.603, which was the second at the time.\n  - Use the top50 of each prediction.\n  - And the each prediction was simply outer joined to create a pair of user items.\n  - fillna with 1 / 50\n  - rank blending with formula `1/rank_a + 1/rank_b + 1/rank_c ` and sort descending.\n- As the number of models increased, it became difficult to find the optimal weights, so stacking was performed.\n  - Dataset is the same as above.\n  - Feature: prediction score and rank\n- This resulted in a score of 0.604 by LightGBM. (CV: 0.59624)\n- The same was created in CatBoost (CV: 0.59592) and averaged.\n- Finally, I interpolated predictions by aid with the highest number of actions for each type in the test period.\n- Random Thought\n  - Too many more models could lower the score. This is probably due to too many negative examples.\n  - Stacking models created by different candidates tended to increase the overall score. \n\n----------------------------------------------------------------------- \n\n## Sirius part\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F397595%2Ffa7ab3c2005f4af8929e6ed5e7b6f42b%2Fsolution_sirius.png?generation=1675249468500843&alt=media)\n\n### Recalling\n- All the recalling methods are listed in the figure.\n- Avg candidates cnt for each user: 224.74\n- Overall recall: 0.677232,\n    - Clicks recall:0.688493\n    - Carts recall:0.557334\n    - Orders recall:0.735304\n### Ranking\n#### Samples\n- Click model: the generated candidates which the user clicked are labeled with 1\n- Cart/Order model: the generated candidates which the user carted or ordered are labeled with 1. I merged cart and order labels because of their sharing of similar semantic\n- Downsampling of negtive samples is abopted. I just keep the negative samples with amount of `20*len(pos_samples)`\n\n#### Features\n- All the features are listed in the figure.\n- The feature with highest importance in my model is BPR①, implemented with Implicit package.  ([Notebook](https://www.kaggle.com/code/sirius81/bpr-feature/notebook) will be public after everything is verified.)\n- Bi-gram② is a terminology of NLP, meaning the successive visiting of aid1, aid2 here. It is similar to co-visit matrix, but it only counts the next one action. And I normlize the count by the hotness of item then. This group of features gave the biggest boost among all my experiments, ~0.006.  ([Notebook](https://www.kaggle.com/code/sirius81/otto-bigram-feature/notebook) will be public after everything is verified.)\n- OOF of other action③ help improve my score ~0.001.\n- Features ①+③+ shimacos’s w2v improved Makotu's CV ~0.001 and features ①+② improved Alvor's CV ~0.003.\n\n#### Models\n- Catboost is trained with Logloss. When training Cart/Order model, I set different weights to positive sample, 5 and 10 for is_carts and is_orders respectively. Along with the strategy of mergeing Cart labels and Order labels, it boosts CV ~0.002, compared with training Cart model and Order model separately.\n- MLP is trained with the same features as catboost, plus an standard scaler to normlize the features. But it is not accurate and diverse enough to improve the blending score. So in our final submission, it is not included. Beside MLP, I have tried to introduce GRU and pretrained item embedding into the NN model, but all failed.\n### Ensemble\nThat’s @shimacos ’s credit. Very impressive work!",
    "2125148": "Thanks for the write up\n\n`As the number of models increased, it became difficult to find the optimal weights, so stacking was performed.\nDataset is the same as above.\nFeature: prediction score and rank`\n\nWhich dataset was used for stacking? Was there a OOF set?",
    "2125159": "Yes, I used oof of radek dataset.\nThis was possible because our team all used Redek's dataset.",
    "2125165": "Are you not using Radek's 4th week truncated dataset to train your individual reranker models? Makotu, Alvor and Sirius Model? On which dataset do you train these models and on which is the ensembled trained?",
    "2125172": "We used Radek's 4th week truncated dataset to train by GroupKFold.\nThen built stacking model by the oof prediction.\nThe following is a section on how to create the dataset.\n> Use the top50 of each prediction.\nAnd the each prediction was simply outer joined to create a pair of user items.",
    "2125188": "Wow, so much content here. Great solution, congrats guys!",
    "2125266": "No matter what the final result will be, cooperating with you guys is super pleasant. Thank you very much!",
    "2125526": "Great explanation a lot to learn here .. specially use of angular distance (which we havent used and weren't aware of). Lessons to take in future comps . Also thanks @sirius81 for your tips from your previous H&M post helped us a lot for this competition :)",
    "2125861": "Hi @gauravbrills , thanks for your comments. Very glad to hear that!",
    "2125944": "I admire your team! Thank you for your hard work and nice sharing.",
    "2125996": "Will you provide the code later? I would like to learn the details of the processing inside",
    "2126073": "Thanks for sharing and great team from 4 of you!! I hope Kaggle makes fair decision. \n>Bi-gram② is a terminology of NLP, meaning the successive visiting of here. It is similar to co-visit matrix, but it only counts the next one action.\n\n\nIs it similar(or equal) to co-occurence i.e. count of (item, next_item)?",
    "2126081": "Yes, Exactly.",
    "2127822": "Thank you for sharing! Really impressed by stacking, bi-gram, and the speed of BigQuery. \nBesides BigQuery, what computational resources did you use, and what were the main considerations?",
    "2127855": "Thanks for sharing. Quite Remarkable.",
    "2127864": "Impressive !",
    "2129644": "congratulations @mhyodo !!",
    "2131791": "https://www.kaggle.com/code/sirius81/otto-bigram-feature/notebook  404😂",
    "2131801": "\"[Notebook](https://www.kaggle.com/code/sirius81/otto-bigram-feature/notebook) will be public after everything is verified.\"  But I have added you in the sharing list now.",
    "2132694": "Thanks, I learned a lot of new things from your code",
    "2143194": "excuse me, may i ask how to determine the w2v model trained well, i train the w2v model for recall and cv top 200 get round 0.600, and i use the w2v embedding for rerank but my rerank doesn't work, i have not find the reason,and i think my w2v embedding is not good, so how to consider it"
  },
  "source": "meta"
}