{
  "id": 383769,
  "title": "7th Place Solution",
  "url": "/competitions/otto-recommender-system/writeups/jack-toshi-k-7th-place-solution",
  "author_name": "",
  "post_date": "2023-02-14T02:45:37.633Z",
  "votes": 84,
  "comment_count": 11,
  "views": 0,
  "content": "<p>First of all, thank you for hosting this super exciting competition! And thank you to everyone for sharing many important insights in this competition. Discussions also have been very beneficial to us in our efforts to achieve this grade.</p>\n<h2>1. Overview</h2>\n<p>Our team consists of two members, Jack and toshi_k. Although both of us are competitions grandmaster, we have different strengths. Before making up a team, our approaches were totally different. Ensemble of two approaches cancelled out each weakness and boosted our team to the gold medal.</p>\n<p>Our solution is composed of LightGBM part, GNN (Graph Neural Network) part and ensemble. Jack was in charge of LightGBM part. He trained the best solo model in our team. toshi_k was in charge of GNN part and ensemble. He trained unique models by modern deep learning.  It contributed +0.003 to the team score by ensemble.</p>\n<p>The details of LightGBM part is described in section 2. GNN part is described in section 3. Ensemble method and result are described in section 4.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F169364%2Fd46e627f90f7d1565ff24df3f9076e9f%2Foverall_03.jpg?generation=1675573670780068&amp;alt=media\" alt=\"\"></p>\n<h2>2. LightGBM part (by Jack)</h2>\n<h3>2.1 Candidate Generation</h3>\n<p>Before I go into details, let me just say that the method described here is my own, but when I optimized Recall@20, there was little difference of performance between it and the <a href=\"https://www.kaggle.com/code/cdeotte/compute-validation-score-cv-565\" target=\"_blank\">Chris's Public Notebook</a> on validation. Thus, it is unclear how much of an advantage it may have.</p>\n<p>The basic idea is to approximate a kind of posterior probability (let's call it a candidate score) of aids based on the co-visitation matrix, and the candidates were selected from the top ones. These candidates are common to all three action types.</p>\n<p>The last 10 aids and action types of the session were used to calculate the candidate score. That is, the score can be calculated as follows:</p>\n<p>$$Score(aid) = P(aid | (aid_1, type_1), (aid_2, type_2), …, (aid_{10}, type_{10}))$$</p>\n<p>Here, we forcibly assume independence and conditional independence of each action, similar in concept to Naive Bayes,</p>\n<p>$$\\begin{eqnarray}<br>\nScore(aid)<br>\n&amp;=&amp; P(aid) \\cdot \\frac{P((aid_1, type_1), …, (aid_{10}, type_{10}) | aid)}{P((aid_1, type_1), …, (aid_{10}, type_{10}))} \\\\<br>\n&amp;\\approx&amp; P(aid) \\cdot \\prod_{i=1}^{10} \\frac{P((aid_i, type_i) | aid)}{P(aid_i, type_i)} \\\\<br>\n&amp;=&amp; P(aid) \\cdot \\prod_{i=1}^{10} \\frac{P((aid_i, type_i), aid)}{P(aid_i, type_i) P(aid)}<br>\n\\end{eqnarray}$$</p>\n<p>The term inside the product is the ratio of the joint probability to the product of individual probability, that represents the extent to which P(aid) is enhanced by the observation of (aid_i, type_i). It is quite impossible to assume the above independence with this data, and therefore this value can exceed 1 and is no longer a probability. However, I expected it to work reasonably well in prioritizing candidates.</p>\n<p>Each term in the above equation is obtained by counting the frequency of each aid and co-visitation of aid pairs and dividing by the total number of sessions. In calculating the co-visitation matrix, only interactions within a 24-hour period are counted, and no multiple counts are made within the same session. The period of calculation was the entire period including test data (in inference phase), and co-visitation in both directions was to be counted.</p>\n<p>The co-visitation matrix does not hold for all pairs of aids, but only those that have many co-visitation for each aid. For aid pairs that are not in the co-visitation matrix, the ratio of the joint probability on the right side of the above formula is set to 1, so that they do not affect the score calculation.</p>\n<p>In the actual calculation, the logarithm is taken and further weighted to the most recent action, as follows:<br>\n$$Score(aid) = \\log(P(aid)) + \\sum_{i=1}^{10} \\frac{11-i}{10} \\cdot \\log \\left( \\frac{P((aid_i, type_i), aid)}{P(aid_i, type_i) P(aid)} \\right)$$</p>\n<p>Many other heuristics, such as adding pseudo counts and adjusting by action type, have been incorporated, but they are too complicated to mention here.</p>\n<p>Starting with those with the highest candidate score, the top 200 were taken for training data, the top 300 for inference of test data, and then the already visited aids were added to make the final candidates.</p>\n<p>The recall on validation of the top 200 candidates thus obtained was as follows:</p>\n<ul>\n<li>clicks: 0.697</li>\n<li>carts: 0.559</li>\n<li>orders: 0.736</li>\n</ul>\n<h3>2.2 LightGBM Rerank Model</h3>\n<p>This part is not much different from the methods already shared by others. The rerank model was trained by LightGBM (LamabdaRank), and separate models were built for each action type.</p>\n<p>On validation, the second last week of train set (truncated) was used as training data and the last week of train set (truncated) as validation data.<br>\nWhen inference was made on the test data, a model trained on the last week's data (LightGBM1) and a model trained on the second last week's data (LightGBM2) were built, and their outputs were ensembled by simple average. Since the training data were completely swapped, I expected a reasonable ensemble effect, but in fact it seems that the effect was only slight.</p>\n<p>Most of the features are based on co-visitation matrix, but each aggregation period is separate for training, validation, and test. That is, the co-visitation matrix is created for each of the three different periods, and the features are created, so they are leakage free.</p>\n<p>The total number of features in the final model is 344, as follows:</p>\n<ul>\n<li>session features (32)<ul>\n<li>the number of all actions (1)</li>\n<li>the number of each action (3)</li>\n<li>the number of unique aids in the session (1)</li>\n<li>the number of unique aids of (carts/orders) and the ratio to the above (4)</li>\n<li>the last action type (1)</li>\n<li>the last relative timestamp from the start of the test period (1)</li>\n<li>the number of actions from the last of each action type to the last action of the session (3)</li>\n<li>elapsed time from i-th last action (i=2, …, 10) to the last action of the session (9)</li>\n<li>revisit ratio of all aids by pair of action types (9)</li></ul></li>\n<li>aid features (50)<ul>\n<li>count of (any/buy/click/cart/order) (5)</li>\n<li>exponential decay count of (any/buy/order) (3)</li>\n<li>count of (any/buy/order) in last n days (n=1~7) (21)</li>\n<li>count of (any/buy/order) in last n weeks (n=1~4) (12)</li>\n<li>revisit count in all sessions by pair of action types (9)</li></ul></li>\n<li>session*aid features (12)<ul>\n<li>the latest action of that aid (1)</li>\n<li>the number of each action of that aid (3)</li>\n<li>the number of actions from last visit to that aid to the last action of the session (1)</li>\n<li>elapsed time from last visit to that aid to the last action of the session (1)</li>\n<li>the above two features for each action type (6)</li></ul></li>\n<li>co-visitation features (250)<ul>\n<li>the number of co-visitation of aid with aid_i (i=1, …, 10) devided by the global count of aid_i<ul>\n<li>any to any (both direction/oneway) (20)</li>\n<li>any to buy (both direction/oneway) (20)</li>\n<li>buy to any (both direction/oneway) (20)</li>\n<li>buy to buy (both direction/oneway) (20)</li>\n<li>type_i to any (both direction/oneway) (20)</li>\n<li>type_i to buy (both direction/oneway) (20)</li>\n<li>click to click (both direction) (10)</li>\n<li>click to cart (both direction) (10)</li>\n<li>cart to click (both direction) (10)</li>\n<li>cart to cart (both direction) (10)</li></ul></li>\n<li>the rank of co-visitation of aid with aid_i (i=1, …, 10)<ul>\n<li>any to any (both direction) (10)</li>\n<li>any to buy (both direction) (10)</li>\n<li>buy to any (both direction) (10)</li>\n<li>buy to buy (both direction) (10)</li></ul></li>\n<li>global count of aid_i (any/buy/click/cart/type_i) (50)</li></ul></li>\n</ul>\n<p>* \"any\" means the action clicks or carts or orders, and \"buy\" means the action carts or orders.</p>\n<h2>3. GNN part (by toshi_k)</h2>\n<h3>3.1 Basic Idea</h3>\n<p>I considered using DL (Deep Learning) in this competition. Since the datasets are relatively simple, E2E approach of DL seemed like a desirable solution for me. Another advantage is that multi-dimensional interactions and outputs for clicks/carts/orders are easily designed as a DL model architecture.</p>\n<p>DL based recommendation was initially proposed as a kind of non-linear collaborative filtering. The typical one is training AutoEncoder model and using the reconstruction methodology to evaluate missing ratings.</p>\n<ul>\n<li>Training Deep AutoEncoders for Collaborative Filtering<ul>\n<li><a href=\"https://arxiv.org/abs/1708.01715\" target=\"_blank\">https://arxiv.org/abs/1708.01715</a></li></ul></li>\n</ul>\n<p>Although I implemented this type of method as a prototype, it didn't work well. The number of items was so large that it made input vectors ultra sparse. It also yielded the heavy requirements of GPU memory for FC (Fully Connected) layers and made hidden layers shallower and thinner.</p>\n<p>The disadvantage of FC layers is having weights between all combinations of aids even if most of them have nothing to do with each other. After some considerations, I figured out GNN (Graph Neural Networks) can solve this issue. The graph for GNN can represents aid relations and GNN can predict attributions of aids based on the nearly connected aids.</p>\n<p>Using GNN for session based recommendation is also reported in the below study. According to the paper, their method is developed to explore rich transitions among items and generate accurate latent vectors of items. Their experiments on two datasets including thousands of items show that their method outperforms the state-of-the-art methods.</p>\n<ul>\n<li>Session-based Recommendation with Graph Neural Networks<ul>\n<li><a href=\"https://arxiv.org/abs/1811.00855v4\" target=\"_blank\">https://arxiv.org/abs/1811.00855v4</a></li></ul></li>\n</ul>\n<p>My approach is similar to the previous study. One of the biggest differences of problem setting is the number of items. In this competition, the datasets contain millions of items. To handle all items, I built a simpler workflow and installed the subgraph extraction from the global session graph.</p>\n<p>Basically, my approach has 3 steps.</p>\n<ol>\n<li>Construct the global graph that represents aid relations</li>\n<li>Extract the subgraph from the global graph for each session</li>\n<li>Use GNN to predict which aid will be taken</li>\n</ol>\n<p>The conceptual diagram is as below.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F169364%2F2327fdf46f8bcd96370e1f7e0ebaaf4d%2Fgnn_solution_01.jpg?generation=1675573719262138&amp;alt=media\" alt=\"\"></p>\n<p>In the first step, the global graph that represents aid relations is constructed. A subset of training data is used to create this graph. All transitions in this data are counted up and top P transitions (P=20, 30 or 40) from each aid are adopted as the edges of the graph. This graph is roughly corresponding to the co-visitation matrix that other participants call.</p>\n<p>Secondly, the subgraph is extracted from the global graph for each session. The history of each session is traced and nearly connected nodes (=aids) are listed up. Closer nodes and major transitions are prioritised and hundreds aids are filtered for extraction. This process is roughly corresponding to the candidates generation other participants call.</p>\n<p>Thirdly, subgraphs are used to train and test GNN. The last layer of GNN has three channels. They predict if clicks/carts/orders will be taken in the future of each session. The inputs of GNN are the structures of subgraphs and features of nodes and edges. More details of features and GNN model are described in the next two subsections.</p>\n<h3>3.2 Features</h3>\n<p>The input features for my GNN consist of \"node features\" and \"edge features\". Node features represent the characteristics of each aid. Edge features represent the relations between each pair of aids.</p>\n<p>Basically the total number of node features is 18. Nine of them is the global characteristics of aids. These features are shared among all sessions. The other nine features represent the history of sessions. These features are calculated on the session history and different for every sessions. The list of node features is as below.</p>\n<ul>\n<li>Node features (18)<ul>\n<li>Global aid features (9)<ul>\n<li>Popularity counts (3)</li>\n<li>Repeat counts (3)</li>\n<li>Type transition counts (3)</li></ul></li>\n<li>Session history features (9)<ul>\n<li>Distance from the session history (2)</li>\n<li>Number of counts in the session history (3)</li>\n<li>Visited order features (2)</li>\n<li>Visited time features (2)</li></ul></li></ul></li>\n</ul>\n<p>Total number of edge features is 14. Twelve of them is the global characteristics of transitions. These features are shared among all sessions. The other two features represent the history of sessions. These features are calculated on the session history and different for every sessions. The list of edge features is as below.</p>\n<ul>\n<li>Edge features (14)<ul>\n<li>Global transition features (12)<ul>\n<li>Transition count not considering types (2)</li>\n<li>Transition rank not considering types (2)</li>\n<li>Cart-to-cart transition count (2)</li>\n<li>Cart-to-cart transition rank (2)</li>\n<li>Order-to-order transition count (2)</li>\n<li>Order-to-order transition rank (2)</li></ul></li>\n<li>Session history features (2)<ul>\n<li>Self loop or not (1)</li>\n<li>Stepped in session history or not (1)</li></ul></li></ul></li>\n</ul>\n<p>Any combinations of multiple features and higher dimension features are not added. It was expected that such complex features were automatically captured by the representation capability of GNN.</p>\n<p>All missing values are filled with zero and logarithmic transformation (log1p) is applied to most features for the stability of GNN.</p>\n<h3>3.3 Model and Loss function</h3>\n<p>My GNN has 8 GCN (Graph Convolution) layers. This implies GNN model can consider aids located within 8 steps from the history aids for prediction. In some trial experiments, 8 layers model was better than 4 layers one a little, but it was unclear more layers helped or not. </p>\n<p>Aside from GCN layers, my model employs non linear activation functions, skip connections, and normalization layers. These component made training faster and yielded less training loss.</p>\n<p>As mentioned above, the last layer has three channels for clicks/carts/orders. The channels for carts and orders are connected with sigmoid functions and trained by binary cross entropy loss. The channel for clicks is connected with softmax function and trained by softmax cross entropy loss.</p>\n<h3>3.4 Performance boosting</h3>\n<p>I noticed that increasing the number of candidates for inference boosted the LB score. For example, my best model improved from 0.59474 to 0.59750 on public LB by increasing the candidates from 384 to 1024.</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>num of candidates</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Model1</td>\n<td>384</td>\n<td>0.59474</td>\n<td>0.59451</td>\n</tr>\n<tr>\n<td>Model1</td>\n<td>1024</td>\n<td>0.59750</td>\n<td>0.59719</td>\n</tr>\n</tbody>\n</table>\n<p>Unfortunately, more candidates required more computational calculation, 1024 candidates took several days for inference. I increased candidates as much as possible and save the outputs of several models.</p>\n<p>As the final stage of the competition, ensemble of multiple predictions is attempted. Four GNN models are used for this ensemble. Although they were trained with slightly different setting, their basic approaches were the same. Ensemble of 4 models achieved 0.59894 on public LB and 0.59874 on private LB.</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>num of candidates</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Model1</td>\n<td>1024</td>\n<td>0.59750</td>\n<td>0.59719</td>\n</tr>\n<tr>\n<td>Model2</td>\n<td>1024</td>\n<td>0.59653</td>\n<td>0.59613</td>\n</tr>\n<tr>\n<td>Model3</td>\n<td>512</td>\n<td>0.59320</td>\n<td>0.59318</td>\n</tr>\n<tr>\n<td>Model4</td>\n<td>768</td>\n<td>0.59529</td>\n<td>0.59478</td>\n</tr>\n<tr>\n<td>Ensemble of 4 models</td>\n<td>-</td>\n<td>0.59894</td>\n<td>0.59874</td>\n</tr>\n</tbody>\n</table>\n<h2>4. Team Ensemble (by toshi_k)</h2>\n<p>After we made up a team, we tried several ways to merge our predictions. We started from the simple rerank by arithmetic mean of ranks in submission files. It improved the public score 0.597→0.600 then.</p>\n<p>The next attempt was using raw prediction to boost LB score more. We saved the raw prediction of top 50 aids of each part. The biggest issue was the outputs from LambdaRank range in any real number while the output of GNN range in (0, 1).</p>\n<p>Since we didn't have enough time to lead the best theoretical way, we tried some ensemble method experimentally. One interesting finding was the logit transformed value of GNN seemed to have the proportional relationship with LambdaRank with constant shift.</p>\n<p>$$ \\mathrm{logit}(p^\\text{Binary Prediction}) \\propto v^\\text{LambdaRank prediction} + C $$</p>\n<p>The value of constant shift is different on every session_types. This may be just a brute force approximation, we estimate C for each session_type and mapped the output of LambdaRank to 0-to-1 value.</p>\n<p>$$\\begin{eqnarray}<br>\n\\hat{C} &amp;=&amp; \\arg \\min_C \\{ \\frac{1}{R} \\sum_r^R {^Gp_r} - \\frac{1}{R} \\sum_r^R \\sigma(^Lv_r + C) \\}^2 \\\\<br>\n^Lp_r &amp;=&amp; \\sigma (v^L_r + \\hat{C}) \\\\<br>\n\\mathrm{where:} \\\\<br>\n^Gp_r &amp;=&amp; \\text{rth output of GNN} \\in (0, 1) \\\\\\<br>\n^Lv_r &amp;=&amp; \\text{rth output of LambdaRank LGBM } \\in \\mathbb{R} \\\\\\<br>\n^Lp_r &amp;=&amp; \\text{0-1 calibrated value of } ^Lv_r \\in (0, 1)<br>\n\\end{eqnarray}$$</p>\n<p>After this transformation, simple weighted averaging was calculated. Since LightGBM part achieved better score on PublicLB, the weight of LightGBM is set to be larger than GNN part. We tried two patterns of weight settings and chose both of them as final submissions.</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>weight</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>LightGBM Part</td>\n<td>1.000 (LightGBM) : 0.000 (GNN)</td>\n<td>0.60025</td>\n<td>0.60008</td>\n</tr>\n<tr>\n<td>GNN part</td>\n<td>0.000 (LightGBM) : 1.000 (GNN)</td>\n<td>0.59894</td>\n<td>0.59874</td>\n</tr>\n<tr>\n<td>Final Submission 1</td>\n<td>0.600 (LightGBM) : 0.400 (GNN)</td>\n<td>0.60311</td>\n<td>0.60307</td>\n</tr>\n<tr>\n<td>Final Submission 2</td>\n<td>0.525 (LightGBM) : 0.475 (GNN)</td>\n<td>0.60302</td>\n<td>0.60313</td>\n</tr>\n</tbody>\n</table>\n<p>Both of final submissions improved LB score from the best of two parts. While Final Submission 1 was the best on public LB, Final Submission 2 was the best on private LB. Even though all single models got worse on private LB, the score of Final Submission 2 on private LB is better than public LB. Out team merge and ensemble was successful in this sense.</p>\n<h2>5. Conclusion</h2>\n<p>Our team employed LightGBM and GNN for this competition. Each approach took different way for candidate generation, feature engineering and model design. Ensemble of two types of approaches boosted our team score a lot.</p>\n<p>This competition gave 6th gold model to toshi_k and 7th gold to Jack. It motivate us for the further success. What we learn in this competition can be applied to not only future competitions but real world projects. It was confirmed that Kaggle is the practical platform of data science again.</p>\n<p>Thank you for reading this to the end!</p>",
  "messages": [
    {
      "id": "2130054",
      "postDate": "02/05/2023 05:21:03",
      "content": "<p>First of all, thank you for hosting this super exciting competition! And thank you to everyone for sharing many important insights in this competition. Discussions also have been very beneficial to us in our efforts to achieve this grade.</p>\n<h2>1. Overview</h2>\n<p>Our team consists of two members, Jack and toshi_k. Although both of us are competitions grandmaster, we have different strengths. Before making up a team, our approaches were totally different. Ensemble of two approaches cancelled out each weakness and boosted our team to the gold medal.</p>\n<p>Our solution is composed of LightGBM part, GNN (Graph Neural Network) part and ensemble. Jack was in charge of LightGBM part. He trained the best solo model in our team. toshi_k was in charge of GNN part and ensemble. He trained unique models by modern deep learning.  It contributed +0.003 to the team score by ensemble.</p>\n<p>The details of LightGBM part is described in section 2. GNN part is described in section 3. Ensemble method and result are described in section 4.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F169364%2Fd46e627f90f7d1565ff24df3f9076e9f%2Foverall_03.jpg?generation=1675573670780068&amp;alt=media\" alt=\"\"></p>\n<h2>2. LightGBM part (by Jack)</h2>\n<h3>2.1 Candidate Generation</h3>\n<p>Before I go into details, let me just say that the method described here is my own, but when I optimized Recall@20, there was little difference of performance between it and the <a href=\"https://www.kaggle.com/code/cdeotte/compute-validation-score-cv-565\" target=\"_blank\">Chris's Public Notebook</a> on validation. Thus, it is unclear how much of an advantage it may have.</p>\n<p>The basic idea is to approximate a kind of posterior probability (let's call it a candidate score) of aids based on the co-visitation matrix, and the candidates were selected from the top ones. These candidates are common to all three action types.</p>\n<p>The last 10 aids and action types of the session were used to calculate the candidate score. That is, the score can be calculated as follows:</p>\n<p>$$Score(aid) = P(aid | (aid_1, type_1), (aid_2, type_2), …, (aid_{10}, type_{10}))$$</p>\n<p>Here, we forcibly assume independence and conditional independence of each action, similar in concept to Naive Bayes,</p>\n<p>$$\\begin{eqnarray}<br>\nScore(aid)<br>\n&amp;=&amp; P(aid) \\cdot \\frac{P((aid_1, type_1), …, (aid_{10}, type_{10}) | aid)}{P((aid_1, type_1), …, (aid_{10}, type_{10}))} \\\\<br>\n&amp;\\approx&amp; P(aid) \\cdot \\prod_{i=1}^{10} \\frac{P((aid_i, type_i) | aid)}{P(aid_i, type_i)} \\\\<br>\n&amp;=&amp; P(aid) \\cdot \\prod_{i=1}^{10} \\frac{P((aid_i, type_i), aid)}{P(aid_i, type_i) P(aid)}<br>\n\\end{eqnarray}$$</p>\n<p>The term inside the product is the ratio of the joint probability to the product of individual probability, that represents the extent to which P(aid) is enhanced by the observation of (aid_i, type_i). It is quite impossible to assume the above independence with this data, and therefore this value can exceed 1 and is no longer a probability. However, I expected it to work reasonably well in prioritizing candidates.</p>\n<p>Each term in the above equation is obtained by counting the frequency of each aid and co-visitation of aid pairs and dividing by the total number of sessions. In calculating the co-visitation matrix, only interactions within a 24-hour period are counted, and no multiple counts are made within the same session. The period of calculation was the entire period including test data (in inference phase), and co-visitation in both directions was to be counted.</p>\n<p>The co-visitation matrix does not hold for all pairs of aids, but only those that have many co-visitation for each aid. For aid pairs that are not in the co-visitation matrix, the ratio of the joint probability on the right side of the above formula is set to 1, so that they do not affect the score calculation.</p>\n<p>In the actual calculation, the logarithm is taken and further weighted to the most recent action, as follows:<br>\n$$Score(aid) = \\log(P(aid)) + \\sum_{i=1}^{10} \\frac{11-i}{10} \\cdot \\log \\left( \\frac{P((aid_i, type_i), aid)}{P(aid_i, type_i) P(aid)} \\right)$$</p>\n<p>Many other heuristics, such as adding pseudo counts and adjusting by action type, have been incorporated, but they are too complicated to mention here.</p>\n<p>Starting with those with the highest candidate score, the top 200 were taken for training data, the top 300 for inference of test data, and then the already visited aids were added to make the final candidates.</p>\n<p>The recall on validation of the top 200 candidates thus obtained was as follows:</p>\n<ul>\n<li>clicks: 0.697</li>\n<li>carts: 0.559</li>\n<li>orders: 0.736</li>\n</ul>\n<h3>2.2 LightGBM Rerank Model</h3>\n<p>This part is not much different from the methods already shared by others. The rerank model was trained by LightGBM (LamabdaRank), and separate models were built for each action type.</p>\n<p>On validation, the second last week of train set (truncated) was used as training data and the last week of train set (truncated) as validation data.<br>\nWhen inference was made on the test data, a model trained on the last week's data (LightGBM1) and a model trained on the second last week's data (LightGBM2) were built, and their outputs were ensembled by simple average. Since the training data were completely swapped, I expected a reasonable ensemble effect, but in fact it seems that the effect was only slight.</p>\n<p>Most of the features are based on co-visitation matrix, but each aggregation period is separate for training, validation, and test. That is, the co-visitation matrix is created for each of the three different periods, and the features are created, so they are leakage free.</p>\n<p>The total number of features in the final model is 344, as follows:</p>\n<ul>\n<li>session features (32)<ul>\n<li>the number of all actions (1)</li>\n<li>the number of each action (3)</li>\n<li>the number of unique aids in the session (1)</li>\n<li>the number of unique aids of (carts/orders) and the ratio to the above (4)</li>\n<li>the last action type (1)</li>\n<li>the last relative timestamp from the start of the test period (1)</li>\n<li>the number of actions from the last of each action type to the last action of the session (3)</li>\n<li>elapsed time from i-th last action (i=2, …, 10) to the last action of the session (9)</li>\n<li>revisit ratio of all aids by pair of action types (9)</li></ul></li>\n<li>aid features (50)<ul>\n<li>count of (any/buy/click/cart/order) (5)</li>\n<li>exponential decay count of (any/buy/order) (3)</li>\n<li>count of (any/buy/order) in last n days (n=1~7) (21)</li>\n<li>count of (any/buy/order) in last n weeks (n=1~4) (12)</li>\n<li>revisit count in all sessions by pair of action types (9)</li></ul></li>\n<li>session*aid features (12)<ul>\n<li>the latest action of that aid (1)</li>\n<li>the number of each action of that aid (3)</li>\n<li>the number of actions from last visit to that aid to the last action of the session (1)</li>\n<li>elapsed time from last visit to that aid to the last action of the session (1)</li>\n<li>the above two features for each action type (6)</li></ul></li>\n<li>co-visitation features (250)<ul>\n<li>the number of co-visitation of aid with aid_i (i=1, …, 10) devided by the global count of aid_i<ul>\n<li>any to any (both direction/oneway) (20)</li>\n<li>any to buy (both direction/oneway) (20)</li>\n<li>buy to any (both direction/oneway) (20)</li>\n<li>buy to buy (both direction/oneway) (20)</li>\n<li>type_i to any (both direction/oneway) (20)</li>\n<li>type_i to buy (both direction/oneway) (20)</li>\n<li>click to click (both direction) (10)</li>\n<li>click to cart (both direction) (10)</li>\n<li>cart to click (both direction) (10)</li>\n<li>cart to cart (both direction) (10)</li></ul></li>\n<li>the rank of co-visitation of aid with aid_i (i=1, …, 10)<ul>\n<li>any to any (both direction) (10)</li>\n<li>any to buy (both direction) (10)</li>\n<li>buy to any (both direction) (10)</li>\n<li>buy to buy (both direction) (10)</li></ul></li>\n<li>global count of aid_i (any/buy/click/cart/type_i) (50)</li></ul></li>\n</ul>\n<p>* \"any\" means the action clicks or carts or orders, and \"buy\" means the action carts or orders.</p>\n<h2>3. GNN part (by toshi_k)</h2>\n<h3>3.1 Basic Idea</h3>\n<p>I considered using DL (Deep Learning) in this competition. Since the datasets are relatively simple, E2E approach of DL seemed like a desirable solution for me. Another advantage is that multi-dimensional interactions and outputs for clicks/carts/orders are easily designed as a DL model architecture.</p>\n<p>DL based recommendation was initially proposed as a kind of non-linear collaborative filtering. The typical one is training AutoEncoder model and using the reconstruction methodology to evaluate missing ratings.</p>\n<ul>\n<li>Training Deep AutoEncoders for Collaborative Filtering<ul>\n<li><a href=\"https://arxiv.org/abs/1708.01715\" target=\"_blank\">https://arxiv.org/abs/1708.01715</a></li></ul></li>\n</ul>\n<p>Although I implemented this type of method as a prototype, it didn't work well. The number of items was so large that it made input vectors ultra sparse. It also yielded the heavy requirements of GPU memory for FC (Fully Connected) layers and made hidden layers shallower and thinner.</p>\n<p>The disadvantage of FC layers is having weights between all combinations of aids even if most of them have nothing to do with each other. After some considerations, I figured out GNN (Graph Neural Networks) can solve this issue. The graph for GNN can represents aid relations and GNN can predict attributions of aids based on the nearly connected aids.</p>\n<p>Using GNN for session based recommendation is also reported in the below study. According to the paper, their method is developed to explore rich transitions among items and generate accurate latent vectors of items. Their experiments on two datasets including thousands of items show that their method outperforms the state-of-the-art methods.</p>\n<ul>\n<li>Session-based Recommendation with Graph Neural Networks<ul>\n<li><a href=\"https://arxiv.org/abs/1811.00855v4\" target=\"_blank\">https://arxiv.org/abs/1811.00855v4</a></li></ul></li>\n</ul>\n<p>My approach is similar to the previous study. One of the biggest differences of problem setting is the number of items. In this competition, the datasets contain millions of items. To handle all items, I built a simpler workflow and installed the subgraph extraction from the global session graph.</p>\n<p>Basically, my approach has 3 steps.</p>\n<ol>\n<li>Construct the global graph that represents aid relations</li>\n<li>Extract the subgraph from the global graph for each session</li>\n<li>Use GNN to predict which aid will be taken</li>\n</ol>\n<p>The conceptual diagram is as below.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F169364%2F2327fdf46f8bcd96370e1f7e0ebaaf4d%2Fgnn_solution_01.jpg?generation=1675573719262138&amp;alt=media\" alt=\"\"></p>\n<p>In the first step, the global graph that represents aid relations is constructed. A subset of training data is used to create this graph. All transitions in this data are counted up and top P transitions (P=20, 30 or 40) from each aid are adopted as the edges of the graph. This graph is roughly corresponding to the co-visitation matrix that other participants call.</p>\n<p>Secondly, the subgraph is extracted from the global graph for each session. The history of each session is traced and nearly connected nodes (=aids) are listed up. Closer nodes and major transitions are prioritised and hundreds aids are filtered for extraction. This process is roughly corresponding to the candidates generation other participants call.</p>\n<p>Thirdly, subgraphs are used to train and test GNN. The last layer of GNN has three channels. They predict if clicks/carts/orders will be taken in the future of each session. The inputs of GNN are the structures of subgraphs and features of nodes and edges. More details of features and GNN model are described in the next two subsections.</p>\n<h3>3.2 Features</h3>\n<p>The input features for my GNN consist of \"node features\" and \"edge features\". Node features represent the characteristics of each aid. Edge features represent the relations between each pair of aids.</p>\n<p>Basically the total number of node features is 18. Nine of them is the global characteristics of aids. These features are shared among all sessions. The other nine features represent the history of sessions. These features are calculated on the session history and different for every sessions. The list of node features is as below.</p>\n<ul>\n<li>Node features (18)<ul>\n<li>Global aid features (9)<ul>\n<li>Popularity counts (3)</li>\n<li>Repeat counts (3)</li>\n<li>Type transition counts (3)</li></ul></li>\n<li>Session history features (9)<ul>\n<li>Distance from the session history (2)</li>\n<li>Number of counts in the session history (3)</li>\n<li>Visited order features (2)</li>\n<li>Visited time features (2)</li></ul></li></ul></li>\n</ul>\n<p>Total number of edge features is 14. Twelve of them is the global characteristics of transitions. These features are shared among all sessions. The other two features represent the history of sessions. These features are calculated on the session history and different for every sessions. The list of edge features is as below.</p>\n<ul>\n<li>Edge features (14)<ul>\n<li>Global transition features (12)<ul>\n<li>Transition count not considering types (2)</li>\n<li>Transition rank not considering types (2)</li>\n<li>Cart-to-cart transition count (2)</li>\n<li>Cart-to-cart transition rank (2)</li>\n<li>Order-to-order transition count (2)</li>\n<li>Order-to-order transition rank (2)</li></ul></li>\n<li>Session history features (2)<ul>\n<li>Self loop or not (1)</li>\n<li>Stepped in session history or not (1)</li></ul></li></ul></li>\n</ul>\n<p>Any combinations of multiple features and higher dimension features are not added. It was expected that such complex features were automatically captured by the representation capability of GNN.</p>\n<p>All missing values are filled with zero and logarithmic transformation (log1p) is applied to most features for the stability of GNN.</p>\n<h3>3.3 Model and Loss function</h3>\n<p>My GNN has 8 GCN (Graph Convolution) layers. This implies GNN model can consider aids located within 8 steps from the history aids for prediction. In some trial experiments, 8 layers model was better than 4 layers one a little, but it was unclear more layers helped or not. </p>\n<p>Aside from GCN layers, my model employs non linear activation functions, skip connections, and normalization layers. These component made training faster and yielded less training loss.</p>\n<p>As mentioned above, the last layer has three channels for clicks/carts/orders. The channels for carts and orders are connected with sigmoid functions and trained by binary cross entropy loss. The channel for clicks is connected with softmax function and trained by softmax cross entropy loss.</p>\n<h3>3.4 Performance boosting</h3>\n<p>I noticed that increasing the number of candidates for inference boosted the LB score. For example, my best model improved from 0.59474 to 0.59750 on public LB by increasing the candidates from 384 to 1024.</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>num of candidates</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Model1</td>\n<td>384</td>\n<td>0.59474</td>\n<td>0.59451</td>\n</tr>\n<tr>\n<td>Model1</td>\n<td>1024</td>\n<td>0.59750</td>\n<td>0.59719</td>\n</tr>\n</tbody>\n</table>\n<p>Unfortunately, more candidates required more computational calculation, 1024 candidates took several days for inference. I increased candidates as much as possible and save the outputs of several models.</p>\n<p>As the final stage of the competition, ensemble of multiple predictions is attempted. Four GNN models are used for this ensemble. Although they were trained with slightly different setting, their basic approaches were the same. Ensemble of 4 models achieved 0.59894 on public LB and 0.59874 on private LB.</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>num of candidates</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Model1</td>\n<td>1024</td>\n<td>0.59750</td>\n<td>0.59719</td>\n</tr>\n<tr>\n<td>Model2</td>\n<td>1024</td>\n<td>0.59653</td>\n<td>0.59613</td>\n</tr>\n<tr>\n<td>Model3</td>\n<td>512</td>\n<td>0.59320</td>\n<td>0.59318</td>\n</tr>\n<tr>\n<td>Model4</td>\n<td>768</td>\n<td>0.59529</td>\n<td>0.59478</td>\n</tr>\n<tr>\n<td>Ensemble of 4 models</td>\n<td>-</td>\n<td>0.59894</td>\n<td>0.59874</td>\n</tr>\n</tbody>\n</table>\n<h2>4. Team Ensemble (by toshi_k)</h2>\n<p>After we made up a team, we tried several ways to merge our predictions. We started from the simple rerank by arithmetic mean of ranks in submission files. It improved the public score 0.597→0.600 then.</p>\n<p>The next attempt was using raw prediction to boost LB score more. We saved the raw prediction of top 50 aids of each part. The biggest issue was the outputs from LambdaRank range in any real number while the output of GNN range in (0, 1).</p>\n<p>Since we didn't have enough time to lead the best theoretical way, we tried some ensemble method experimentally. One interesting finding was the logit transformed value of GNN seemed to have the proportional relationship with LambdaRank with constant shift.</p>\n<p>$$ \\mathrm{logit}(p^\\text{Binary Prediction}) \\propto v^\\text{LambdaRank prediction} + C $$</p>\n<p>The value of constant shift is different on every session_types. This may be just a brute force approximation, we estimate C for each session_type and mapped the output of LambdaRank to 0-to-1 value.</p>\n<p>$$\\begin{eqnarray}<br>\n\\hat{C} &amp;=&amp; \\arg \\min_C \\{ \\frac{1}{R} \\sum_r^R {^Gp_r} - \\frac{1}{R} \\sum_r^R \\sigma(^Lv_r + C) \\}^2 \\\\<br>\n^Lp_r &amp;=&amp; \\sigma (v^L_r + \\hat{C}) \\\\<br>\n\\mathrm{where:} \\\\<br>\n^Gp_r &amp;=&amp; \\text{rth output of GNN} \\in (0, 1) \\\\\\<br>\n^Lv_r &amp;=&amp; \\text{rth output of LambdaRank LGBM } \\in \\mathbb{R} \\\\\\<br>\n^Lp_r &amp;=&amp; \\text{0-1 calibrated value of } ^Lv_r \\in (0, 1)<br>\n\\end{eqnarray}$$</p>\n<p>After this transformation, simple weighted averaging was calculated. Since LightGBM part achieved better score on PublicLB, the weight of LightGBM is set to be larger than GNN part. We tried two patterns of weight settings and chose both of them as final submissions.</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>weight</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>LightGBM Part</td>\n<td>1.000 (LightGBM) : 0.000 (GNN)</td>\n<td>0.60025</td>\n<td>0.60008</td>\n</tr>\n<tr>\n<td>GNN part</td>\n<td>0.000 (LightGBM) : 1.000 (GNN)</td>\n<td>0.59894</td>\n<td>0.59874</td>\n</tr>\n<tr>\n<td>Final Submission 1</td>\n<td>0.600 (LightGBM) : 0.400 (GNN)</td>\n<td>0.60311</td>\n<td>0.60307</td>\n</tr>\n<tr>\n<td>Final Submission 2</td>\n<td>0.525 (LightGBM) : 0.475 (GNN)</td>\n<td>0.60302</td>\n<td>0.60313</td>\n</tr>\n</tbody>\n</table>\n<p>Both of final submissions improved LB score from the best of two parts. While Final Submission 1 was the best on public LB, Final Submission 2 was the best on private LB. Even though all single models got worse on private LB, the score of Final Submission 2 on private LB is better than public LB. Out team merge and ensemble was successful in this sense.</p>\n<h2>5. Conclusion</h2>\n<p>Our team employed LightGBM and GNN for this competition. Each approach took different way for candidate generation, feature engineering and model design. Ensemble of two types of approaches boosted our team score a lot.</p>\n<p>This competition gave 6th gold model to toshi_k and 7th gold to Jack. It motivate us for the further success. What we learn in this competition can be applied to not only future competitions but real world projects. It was confirmed that Kaggle is the practical platform of data science again.</p>\n<p>Thank you for reading this to the end!</p>",
      "rawMarkdown": "First of all, thank you for hosting this super exciting competition! And thank you to everyone for sharing many important insights in this competition. Discussions also have been very beneficial to us in our efforts to achieve this grade.\n\n## 1. Overview\n \nOur team consists of two members, Jack and toshi_k. Although both of us are competitions grandmaster, we have different strengths. Before making up a team, our approaches were totally different. Ensemble of two approaches cancelled out each weakness and boosted our team to the gold medal.\n\nOur solution is composed of LightGBM part, GNN (Graph Neural Network) part and ensemble. Jack was in charge of LightGBM part. He trained the best solo model in our team. toshi_k was in charge of GNN part and ensemble. He trained unique models by modern deep learning.  It contributed +0.003 to the team score by ensemble.\n\nThe details of LightGBM part is described in section 2. GNN part is described in section 3. Ensemble method and result are described in section 4.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F169364%2Fd46e627f90f7d1565ff24df3f9076e9f%2Foverall_03.jpg?generation=1675573670780068&alt=media)\n\n## 2. LightGBM part (by Jack)\n\n### 2.1 Candidate Generation\n\nBefore I go into details, let me just say that the method described here is my own, but when I optimized Recall@20, there was little difference of performance between it and the [Chris's Public Notebook](https://www.kaggle.com/code/cdeotte/compute-validation-score-cv-565) on validation. Thus, it is unclear how much of an advantage it may have.\n\nThe basic idea is to approximate a kind of posterior probability (let's call it a candidate score) of aids based on the co-visitation matrix, and the candidates were selected from the top ones. These candidates are common to all three action types.\n\nThe last 10 aids and action types of the session were used to calculate the candidate score. That is, the score can be calculated as follows:\n\n$$Score(aid) = P(aid | (aid_1, type_1), (aid_2, type_2), ..., (aid_{10}, type_{10}))$$\n\nHere, we forcibly assume independence and conditional independence of each action, similar in concept to Naive Bayes,\n\n$$\\begin{eqnarray}\nScore(aid)\n&=& P(aid) \\cdot \\frac{P((aid_1, type_1), ..., (aid_{10}, type_{10}) | aid)}{P((aid_1, type_1), ..., (aid_{10}, type_{10}))} \\\\\\\\\n&\\approx& P(aid) \\cdot \\prod_{i=1}^{10} \\frac{P((aid_i, type_i) | aid)}{P(aid_i, type_i)} \\\\\\\\\n&=& P(aid) \\cdot \\prod_{i=1}^{10} \\frac{P((aid_i, type_i), aid)}{P(aid_i, type_i) P(aid)}\n\\end{eqnarray}$$\n\nThe term inside the product is the ratio of the joint probability to the product of individual probability, that represents the extent to which P(aid) is enhanced by the observation of (aid_i, type_i). It is quite impossible to assume the above independence with this data, and therefore this value can exceed 1 and is no longer a probability. However, I expected it to work reasonably well in prioritizing candidates.\n\nEach term in the above equation is obtained by counting the frequency of each aid and co-visitation of aid pairs and dividing by the total number of sessions. In calculating the co-visitation matrix, only interactions within a 24-hour period are counted, and no multiple counts are made within the same session. The period of calculation was the entire period including test data (in inference phase), and co-visitation in both directions was to be counted.\n\nThe co-visitation matrix does not hold for all pairs of aids, but only those that have many co-visitation for each aid. For aid pairs that are not in the co-visitation matrix, the ratio of the joint probability on the right side of the above formula is set to 1, so that they do not affect the score calculation.\n\nIn the actual calculation, the logarithm is taken and further weighted to the most recent action, as follows:\n$$Score(aid) = \\log(P(aid)) + \\sum_{i=1}^{10} \\frac{11-i}{10} \\cdot \\log \\left( \\frac{P((aid_i, type_i), aid)}{P(aid_i, type_i) P(aid)} \\right)$$\n\nMany other heuristics, such as adding pseudo counts and adjusting by action type, have been incorporated, but they are too complicated to mention here.\n\nStarting with those with the highest candidate score, the top 200 were taken for training data, the top 300 for inference of test data, and then the already visited aids were added to make the final candidates.\n\nThe recall on validation of the top 200 candidates thus obtained was as follows:\n- clicks: 0.697\n- carts: 0.559\n- orders: 0.736\n\n### 2.2 LightGBM Rerank Model\n\nThis part is not much different from the methods already shared by others. The rerank model was trained by LightGBM (LamabdaRank), and separate models were built for each action type.\n\nOn validation, the second last week of train set (truncated) was used as training data and the last week of train set (truncated) as validation data.\nWhen inference was made on the test data, a model trained on the last week's data (LightGBM1) and a model trained on the second last week's data (LightGBM2) were built, and their outputs were ensembled by simple average. Since the training data were completely swapped, I expected a reasonable ensemble effect, but in fact it seems that the effect was only slight.\n\nMost of the features are based on co-visitation matrix, but each aggregation period is separate for training, validation, and test. That is, the co-visitation matrix is created for each of the three different periods, and the features are created, so they are leakage free.\n\nThe total number of features in the final model is 344, as follows:\n- session features (32)\n    - the number of all actions (1)\n    - the number of each action (3)\n    - the number of unique aids in the session (1)\n    - the number of unique aids of (carts/orders) and the ratio to the above (4)\n    - the last action type (1)\n    - the last relative timestamp from the start of the test period (1)\n    - the number of actions from the last of each action type to the last action of the session (3)\n    - elapsed time from i-th last action (i=2, ..., 10) to the last action of the session (9)\n    - revisit ratio of all aids by pair of action types (9)\n- aid features (50)\n    - count of (any/buy/click/cart/order) (5)\n    - exponential decay count of (any/buy/order) (3)\n    - count of (any/buy/order) in last n days (n=1~7) (21)\n    - count of (any/buy/order) in last n weeks (n=1~4) (12)\n    - revisit count in all sessions by pair of action types (9)\n- session*aid features (12)\n    - the latest action of that aid (1)\n    - the number of each action of that aid (3)\n    - the number of actions from last visit to that aid to the last action of the session (1)\n    - elapsed time from last visit to that aid to the last action of the session (1)\n    - the above two features for each action type (6)\n- co-visitation features (250)\n    - the number of co-visitation of aid with aid_i (i=1, ..., 10) devided by the global count of aid_i\n        - any to any (both direction/oneway) (20)\n        - any to buy (both direction/oneway) (20)\n        - buy to any (both direction/oneway) (20)\n        - buy to buy (both direction/oneway) (20)\n        - type_i to any (both direction/oneway) (20)\n        - type_i to buy (both direction/oneway) (20)\n        - click to click (both direction) (10)\n        - click to cart (both direction) (10)\n        - cart to click (both direction) (10)\n        - cart to cart (both direction) (10)\n    - the rank of co-visitation of aid with aid_i (i=1, ..., 10)\n        - any to any (both direction) (10)\n        - any to buy (both direction) (10)\n        - buy to any (both direction) (10)\n        - buy to buy (both direction) (10)\n    - global count of aid_i (any/buy/click/cart/type_i) (50)\n\n\\* \"any\" means the action clicks or carts or orders, and \"buy\" means the action carts or orders.\n\n## 3. GNN part (by toshi_k)\n\n### 3.1 Basic Idea\n\nI considered using DL (Deep Learning) in this competition. Since the datasets are relatively simple, E2E approach of DL seemed like a desirable solution for me. Another advantage is that multi-dimensional interactions and outputs for clicks/carts/orders are easily designed as a DL model architecture.\n\nDL based recommendation was initially proposed as a kind of non-linear collaborative filtering. The typical one is training AutoEncoder model and using the reconstruction methodology to evaluate missing ratings.\n\n- Training Deep AutoEncoders for Collaborative Filtering\n    - https://arxiv.org/abs/1708.01715\n\nAlthough I implemented this type of method as a prototype, it didn't work well. The number of items was so large that it made input vectors ultra sparse. It also yielded the heavy requirements of GPU memory for FC (Fully Connected) layers and made hidden layers shallower and thinner.\n\nThe disadvantage of FC layers is having weights between all combinations of aids even if most of them have nothing to do with each other. After some considerations, I figured out GNN (Graph Neural Networks) can solve this issue. The graph for GNN can represents aid relations and GNN can predict attributions of aids based on the nearly connected aids.\n\nUsing GNN for session based recommendation is also reported in the below study. According to the paper, their method is developed to explore rich transitions among items and generate accurate latent vectors of items. Their experiments on two datasets including thousands of items show that their method outperforms the state-of-the-art methods.\n\n- Session-based Recommendation with Graph Neural Networks\n    - https://arxiv.org/abs/1811.00855v4\n\nMy approach is similar to the previous study. One of the biggest differences of problem setting is the number of items. In this competition, the datasets contain millions of items. To handle all items, I built a simpler workflow and installed the subgraph extraction from the global session graph.\n\nBasically, my approach has 3 steps.\n1. Construct the global graph that represents aid relations\n2. Extract the subgraph from the global graph for each session\n3. Use GNN to predict which aid will be taken\n\nThe conceptual diagram is as below.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F169364%2F2327fdf46f8bcd96370e1f7e0ebaaf4d%2Fgnn_solution_01.jpg?generation=1675573719262138&alt=media)\n\nIn the first step, the global graph that represents aid relations is constructed. A subset of training data is used to create this graph. All transitions in this data are counted up and top P transitions (P=20, 30 or 40) from each aid are adopted as the edges of the graph. This graph is roughly corresponding to the co-visitation matrix that other participants call.\n\nSecondly, the subgraph is extracted from the global graph for each session. The history of each session is traced and nearly connected nodes (=aids) are listed up. Closer nodes and major transitions are prioritised and hundreds aids are filtered for extraction. This process is roughly corresponding to the candidates generation other participants call.\n\nThirdly, subgraphs are used to train and test GNN. The last layer of GNN has three channels. They predict if clicks/carts/orders will be taken in the future of each session. The inputs of GNN are the structures of subgraphs and features of nodes and edges. More details of features and GNN model are described in the next two subsections.\n\n### 3.2 Features\n\nThe input features for my GNN consist of \"node features\" and \"edge features\". Node features represent the characteristics of each aid. Edge features represent the relations between each pair of aids.\n\nBasically the total number of node features is 18. Nine of them is the global characteristics of aids. These features are shared among all sessions. The other nine features represent the history of sessions. These features are calculated on the session history and different for every sessions. The list of node features is as below.\n\n- Node features (18)\n    - Global aid features (9)\n        - Popularity counts (3)\n        - Repeat counts (3)\n        - Type transition counts (3)\n    - Session history features (9)\n        - Distance from the session history (2)\n        - Number of counts in the session history (3)\n        - Visited order features (2)\n        - Visited time features (2)\n\nTotal number of edge features is 14. Twelve of them is the global characteristics of transitions. These features are shared among all sessions. The other two features represent the history of sessions. These features are calculated on the session history and different for every sessions. The list of edge features is as below.\n\n- Edge features (14)\n    - Global transition features (12)\n        - Transition count not considering types (2)\n        - Transition rank not considering types (2)\n        - Cart-to-cart transition count (2)\n        - Cart-to-cart transition rank (2)\n        - Order-to-order transition count (2)\n        - Order-to-order transition rank (2)\n    - Session history features (2)\n        - Self loop or not (1)\n        - Stepped in session history or not (1)\n\nAny combinations of multiple features and higher dimension features are not added. It was expected that such complex features were automatically captured by the representation capability of GNN.\n\nAll missing values are filled with zero and logarithmic transformation (log1p) is applied to most features for the stability of GNN.\n\n### 3.3 Model and Loss function\n\nMy GNN has 8 GCN (Graph Convolution) layers. This implies GNN model can consider aids located within 8 steps from the history aids for prediction. In some trial experiments, 8 layers model was better than 4 layers one a little, but it was unclear more layers helped or not. \n\nAside from GCN layers, my model employs non linear activation functions, skip connections, and normalization layers. These component made training faster and yielded less training loss.\n\nAs mentioned above, the last layer has three channels for clicks/carts/orders. The channels for carts and orders are connected with sigmoid functions and trained by binary cross entropy loss. The channel for clicks is connected with softmax function and trained by softmax cross entropy loss.\n\n### 3.4 Performance boosting\n\nI noticed that increasing the number of candidates for inference boosted the LB score. For example, my best model improved from 0.59474 to 0.59750 on public LB by increasing the candidates from 384 to 1024.\n\n| Model | num of candidates | Public LB | Private LB |\n| --- | --- | --- | --- |\n| Model1 | 384 | 0.59474 | 0.59451 | \n| Model1 | 1024 | 0.59750 | 0.59719 |\n\nUnfortunately, more candidates required more computational calculation, 1024 candidates took several days for inference. I increased candidates as much as possible and save the outputs of several models.\n\nAs the final stage of the competition, ensemble of multiple predictions is attempted. Four GNN models are used for this ensemble. Although they were trained with slightly different setting, their basic approaches were the same. Ensemble of 4 models achieved 0.59894 on public LB and 0.59874 on private LB.\n\n| Model | num of candidates | Public LB | Private LB |\n| --- | --- | --- | --- |\n| Model1 | 1024 | 0.59750 | 0.59719 | \n| Model2 | 1024 | 0.59653 | 0.59613 |\n| Model3 | 512 | 0.59320 | 0.59318 |\n| Model4 | 768 | 0.59529 | 0.59478 |\n| Ensemble of 4 models | - | 0.59894 | 0.59874 |\n\n## 4. Team Ensemble (by toshi_k)\n\nAfter we made up a team, we tried several ways to merge our predictions. We started from the simple rerank by arithmetic mean of ranks in submission files. It improved the public score 0.597→0.600 then.\n\nThe next attempt was using raw prediction to boost LB score more. We saved the raw prediction of top 50 aids of each part. The biggest issue was the outputs from LambdaRank range in any real number while the output of GNN range in (0, 1).\n\nSince we didn't have enough time to lead the best theoretical way, we tried some ensemble method experimentally. One interesting finding was the logit transformed value of GNN seemed to have the proportional relationship with LambdaRank with constant shift.\n\n$$ \\mathrm{logit}(p^\\text{Binary Prediction}) \\propto v^\\text{LambdaRank prediction} + C $$\n\nThe value of constant shift is different on every session_types. This may be just a brute force approximation, we estimate C for each session_type and mapped the output of LambdaRank to 0-to-1 value.\n\n$$\\begin{eqnarray}\n\\hat{C} &=& \\arg \\min_C \\\\{ \\frac{1}{R} \\sum_r^R {^Gp_r} - \\frac{1}{R} \\sum_r^R \\sigma(^Lv_r + C) \\\\}^2 \\\\\\\\\n^Lp_r &=& \\sigma (v^L_r + \\hat{C}) \\\\\\\\\n\\mathrm{where:} \\\\\\\\\n^Gp_r &=& \\text{rth output of GNN} \\in (0, 1) \\\\\\\\\\\\\n^Lv_r &=& \\text{rth output of LambdaRank LGBM } \\in \\mathbb{R} \\\\\\\\\\\\\n^Lp_r &=& \\text{0-1 calibrated value of } ^Lv_r \\in (0, 1)\n\\end{eqnarray}$$\n\nAfter this transformation, simple weighted averaging was calculated. Since LightGBM part achieved better score on PublicLB, the weight of LightGBM is set to be larger than GNN part. We tried two patterns of weight settings and chose both of them as final submissions.\n\n|  | weight | Public LB | Private LB |\n| --- | --- | --- | --- |\n| LightGBM Part | 1.000 (LightGBM) : 0.000 (GNN) | 0.60025 | 0.60008 | \n| GNN part | 0.000 (LightGBM) : 1.000 (GNN) | 0.59894 | 0.59874 | \n| Final Submission 1 | 0.600 (LightGBM) : 0.400 (GNN) | 0.60311 | 0.60307 | \n| Final Submission 2 | 0.525 (LightGBM) : 0.475 (GNN) | 0.60302 | 0.60313 |\n\nBoth of final submissions improved LB score from the best of two parts. While Final Submission 1 was the best on public LB, Final Submission 2 was the best on private LB. Even though all single models got worse on private LB, the score of Final Submission 2 on private LB is better than public LB. Out team merge and ensemble was successful in this sense.\n\n## 5. Conclusion\n\nOur team employed LightGBM and GNN for this competition. Each approach took different way for candidate generation, feature engineering and model design. Ensemble of two types of approaches boosted our team score a lot.\n\nThis competition gave 6th gold model to toshi_k and 7th gold to Jack. It motivate us for the further success. What we learn in this competition can be applied to not only future competitions but real world projects. It was confirmed that Kaggle is the practical platform of data science again.\n\nThank you for reading this to the end!",
      "votes": null
    },
    {
      "id": "2130095",
      "postDate": "02/05/2023 06:16:01",
      "content": "<p>Fantastic result and great approach note <a href=\"https://www.kaggle.com/rsakata\" target=\"_blank\">@rsakata</a>! All the best and best regards!</p>",
      "rawMarkdown": "Fantastic result and great approach note @rsakata! All the best and best regards!",
      "votes": null
    },
    {
      "id": "2130310",
      "postDate": "02/05/2023 11:03:35",
      "content": "<p><a href=\"https://www.kaggle.com/toshik\" target=\"_blank\">@toshik</a> Congratulation for gold medals. I think gold medals of this tough competition has special values.</p>\n<p>I have a question, about ensemble of each single GNN models, what “calibrated average” actually calculates?<br>\nAlso, since I’m not familiar with this technique, I will appreciate if you share some references if any.</p>\n<p>Thank you for sharing this unique &amp; special approach. Although I don’t understand all the specifics, I guess utilizing GNN as such extent requires deep understanding of this literature.</p>",
      "rawMarkdown": "toshik Congratulation for gold medals. I think gold medals of this tough competition has special values.\n\nI have a question, about ensemble of each single GNN models, what “calibrated average” actually calculates?\nAlso, since I’m not familiar with this technique, I will appreciate if you share some references if any.\n\nThank you for sharing this unique & special approach. Although I don’t understand all the specifics, I guess utilizing GNN as such extent requires deep understanding of this literature.",
      "votes": null
    },
    {
      "id": "2130516",
      "postDate": "02/05/2023 14:00:24",
      "content": "<p><a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> Thank you for your comment. I'm not ready to share details of that part, but It was just a trivial technique. When some of my models output confident values (e.g. 0.4-1.0) and the others output less confident values (e.g. 0.0-0.6), they are calibrated to the same range (e.g 0.2-0.8) before averaging.</p>",
      "rawMarkdown": "tatamikenn Thank you for your comment. I'm not ready to share details of that part, but It was just a trivial technique. When some of my models output confident values (e.g. 0.4-1.0) and the others output less confident values (e.g. 0.0-0.6), they are calibrated to the same range (e.g 0.2-0.8) before averaging.",
      "votes": null
    },
    {
      "id": "2130675",
      "postDate": "02/05/2023 16:08:29",
      "content": "<p>Thank you for sharing very useful content.</p>",
      "rawMarkdown": "Thank you for sharing very useful content.",
      "votes": null
    },
    {
      "id": "2130798",
      "postDate": "02/05/2023 17:35:26",
      "content": "<p>Great ! What did you use to train the GNN ? I thought it requires a lot of resources and should be slower to train. I personally used an heterogenous graph  but I could use only kaggle kernels. I build a kind of GCN Autoencoder by reconstructing the adjacency matrix and then:</p>\n<ul>\n<li>use the encoding of both entities as features.</li>\n<li>apply a KNN Graph on the Users Embeddings to extract the nearest K  AIDs.</li>\n</ul>\n<p>Now I think that I should used it to extract other candidates ( another matrix ) to improve my classifications k by providing potential more true positives.</p>\n<p>Thanks for your great topic !</p>",
      "rawMarkdown": "Great ! What did you use to train the GNN ? I thought it requires a lot of resources and should be slower to train. I personally used an heterogenous graph  but I could use only kaggle kernels. I build a kind of GCN Autoencoder by reconstructing the adjacency matrix and then:\n- use the encoding of both entities as features.\n- apply a KNN Graph on the Users Embeddings to extract the nearest K  AIDs.\n\nNow I think that I should used it to extract other candidates ( another matrix ) to improve my classifications k by providing potential more true positives.\n\nThanks for your great topic !",
      "votes": null
    },
    {
      "id": "2130866",
      "postDate": "02/05/2023 18:31:38",
      "content": "<p>This is brilliant - congratulations both of you! 🎉 It's so interesting to see such an in-depth explanation of a GNN approach to this competition, and great to see the boost in performance you got from ensembling the two methods together :)</p>",
      "rawMarkdown": "This is brilliant - congratulations both of you! 🎉 It's so interesting to see such an in-depth explanation of a GNN approach to this competition, and great to see the boost in performance you got from ensembling the two methods together :)",
      "votes": null
    },
    {
      "id": "2132075",
      "postDate": "02/06/2023 15:35:33",
      "content": "<p>Thats amazing. Using GNN approach to solve this !!</p>",
      "rawMarkdown": "Thats amazing. Using GNN approach to solve this !!",
      "votes": null
    },
    {
      "id": "2132671",
      "postDate": "02/07/2023 01:26:00",
      "content": "<p><a href=\"https://www.kaggle.com/rayanaay\" target=\"_blank\">@rayanaay</a> My machine has Core-i9 10980XE and RTX A6000. It was actually resource consuming to train models. Although using GCN autoencoder may be a possible approach, extract candidates and ML based reranking would be important regardless of whether it is used.</p>",
      "rawMarkdown": "rayanaay My machine has Core-i9 10980XE and RTX A6000. It was actually resource consuming to train models. Although using GCN autoencoder may be a possible approach, extract candidates and ML based reranking would be important regardless of whether it is used.",
      "votes": null
    },
    {
      "id": "2134525",
      "postDate": "02/08/2023 03:56:22",
      "content": "<p>Amazing result!!<br>\n<a href=\"https://www.kaggle.com/toshik\" target=\"_blank\">@toshik</a> Could you share your gnn training/inference code?</p>",
      "rawMarkdown": "Amazing result!!\n@toshik Could you share your gnn training/inference code?",
      "votes": null
    },
    {
      "id": "2141829",
      "postDate": "02/13/2023 05:36:19",
      "content": "<p>Your solution is quite unique and well thought out, combining the strengths of both LightGBM and GNN. <a href=\"https://www.kaggle.com/rsakata\" target=\"_blank\">@rsakata</a> your candidate generation approach is quite impressive, utilizing the co-visitation matrix and a posterior probability approximation to prioritize the candidate aids. And <a href=\"https://www.kaggle.com/toshik\" target=\"_blank\">@toshik</a> your use of GNN to handle the large number of items in the datasets is very clever and effective, especially with the added step of subgraph extraction from the global session graph.</p>\n<p><a href=\"https://www.kaggle.com/toshik\" target=\"_blank\">@toshik</a> <a href=\"https://www.kaggle.com/rsakata\" target=\"_blank\">@rsakata</a> I have a few questions regarding your solution:</p>\n<ol>\n<li>Can you give more detail on the subgraph extraction step and how it helps in handling the large number of items?</li>\n<li>How did you determine the best hyperparameters for your GNN model?</li>\n</ol>",
      "rawMarkdown": "Your solution is quite unique and well thought out, combining the strengths of both LightGBM and GNN. @rsakata your candidate generation approach is quite impressive, utilizing the co-visitation matrix and a posterior probability approximation to prioritize the candidate aids. And @toshik your use of GNN to handle the large number of items in the datasets is very clever and effective, especially with the added step of subgraph extraction from the global session graph.\n\n@toshik @rsakata I have a few questions regarding your solution:\n\n1. Can you give more detail on the subgraph extraction step and how it helps in handling the large number of items?\n2. How did you determine the best hyperparameters for your GNN model?",
      "votes": null
    },
    {
      "id": "2143011",
      "postDate": "02/14/2023 00:55:12",
      "content": "<p><a href=\"https://www.kaggle.com/giranntu\" target=\"_blank\">@giranntu</a> The subgraph extraction step helps the model to focus on prospective items. While processing whole graph on GPU is impossible due to the memory limitation, subgraphs can be processed on GPU.<br>\nI tuned hyperparameters manually. Hyperparameters tuning for GNN is so severe that defining search configuration is still difficult.</p>",
      "rawMarkdown": "giranntu The subgraph extraction step helps the model to focus on prospective items. While processing whole graph on GPU is impossible due to the memory limitation, subgraphs can be processed on GPU.\nI tuned hyperparameters manually. Hyperparameters tuning for GNN is so severe that defining search configuration is still difficult.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2130095,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "02/05/2023 06:16:01",
      "content": "<p>Fantastic result and great approach note <a href=\"https://www.kaggle.com/rsakata\" target=\"_blank\">@rsakata</a>! All the best and best regards!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2130310,
      "author_name": "tatamikenn",
      "author_url": "",
      "post_date": "02/05/2023 11:03:35",
      "content": "<p><a href=\"https://www.kaggle.com/toshik\" target=\"_blank\">@toshik</a> Congratulation for gold medals. I think gold medals of this tough competition has special values.</p>\n<p>I have a question, about ensemble of each single GNN models, what “calibrated average” actually calculates?<br>\nAlso, since I’m not familiar with this technique, I will appreciate if you share some references if any.</p>\n<p>Thank you for sharing this unique &amp; special approach. Although I don’t understand all the specifics, I guess utilizing GNN as such extent requires deep understanding of this literature.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2130516,
          "author_name": "toshik",
          "author_url": "",
          "post_date": "02/05/2023 14:00:24",
          "content": "<p><a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> Thank you for your comment. I'm not ready to share details of that part, but It was just a trivial technique. When some of my models output confident values (e.g. 0.4-1.0) and the others output less confident values (e.g. 0.0-0.6), they are calibrated to the same range (e.g 0.2-0.8) before averaging.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2130675,
      "author_name": "ibrahimkaratas",
      "author_url": "",
      "post_date": "02/05/2023 16:08:29",
      "content": "<p>Thank you for sharing very useful content.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2130798,
      "author_name": "rayanaay",
      "author_url": "",
      "post_date": "02/05/2023 17:35:26",
      "content": "<p>Great ! What did you use to train the GNN ? I thought it requires a lot of resources and should be slower to train. I personally used an heterogenous graph  but I could use only kaggle kernels. I build a kind of GCN Autoencoder by reconstructing the adjacency matrix and then:</p>\n<ul>\n<li>use the encoding of both entities as features.</li>\n<li>apply a KNN Graph on the Users Embeddings to extract the nearest K  AIDs.</li>\n</ul>\n<p>Now I think that I should used it to extract other candidates ( another matrix ) to improve my classifications k by providing potential more true positives.</p>\n<p>Thanks for your great topic !</p>",
      "votes": null,
      "replies": [
        {
          "id": 2132671,
          "author_name": "toshik",
          "author_url": "",
          "post_date": "02/07/2023 01:26:00",
          "content": "<p><a href=\"https://www.kaggle.com/rayanaay\" target=\"_blank\">@rayanaay</a> My machine has Core-i9 10980XE and RTX A6000. It was actually resource consuming to train models. Although using GCN autoencoder may be a possible approach, extract candidates and ML based reranking would be important regardless of whether it is used.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2130866,
      "author_name": "judehunt23",
      "author_url": "",
      "post_date": "02/05/2023 18:31:38",
      "content": "<p>This is brilliant - congratulations both of you! 🎉 It's so interesting to see such an in-depth explanation of a GNN approach to this competition, and great to see the boost in performance you got from ensembling the two methods together :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2132075,
      "author_name": "sohilsharma1996",
      "author_url": "",
      "post_date": "02/06/2023 15:35:33",
      "content": "<p>Thats amazing. Using GNN approach to solve this !!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2134525,
      "author_name": "no9more9ria10",
      "author_url": "",
      "post_date": "02/08/2023 03:56:22",
      "content": "<p>Amazing result!!<br>\n<a href=\"https://www.kaggle.com/toshik\" target=\"_blank\">@toshik</a> Could you share your gnn training/inference code?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2141829,
      "author_name": "giranntu",
      "author_url": "",
      "post_date": "02/13/2023 05:36:19",
      "content": "<p>Your solution is quite unique and well thought out, combining the strengths of both LightGBM and GNN. <a href=\"https://www.kaggle.com/rsakata\" target=\"_blank\">@rsakata</a> your candidate generation approach is quite impressive, utilizing the co-visitation matrix and a posterior probability approximation to prioritize the candidate aids. And <a href=\"https://www.kaggle.com/toshik\" target=\"_blank\">@toshik</a> your use of GNN to handle the large number of items in the datasets is very clever and effective, especially with the added step of subgraph extraction from the global session graph.</p>\n<p><a href=\"https://www.kaggle.com/toshik\" target=\"_blank\">@toshik</a> <a href=\"https://www.kaggle.com/rsakata\" target=\"_blank\">@rsakata</a> I have a few questions regarding your solution:</p>\n<ol>\n<li>Can you give more detail on the subgraph extraction step and how it helps in handling the large number of items?</li>\n<li>How did you determine the best hyperparameters for your GNN model?</li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 2143011,
          "author_name": "toshik",
          "author_url": "",
          "post_date": "02/14/2023 00:55:12",
          "content": "<p><a href=\"https://www.kaggle.com/giranntu\" target=\"_blank\">@giranntu</a> The subgraph extraction step helps the model to focus on prospective items. While processing whole graph on GPU is impossible due to the memory limitation, subgraphs can be processed on GPU.<br>\nI tuned hyperparameters manually. Hyperparameters tuning for GNN is so severe that defining search configuration is still difficult.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2130054": "First of all, thank you for hosting this super exciting competition! And thank you to everyone for sharing many important insights in this competition. Discussions also have been very beneficial to us in our efforts to achieve this grade.\n\n## 1. Overview\n \nOur team consists of two members, Jack and toshi_k. Although both of us are competitions grandmaster, we have different strengths. Before making up a team, our approaches were totally different. Ensemble of two approaches cancelled out each weakness and boosted our team to the gold medal.\n\nOur solution is composed of LightGBM part, GNN (Graph Neural Network) part and ensemble. Jack was in charge of LightGBM part. He trained the best solo model in our team. toshi_k was in charge of GNN part and ensemble. He trained unique models by modern deep learning.  It contributed +0.003 to the team score by ensemble.\n\nThe details of LightGBM part is described in section 2. GNN part is described in section 3. Ensemble method and result are described in section 4.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F169364%2Fd46e627f90f7d1565ff24df3f9076e9f%2Foverall_03.jpg?generation=1675573670780068&alt=media)\n\n## 2. LightGBM part (by Jack)\n\n### 2.1 Candidate Generation\n\nBefore I go into details, let me just say that the method described here is my own, but when I optimized Recall@20, there was little difference of performance between it and the [Chris's Public Notebook](https://www.kaggle.com/code/cdeotte/compute-validation-score-cv-565) on validation. Thus, it is unclear how much of an advantage it may have.\n\nThe basic idea is to approximate a kind of posterior probability (let's call it a candidate score) of aids based on the co-visitation matrix, and the candidates were selected from the top ones. These candidates are common to all three action types.\n\nThe last 10 aids and action types of the session were used to calculate the candidate score. That is, the score can be calculated as follows:\n\n$$Score(aid) = P(aid | (aid_1, type_1), (aid_2, type_2), ..., (aid_{10}, type_{10}))$$\n\nHere, we forcibly assume independence and conditional independence of each action, similar in concept to Naive Bayes,\n\n$$\\begin{eqnarray}\nScore(aid)\n&=& P(aid) \\cdot \\frac{P((aid_1, type_1), ..., (aid_{10}, type_{10}) | aid)}{P((aid_1, type_1), ..., (aid_{10}, type_{10}))} \\\\\\\\\n&\\approx& P(aid) \\cdot \\prod_{i=1}^{10} \\frac{P((aid_i, type_i) | aid)}{P(aid_i, type_i)} \\\\\\\\\n&=& P(aid) \\cdot \\prod_{i=1}^{10} \\frac{P((aid_i, type_i), aid)}{P(aid_i, type_i) P(aid)}\n\\end{eqnarray}$$\n\nThe term inside the product is the ratio of the joint probability to the product of individual probability, that represents the extent to which P(aid) is enhanced by the observation of (aid_i, type_i). It is quite impossible to assume the above independence with this data, and therefore this value can exceed 1 and is no longer a probability. However, I expected it to work reasonably well in prioritizing candidates.\n\nEach term in the above equation is obtained by counting the frequency of each aid and co-visitation of aid pairs and dividing by the total number of sessions. In calculating the co-visitation matrix, only interactions within a 24-hour period are counted, and no multiple counts are made within the same session. The period of calculation was the entire period including test data (in inference phase), and co-visitation in both directions was to be counted.\n\nThe co-visitation matrix does not hold for all pairs of aids, but only those that have many co-visitation for each aid. For aid pairs that are not in the co-visitation matrix, the ratio of the joint probability on the right side of the above formula is set to 1, so that they do not affect the score calculation.\n\nIn the actual calculation, the logarithm is taken and further weighted to the most recent action, as follows:\n$$Score(aid) = \\log(P(aid)) + \\sum_{i=1}^{10} \\frac{11-i}{10} \\cdot \\log \\left( \\frac{P((aid_i, type_i), aid)}{P(aid_i, type_i) P(aid)} \\right)$$\n\nMany other heuristics, such as adding pseudo counts and adjusting by action type, have been incorporated, but they are too complicated to mention here.\n\nStarting with those with the highest candidate score, the top 200 were taken for training data, the top 300 for inference of test data, and then the already visited aids were added to make the final candidates.\n\nThe recall on validation of the top 200 candidates thus obtained was as follows:\n- clicks: 0.697\n- carts: 0.559\n- orders: 0.736\n\n### 2.2 LightGBM Rerank Model\n\nThis part is not much different from the methods already shared by others. The rerank model was trained by LightGBM (LamabdaRank), and separate models were built for each action type.\n\nOn validation, the second last week of train set (truncated) was used as training data and the last week of train set (truncated) as validation data.\nWhen inference was made on the test data, a model trained on the last week's data (LightGBM1) and a model trained on the second last week's data (LightGBM2) were built, and their outputs were ensembled by simple average. Since the training data were completely swapped, I expected a reasonable ensemble effect, but in fact it seems that the effect was only slight.\n\nMost of the features are based on co-visitation matrix, but each aggregation period is separate for training, validation, and test. That is, the co-visitation matrix is created for each of the three different periods, and the features are created, so they are leakage free.\n\nThe total number of features in the final model is 344, as follows:\n- session features (32)\n    - the number of all actions (1)\n    - the number of each action (3)\n    - the number of unique aids in the session (1)\n    - the number of unique aids of (carts/orders) and the ratio to the above (4)\n    - the last action type (1)\n    - the last relative timestamp from the start of the test period (1)\n    - the number of actions from the last of each action type to the last action of the session (3)\n    - elapsed time from i-th last action (i=2, ..., 10) to the last action of the session (9)\n    - revisit ratio of all aids by pair of action types (9)\n- aid features (50)\n    - count of (any/buy/click/cart/order) (5)\n    - exponential decay count of (any/buy/order) (3)\n    - count of (any/buy/order) in last n days (n=1~7) (21)\n    - count of (any/buy/order) in last n weeks (n=1~4) (12)\n    - revisit count in all sessions by pair of action types (9)\n- session*aid features (12)\n    - the latest action of that aid (1)\n    - the number of each action of that aid (3)\n    - the number of actions from last visit to that aid to the last action of the session (1)\n    - elapsed time from last visit to that aid to the last action of the session (1)\n    - the above two features for each action type (6)\n- co-visitation features (250)\n    - the number of co-visitation of aid with aid_i (i=1, ..., 10) devided by the global count of aid_i\n        - any to any (both direction/oneway) (20)\n        - any to buy (both direction/oneway) (20)\n        - buy to any (both direction/oneway) (20)\n        - buy to buy (both direction/oneway) (20)\n        - type_i to any (both direction/oneway) (20)\n        - type_i to buy (both direction/oneway) (20)\n        - click to click (both direction) (10)\n        - click to cart (both direction) (10)\n        - cart to click (both direction) (10)\n        - cart to cart (both direction) (10)\n    - the rank of co-visitation of aid with aid_i (i=1, ..., 10)\n        - any to any (both direction) (10)\n        - any to buy (both direction) (10)\n        - buy to any (both direction) (10)\n        - buy to buy (both direction) (10)\n    - global count of aid_i (any/buy/click/cart/type_i) (50)\n\n\\* \"any\" means the action clicks or carts or orders, and \"buy\" means the action carts or orders.\n\n## 3. GNN part (by toshi_k)\n\n### 3.1 Basic Idea\n\nI considered using DL (Deep Learning) in this competition. Since the datasets are relatively simple, E2E approach of DL seemed like a desirable solution for me. Another advantage is that multi-dimensional interactions and outputs for clicks/carts/orders are easily designed as a DL model architecture.\n\nDL based recommendation was initially proposed as a kind of non-linear collaborative filtering. The typical one is training AutoEncoder model and using the reconstruction methodology to evaluate missing ratings.\n\n- Training Deep AutoEncoders for Collaborative Filtering\n    - https://arxiv.org/abs/1708.01715\n\nAlthough I implemented this type of method as a prototype, it didn't work well. The number of items was so large that it made input vectors ultra sparse. It also yielded the heavy requirements of GPU memory for FC (Fully Connected) layers and made hidden layers shallower and thinner.\n\nThe disadvantage of FC layers is having weights between all combinations of aids even if most of them have nothing to do with each other. After some considerations, I figured out GNN (Graph Neural Networks) can solve this issue. The graph for GNN can represents aid relations and GNN can predict attributions of aids based on the nearly connected aids.\n\nUsing GNN for session based recommendation is also reported in the below study. According to the paper, their method is developed to explore rich transitions among items and generate accurate latent vectors of items. Their experiments on two datasets including thousands of items show that their method outperforms the state-of-the-art methods.\n\n- Session-based Recommendation with Graph Neural Networks\n    - https://arxiv.org/abs/1811.00855v4\n\nMy approach is similar to the previous study. One of the biggest differences of problem setting is the number of items. In this competition, the datasets contain millions of items. To handle all items, I built a simpler workflow and installed the subgraph extraction from the global session graph.\n\nBasically, my approach has 3 steps.\n1. Construct the global graph that represents aid relations\n2. Extract the subgraph from the global graph for each session\n3. Use GNN to predict which aid will be taken\n\nThe conceptual diagram is as below.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F169364%2F2327fdf46f8bcd96370e1f7e0ebaaf4d%2Fgnn_solution_01.jpg?generation=1675573719262138&alt=media)\n\nIn the first step, the global graph that represents aid relations is constructed. A subset of training data is used to create this graph. All transitions in this data are counted up and top P transitions (P=20, 30 or 40) from each aid are adopted as the edges of the graph. This graph is roughly corresponding to the co-visitation matrix that other participants call.\n\nSecondly, the subgraph is extracted from the global graph for each session. The history of each session is traced and nearly connected nodes (=aids) are listed up. Closer nodes and major transitions are prioritised and hundreds aids are filtered for extraction. This process is roughly corresponding to the candidates generation other participants call.\n\nThirdly, subgraphs are used to train and test GNN. The last layer of GNN has three channels. They predict if clicks/carts/orders will be taken in the future of each session. The inputs of GNN are the structures of subgraphs and features of nodes and edges. More details of features and GNN model are described in the next two subsections.\n\n### 3.2 Features\n\nThe input features for my GNN consist of \"node features\" and \"edge features\". Node features represent the characteristics of each aid. Edge features represent the relations between each pair of aids.\n\nBasically the total number of node features is 18. Nine of them is the global characteristics of aids. These features are shared among all sessions. The other nine features represent the history of sessions. These features are calculated on the session history and different for every sessions. The list of node features is as below.\n\n- Node features (18)\n    - Global aid features (9)\n        - Popularity counts (3)\n        - Repeat counts (3)\n        - Type transition counts (3)\n    - Session history features (9)\n        - Distance from the session history (2)\n        - Number of counts in the session history (3)\n        - Visited order features (2)\n        - Visited time features (2)\n\nTotal number of edge features is 14. Twelve of them is the global characteristics of transitions. These features are shared among all sessions. The other two features represent the history of sessions. These features are calculated on the session history and different for every sessions. The list of edge features is as below.\n\n- Edge features (14)\n    - Global transition features (12)\n        - Transition count not considering types (2)\n        - Transition rank not considering types (2)\n        - Cart-to-cart transition count (2)\n        - Cart-to-cart transition rank (2)\n        - Order-to-order transition count (2)\n        - Order-to-order transition rank (2)\n    - Session history features (2)\n        - Self loop or not (1)\n        - Stepped in session history or not (1)\n\nAny combinations of multiple features and higher dimension features are not added. It was expected that such complex features were automatically captured by the representation capability of GNN.\n\nAll missing values are filled with zero and logarithmic transformation (log1p) is applied to most features for the stability of GNN.\n\n### 3.3 Model and Loss function\n\nMy GNN has 8 GCN (Graph Convolution) layers. This implies GNN model can consider aids located within 8 steps from the history aids for prediction. In some trial experiments, 8 layers model was better than 4 layers one a little, but it was unclear more layers helped or not. \n\nAside from GCN layers, my model employs non linear activation functions, skip connections, and normalization layers. These component made training faster and yielded less training loss.\n\nAs mentioned above, the last layer has three channels for clicks/carts/orders. The channels for carts and orders are connected with sigmoid functions and trained by binary cross entropy loss. The channel for clicks is connected with softmax function and trained by softmax cross entropy loss.\n\n### 3.4 Performance boosting\n\nI noticed that increasing the number of candidates for inference boosted the LB score. For example, my best model improved from 0.59474 to 0.59750 on public LB by increasing the candidates from 384 to 1024.\n\n| Model | num of candidates | Public LB | Private LB |\n| --- | --- | --- | --- |\n| Model1 | 384 | 0.59474 | 0.59451 | \n| Model1 | 1024 | 0.59750 | 0.59719 |\n\nUnfortunately, more candidates required more computational calculation, 1024 candidates took several days for inference. I increased candidates as much as possible and save the outputs of several models.\n\nAs the final stage of the competition, ensemble of multiple predictions is attempted. Four GNN models are used for this ensemble. Although they were trained with slightly different setting, their basic approaches were the same. Ensemble of 4 models achieved 0.59894 on public LB and 0.59874 on private LB.\n\n| Model | num of candidates | Public LB | Private LB |\n| --- | --- | --- | --- |\n| Model1 | 1024 | 0.59750 | 0.59719 | \n| Model2 | 1024 | 0.59653 | 0.59613 |\n| Model3 | 512 | 0.59320 | 0.59318 |\n| Model4 | 768 | 0.59529 | 0.59478 |\n| Ensemble of 4 models | - | 0.59894 | 0.59874 |\n\n## 4. Team Ensemble (by toshi_k)\n\nAfter we made up a team, we tried several ways to merge our predictions. We started from the simple rerank by arithmetic mean of ranks in submission files. It improved the public score 0.597→0.600 then.\n\nThe next attempt was using raw prediction to boost LB score more. We saved the raw prediction of top 50 aids of each part. The biggest issue was the outputs from LambdaRank range in any real number while the output of GNN range in (0, 1).\n\nSince we didn't have enough time to lead the best theoretical way, we tried some ensemble method experimentally. One interesting finding was the logit transformed value of GNN seemed to have the proportional relationship with LambdaRank with constant shift.\n\n$$ \\mathrm{logit}(p^\\text{Binary Prediction}) \\propto v^\\text{LambdaRank prediction} + C $$\n\nThe value of constant shift is different on every session_types. This may be just a brute force approximation, we estimate C for each session_type and mapped the output of LambdaRank to 0-to-1 value.\n\n$$\\begin{eqnarray}\n\\hat{C} &=& \\arg \\min_C \\\\{ \\frac{1}{R} \\sum_r^R {^Gp_r} - \\frac{1}{R} \\sum_r^R \\sigma(^Lv_r + C) \\\\}^2 \\\\\\\\\n^Lp_r &=& \\sigma (v^L_r + \\hat{C}) \\\\\\\\\n\\mathrm{where:} \\\\\\\\\n^Gp_r &=& \\text{rth output of GNN} \\in (0, 1) \\\\\\\\\\\\\n^Lv_r &=& \\text{rth output of LambdaRank LGBM } \\in \\mathbb{R} \\\\\\\\\\\\\n^Lp_r &=& \\text{0-1 calibrated value of } ^Lv_r \\in (0, 1)\n\\end{eqnarray}$$\n\nAfter this transformation, simple weighted averaging was calculated. Since LightGBM part achieved better score on PublicLB, the weight of LightGBM is set to be larger than GNN part. We tried two patterns of weight settings and chose both of them as final submissions.\n\n|  | weight | Public LB | Private LB |\n| --- | --- | --- | --- |\n| LightGBM Part | 1.000 (LightGBM) : 0.000 (GNN) | 0.60025 | 0.60008 | \n| GNN part | 0.000 (LightGBM) : 1.000 (GNN) | 0.59894 | 0.59874 | \n| Final Submission 1 | 0.600 (LightGBM) : 0.400 (GNN) | 0.60311 | 0.60307 | \n| Final Submission 2 | 0.525 (LightGBM) : 0.475 (GNN) | 0.60302 | 0.60313 |\n\nBoth of final submissions improved LB score from the best of two parts. While Final Submission 1 was the best on public LB, Final Submission 2 was the best on private LB. Even though all single models got worse on private LB, the score of Final Submission 2 on private LB is better than public LB. Out team merge and ensemble was successful in this sense.\n\n## 5. Conclusion\n\nOur team employed LightGBM and GNN for this competition. Each approach took different way for candidate generation, feature engineering and model design. Ensemble of two types of approaches boosted our team score a lot.\n\nThis competition gave 6th gold model to toshi_k and 7th gold to Jack. It motivate us for the further success. What we learn in this competition can be applied to not only future competitions but real world projects. It was confirmed that Kaggle is the practical platform of data science again.\n\nThank you for reading this to the end!",
    "2130095": "Fantastic result and great approach note @rsakata! All the best and best regards!",
    "2130310": "toshik Congratulation for gold medals. I think gold medals of this tough competition has special values.\n\nI have a question, about ensemble of each single GNN models, what “calibrated average” actually calculates?\nAlso, since I’m not familiar with this technique, I will appreciate if you share some references if any.\n\nThank you for sharing this unique & special approach. Although I don’t understand all the specifics, I guess utilizing GNN as such extent requires deep understanding of this literature.",
    "2130516": "tatamikenn Thank you for your comment. I'm not ready to share details of that part, but It was just a trivial technique. When some of my models output confident values (e.g. 0.4-1.0) and the others output less confident values (e.g. 0.0-0.6), they are calibrated to the same range (e.g 0.2-0.8) before averaging.",
    "2130675": "Thank you for sharing very useful content.",
    "2130798": "Great ! What did you use to train the GNN ? I thought it requires a lot of resources and should be slower to train. I personally used an heterogenous graph  but I could use only kaggle kernels. I build a kind of GCN Autoencoder by reconstructing the adjacency matrix and then:\n- use the encoding of both entities as features.\n- apply a KNN Graph on the Users Embeddings to extract the nearest K  AIDs.\n\nNow I think that I should used it to extract other candidates ( another matrix ) to improve my classifications k by providing potential more true positives.\n\nThanks for your great topic !",
    "2130866": "This is brilliant - congratulations both of you! 🎉 It's so interesting to see such an in-depth explanation of a GNN approach to this competition, and great to see the boost in performance you got from ensembling the two methods together :)",
    "2132075": "Thats amazing. Using GNN approach to solve this !!",
    "2132671": "rayanaay My machine has Core-i9 10980XE and RTX A6000. It was actually resource consuming to train models. Although using GCN autoencoder may be a possible approach, extract candidates and ML based reranking would be important regardless of whether it is used.",
    "2134525": "Amazing result!!\n@toshik Could you share your gnn training/inference code?",
    "2141829": "Your solution is quite unique and well thought out, combining the strengths of both LightGBM and GNN. @rsakata your candidate generation approach is quite impressive, utilizing the co-visitation matrix and a posterior probability approximation to prioritize the candidate aids. And @toshik your use of GNN to handle the large number of items in the datasets is very clever and effective, especially with the added step of subgraph extraction from the global session graph.\n\n@toshik @rsakata I have a few questions regarding your solution:\n\n1. Can you give more detail on the subgraph extraction step and how it helps in handling the large number of items?\n2. How did you determine the best hyperparameters for your GNN model?",
    "2143011": "giranntu The subgraph extraction step helps the model to focus on prospective items. While processing whole graph on GPU is impossible due to the memory limitation, subgraphs can be processed on GPU.\nI tuned hyperparameters manually. Hyperparameters tuning for GNN is so severe that defining search configuration is still difficult."
  },
  "source": "meta"
}