{
  "id": 376573,
  "title": "Candidates evaluation tricks from previous winning solutions",
  "url": "/competitions/otto-recommender-system/discussion/376573",
  "author_name": "The Devastator",
  "post_date": "2023-01-07T07:07:36.514000",
  "votes": 13,
  "comment_count": 0,
  "views": 0,
  "content": "<h1>Candidates evaluation tricks from previous winning solutions</h1>\n<p>Since this competition many approaches involve generating candidates and ranking them, It can be useful to take a step back and look at some solutions from past competitions that are based on the same general Idea:</p>\n<blockquote>\n  <p><strong>Candidates Generation</strong> -&gt; <strong>Rank / Match</strong> -&gt; <strong>Final Candidates</strong></p>\n</blockquote>\n<p>I am crossposting a previous short summary I made a couple of month ago about this topic since it can be interesting in this context.</p>\n<p>Enjoy! </p>\n<hr>\n<p>A while ago there had been a competition where the participants needed to match up images/text of the same product. <br>\nOn that competition, there were many ensemble / matching tricks deplyed by the winners, mostly done on embeddings extracted either from the images or from the text description.</p>\n<p>This topic summarize all the matching tricks from the solution writeups.</p>\n<h4><a href=\"https://www.kaggle.com/competitions/shopee-product-matching/discussion/238136\" target=\"_blank\">Shopee's 1st Place Solution</a></h4>\n<p><strong>Iterative Neighborhood Blending (INB)</strong></p>\n<p>Apart from combining image and text matches, it was also crucial to properly utilize embeddings to produce matches. The author made a nontrivial pipeline for searching matches from embeddings.<br>\nBased on QE(Query Expansion) and DBA(DataBase-side feature Augmentation), they created a pipeline called INB(Iterative Neighborhood Blending). <br>\nMost of the ideas are shared with QE and DBA, but some details are different. INB pipeline consists of these components.</p>\n<p><strong>K Nearest Neighbor Search</strong><br>\nThey use <a href=\"https://github.com/facebookresearch/faiss\" target=\"_blank\">faissed</a> for knn search, and set k=51 (maximum 50 non-self matches + 1 self). They used inner product as similarity metric (the embedding is normalized so it is equivalent to cosine similarity)</p>\n<p><strong>To decide about the threshold</strong></p>\n<p>They converted cosine similarity to cosine distance (= 1-cosine similarity) for some convenience in implementation, and obtained (matches, distances) pair that satisfies distance &lt; threshold. For each item x, we call this (matches, distances) pair as \"neighborhood of x\".</p>\n<p><strong>Neighborhood Blending</strong></p>\n<p>The intuition for neighborhood blending is straightforward. After knn and thresholding, they obtain the (matches, similarities) pair for each item, we have a graph, where each node is an item, and the edge weight is the similarity between two nodes. Only the neighborhoods are connected. That is, nodes that didn't pass the threshold condition and min2 condition from the query node, are disconnected.</p>\n<p>We want to use the neighborhood items' information to refine the query item's embedding and make the cluster clearer. In order to do that, they simply weighted-sum the neighborhood embeddings with similarity as weights and add it to the query embedding. So they blend neighborhood embeddings. They call it NB(Neighborhood Blending).</p>\n<p><img src=\"https://i.ibb.co/F4BCKjX/3.jpg\" alt=\"\"></p>\n<p>The image illustrates how one step of neighborhood blending is performed on a toy example. Let's look at node A. Its embedding is [-0.588, 0.784, 0.196] and its similarity to node B, C, D is 0.94, 0.93, 0.52 respectively. Red line means two nodes are neighbors, so they are connected. Dashed line means two nodes didn't pass the threshold, so are disconnected. </p>\n<p>They can apply NB iteratively. After blending neighborhood for stage1 embeddings, we do knn search &amp; 'thresholding with min2' to get stage2 (matches, similarities). <br>\nThey apply NB again, to further refine the embeddings. We can iterate until the evaluation metric stops improving. This is where Iterative comes from.</p>\n<p>Code:</p>\n<pre><code>def blend_neighborhood(emb, match_index_lst, similarities_lst):\n    new_emb = emb.copy()\n    for i in range(emb.shape[0]):\n        cur_emb = emb[match_index_lst[i]]\n        weights = np.expand_dims(similarities_lst[i], 1)\n        new_emb[i] = (cur_emb * weights).sum(axis=0)\n    new_emb = normalize(new_emb, axis=1)\n    return new_emb\n\ndef iterative_neighborhood_blending(emb, threshes):\n    for thresh in threshes:\n        match_index_lst, similarities_lst = neighborhood_search(emb, thresh)\n        emb = blend_neighborhood(emb, match_index_lst, similarities_lst)\n    return match_index_lst\n</code></pre>\n<h4><a href=\"https://www.kaggle.com/competitions/shopee-product-matching/discussion/238022\" target=\"_blank\">2nd place solution (matching prediction by GAT &amp; LGB)</a></h4>\n<p>The second place of this competition used some extra features for their 2nd level models such as:</p>\n<ul>\n<li>Image similarity</li>\n<li>Cosine similarities of NFNet-F0, ViT embeddings</li>\n<li>Loss: CurricularFace (better than ArcFace and others)</li>\n<li>Optimizer: SAM (better than Adam, SGD, and others)</li>\n<li>Concatenate the similarities like F.normalize(torch.cat([F.normalize(emb1), F.normalize(emb2)], axis=1))</li>\n<li>Text similarity</li>\n<li>Cosine similarities of Indonesian-BERT, Multilingual-BERT, and Paraphrase-XLM embeddings</li>\n<li>TF-IDF as in many public kernels</li>\n<li>Multimodal (image + text) similarity</li>\n<li>Trained model with NFNet-F0 and Indonesian BERT (concatenated at final feature layers)</li>\n<li>Graph features</li>\n<li>Avg and std of top-K cosine similarities of each item</li>\n<li>K=5, 10, 15, 30, etc</li>\n<li>Pagerank</li>\n<li>Text length</li>\n<li># of word</li>\n<li>Levenshtein distance</li>\n<li>Image file size</li>\n<li>Width and height of image</li>\n<li>Query Expansion</li>\n<li>Obtain an augmented embedding which weighted average neighbors.</li>\n<li>Concatenate the original and augmented embeddings like F.normalize(torch.cat([F.normalize(orig_emb), F.normalize(qe_emb)], axis=1))</li>\n</ul>\n<p><strong>Graph Attention Networks</strong></p>\n<p>They then constructed a model that based on Graph Attention Networks (GAT)</p>\n<p>They chose GAT due to it's simplicity and customizability</p>\n<p>At the end, they used only 4 features: image/bert/multi-modal/tf-idf similarities<br>\nApply graph attention to \"other edges connected to the node connected to the target edge\" as a neighborhood</p>\n<p><img src=\"https://user-images.githubusercontent.com/27487010/123567967-47ba7880-d7fe-11eb-9624-19fd8fd5d037.png\" alt=\"\"></p>\n<p><strong>Graph-Based Post-processing</strong></p>\n<p>They went further on and recursively removed edges that had the highest betweenness centrality</p>\n<ul>\n<li>The Intuition behind this is that Label groups should form a clique If there is an abundant edge, it should bridge two clique such an edge should have a higher betweenness centrality.</li>\n<li>Recursively removing such an edge until there is no such an edge or connected components become smaller than some threshold pushed their score by a large margin.</li>\n</ul>\n<h4><a href=\"https://www.kaggle.com/competitions/shopee-product-matching/discussion/237972\" target=\"_blank\">18th place solution</a></h4>\n<p><strong>Combining model outputs</strong></p>\n<p>They predicted with all their 1st stage models and then they replaced the predictions of each row with an average of predictions of its nearest neighbors. <br>\nThey set threshold for “nearest” so that 4 neighbors are used on average (this was found experimentally).</p>\n<p><strong>Forcing groups into the desired distribution</strong></p>\n<p>This was the largest single trick that greatly improved their score. <br>\nThey first make an educated guess that the distribution of group targets in the test data is similar to the one in train. Then they try to make their predictions have the same shape.<br>\nThey first decide that groups with 2 elements are going to be those where the third largest element is the lowest. They follow the same logic for all sizes up to 50 and submit the best.</p>\n<p><strong>Cross-mean embedding</strong></p>\n<p>They then applied some simple huristics to boost the score a bit further: </p>\n<ul>\n<li>If there are products(e.g. A,B and C) which have identical image, they are same product. (precision &gt; 99.9%)</li>\n<li>They replace each <em>title</em> embedding to mean of title embedding: (A+B+C)/3</li>\n<li>Same procedure for image embedding, mean by same title. [the competition was about images and texts matching of the same product. </li>\n</ul>\n<p><strong>DBA/QE</strong><br>\nFor image X, I got 3 nearest neighbor(e.g. X,Y,Z) by image embedding, then replace X's image embedding with weighted(logspace) sum of (X,Y,Z). same goes for title.</p>\n<p><strong>Graph merge</strong><br>\nNow we have two embeddings: image and title. They Merge them as graph.</p>\n<p>For each image or title embedding, they got 50 nearest neighbors and euclidean distance.<br>\nOne node represents one embedding. Edge weight is their euclidean distance.<br>\nThen We can merge image and title graph. If there is duplicated edge, pick small(near) one.</p>\n<p>They then employed some kind of <a href=\"https://en.wikipedia.org/wiki/Dijkstra%27s_algorithm\" target=\"_blank\">Dijkstra's algorithm</a> to form clusters.<br>\nAnd chose final threshold from LB score.</p>\n<p><img src=\"https://i.imgur.com/4OLmNIc.png\" alt=\"\"></p>\n<p><img src=\"https://i.imgur.com/Fp2gsrD.png\" alt=\"\"></p>\n<h4><a href=\"https://www.kaggle.com/competitions/shopee-product-matching/discussion/238078\" target=\"_blank\">5th Place Solution</a></h4>\n<p><strong>Matching</strong></p>\n<p><strong>2nd Stage Model</strong><br>\nThey have extracted some features for matched unique pairs and then fed them to XGB 2nd stage model which runs on GPU. The features where \"img_dist\", \"text_dist\", \"dist\", \"dist_rank\", \"cos_sim\", \"cos_sim2\". Basically, vector distances, their ranking within each posting id, 2 different tfidf cosine similarity with different parameters.<br>\nSince test set size was larger than their one fold size, some features were going to have different distribution on the test set due to higher possibility of False Positive matches. <br>\nTherefore they used another XGB model that was trained on percentage rank features and ensembled with the previous.</p>\n<p><strong>FP Features</strong><br>\nSince test set and train set have no maching between them, one can use training set for determining how easy it is for a posting to have FPs. They used closest match distances from training set as features. This method improved our CV score significantly but improved LB relatively less for them. They claim this is again due to the size difference between the train and test.</p>\n<p><strong>Agglomerative Clustering</strong></p>\n<p>Once they have match probability predictions from XGB models, they could then use a threshold and match the postings. But they did something a bit more complex. They sorted all pairs by their match probabilities and started matching from the most likely match. Each match above 0.8 probability, merges their clusters. Each posting with no cluster can be matched to a cluster if the probability is above 0.7. Each posting with no cluster can be matched with each other if the probability is above 0.3. This method allows to have confident clusters and not-so-confident many pairs.</p>\n<h4><a href=\"https://www.kaggle.com/competitions/shopee-product-matching/discussion/238126\" target=\"_blank\">26th Place Solution - Effective Cluster Separation and Neighbour Search</a></h4>\n<p><strong>Balancing recall by the neighbours</strong></p>\n<p>They realized that we they not doing good on lb because simple concatenation of text and image predictions increases the number of predicted values and hence decreased the recall. Thus they came up with this idea to balance the prediction and recall.</p>\n<p>They take the embeddings from each image model and pass it through the pre-processing layer, in the pre-processing layer they take out top 3 neighbours for each row using knns and multiply it with decreasing weights on log space so as to shift these top 3 embeddings into another cluster on the embedding space, while the other embeddings remain the same, (they tried many numbers instead of 3 but taking 3 best neighbours seem to work the best).</p>\n<p>They do the same with text embeddings and at the end they have 4 modified embeddings from four models.</p>\n<p><strong>Neighbourhood Search</strong></p>\n<p>They had five different models and five different embeddings, they tried a lot of different ensemble strategies that didn't help, they came up with this idea:</p>\n<p>They take each model and find 100 neighbours for each posting id and for each model, they make separate dataframes for each. For eg:- they take the nfnet model get 100 neighbours for each posting_id and form a dataframe, so now their dataframe has posting_id , indexes of neighbours and their distances.</p>\n<p>Now they take all the 5 dataframes, they are not the same as different models have different neigbours for different posting_ids, thus they do outer join, so instead of concatenating the embeddings we concat the distances. Now they have our final dataframe, in this they search for the final predictions using some threshold.</p>\n<h4><a href=\"https://www.kaggle.com/competitions/shopee-product-matching/discussion/238149\" target=\"_blank\">62th Place Solution (Stacking Logistic Regression)</a></h4>\n<p><strong>Ensembling</strong></p>\n<p>Concat all image models' outputs to get final image embedding and calculate cosine similarity (X) of all pairs (total N x N pairs). And do the same thing to get text cosine similarity (Y).<br>\nX:  image cosine similarity<br>\nY:  text cosine similarity</p>\n<p><strong>Downsampling</strong></p>\n<p>For each product, they used all products with the same label_group and sampled an equal amount of negative data from top products sorted by X+Y.<br>\nThey they use polynomial transformation (1,X,X^2,Y,Y^2,XY)<br>\nAnd then use Logistic regression to get final decision boundary.</p>\n<h4><a href=\"https://www.kaggle.com/competitions/shopee-product-matching/discussion/238295\" target=\"_blank\">4th Place Solution</a></h4>\n<p><strong>Ensemble</strong><br>\nOn their ensemble, they concatenated and normalized image embeddings, text embeddings (bert based) and tfdif embeddings. On each of those vectors, they calculated pair wise cosine similarity and received three matrices (cossim image, cossim bert, cossim tfidf). They combined those three matrices by first squareing them and then taking a weighted average.</p>\n<p><strong>Post Processing</strong><br>\nWhile the averaged cossin matrix already gave them a very competitive score when using a proper threshold, they still applied several post processing steps to squeeze out a few more points. </p>\n<p>The biggest contributors were:</p>\n<ul>\n<li>Thresholding</li>\n<li>Rank2 matching (if A has B on rank2 and B has A on rank2, add them to each other)</li>\n<li>Rank2 and rank3 difference is large -&gt; add rank2 id</li>\n<li>If there is a group of X members, and we have another row with the X preds on rank 1-X, make one large group.</li>\n<li>At least one other match (except the cosine similarity of rank2 is extremly low)</li>\n<li>Query Expansion</li>\n<li>Rematching unmatched rows</li>\n</ul>\n<h4><a href=\"https://www.kaggle.com/competitions/shopee-product-matching/discussion/238515\" target=\"_blank\">3rd Place Solution (Triplet loss, Boosting, Clustering)</a></h4>\n<p><strong>Postprocessing (clustering)</strong></p>\n<p>They perform clustering, not over embeddings but over pairwise distances. Distance (better say - similarity) for them was GBM probability of duplicate estimation. If there was no forecast for product pair, assume probability of duplicate eq 0. They took the idea of agglomerative clustering and slightly modify it for current task. They start from each single point as a cluster and merge them until average cluster size become equal threshold. If after clustering point stay single - merge it to nearest cluster. This algorithm boost them up the leaderboard after a lot of parameter search</p>",
  "messages": [
    {
      "id": 2090262,
      "postDate": "2023-01-07T07:07:36.513Z",
      "content": "<h1>Candidates evaluation tricks from previous winning solutions</h1>\n<p>Since this competition many approaches involve generating candidates and ranking them, It can be useful to take a step back and look at some solutions from past competitions that are based on the same general Idea:</p>\n<blockquote>\n  <p><strong>Candidates Generation</strong> -&gt; <strong>Rank / Match</strong> -&gt; <strong>Final Candidates</strong></p>\n</blockquote>\n<p>I am crossposting a previous short summary I made a couple of month ago about this topic since it can be interesting in this context.</p>\n<p>Enjoy! </p>\n<hr>\n<p>A while ago there had been a competition where the participants needed to match up images/text of the same product. <br>\nOn that competition, there were many ensemble / matching tricks deplyed by the winners, mostly done on embeddings extracted either from the images or from the text description.</p>\n<p>This topic summarize all the matching tricks from the solution writeups.</p>\n<h4><a href=\"https://www.kaggle.com/competitions/shopee-product-matching/discussion/238136\" target=\"_blank\">Shopee's 1st Place Solution</a></h4>\n<p><strong>Iterative Neighborhood Blending (INB)</strong></p>\n<p>Apart from combining image and text matches, it was also crucial to properly utilize embeddings to produce matches. The author made a nontrivial pipeline for searching matches from embeddings.<br>\nBased on QE(Query Expansion) and DBA(DataBase-side feature Augmentation), they created a pipeline called INB(Iterative Neighborhood Blending). <br>\nMost of the ideas are shared with QE and DBA, but some details are different. INB pipeline consists of these components.</p>\n<p><strong>K Nearest Neighbor Search</strong><br>\nThey use <a href=\"https://github.com/facebookresearch/faiss\" target=\"_blank\">faissed</a> for knn search, and set k=51 (maximum 50 non-self matches + 1 self). They used inner product as similarity metric (the embedding is normalized so it is equivalent to cosine similarity)</p>\n<p><strong>To decide about the threshold</strong></p>\n<p>They converted cosine similarity to cosine distance (= 1-cosine similarity) for some convenience in implementation, and obtained (matches, distances) pair that satisfies distance &lt; threshold. For each item x, we call this (matches, distances) pair as \"neighborhood of x\".</p>\n<p><strong>Neighborhood Blending</strong></p>\n<p>The intuition for neighborhood blending is straightforward. After knn and thresholding, they obtain the (matches, similarities) pair for each item, we have a graph, where each node is an item, and the edge weight is the similarity between two nodes. Only the neighborhoods are connected. That is, nodes that didn't pass the threshold condition and min2 condition from the query node, are disconnected.</p>\n<p>We want to use the neighborhood items' information to refine the query item's embedding and make the cluster clearer. In order to do that, they simply weighted-sum the neighborhood embeddings with similarity as weights and add it to the query embedding. So they blend neighborhood embeddings. They call it NB(Neighborhood Blending).</p>\n<p><img src=\"https://i.ibb.co/F4BCKjX/3.jpg\" alt=\"\"></p>\n<p>The image illustrates how one step of neighborhood blending is performed on a toy example. Let's look at node A. Its embedding is [-0.588, 0.784, 0.196] and its similarity to node B, C, D is 0.94, 0.93, 0.52 respectively. Red line means two nodes are neighbors, so they are connected. Dashed line means two nodes didn't pass the threshold, so are disconnected. </p>\n<p>They can apply NB iteratively. After blending neighborhood for stage1 embeddings, we do knn search &amp; 'thresholding with min2' to get stage2 (matches, similarities). <br>\nThey apply NB again, to further refine the embeddings. We can iterate until the evaluation metric stops improving. This is where Iterative comes from.</p>\n<p>Code:</p>\n<pre><code>def blend_neighborhood(emb, match_index_lst, similarities_lst):\n    new_emb = emb.copy()\n    for i in range(emb.shape[0]):\n        cur_emb = emb[match_index_lst[i]]\n        weights = np.expand_dims(similarities_lst[i], 1)\n        new_emb[i] = (cur_emb * weights).sum(axis=0)\n    new_emb = normalize(new_emb, axis=1)\n    return new_emb\n\ndef iterative_neighborhood_blending(emb, threshes):\n    for thresh in threshes:\n        match_index_lst, similarities_lst = neighborhood_search(emb, thresh)\n        emb = blend_neighborhood(emb, match_index_lst, similarities_lst)\n    return match_index_lst\n</code></pre>\n<h4><a href=\"https://www.kaggle.com/competitions/shopee-product-matching/discussion/238022\" target=\"_blank\">2nd place solution (matching prediction by GAT &amp; LGB)</a></h4>\n<p>The second place of this competition used some extra features for their 2nd level models such as:</p>\n<ul>\n<li>Image similarity</li>\n<li>Cosine similarities of NFNet-F0, ViT embeddings</li>\n<li>Loss: CurricularFace (better than ArcFace and others)</li>\n<li>Optimizer: SAM (better than Adam, SGD, and others)</li>\n<li>Concatenate the similarities like F.normalize(torch.cat([F.normalize(emb1), F.normalize(emb2)], axis=1))</li>\n<li>Text similarity</li>\n<li>Cosine similarities of Indonesian-BERT, Multilingual-BERT, and Paraphrase-XLM embeddings</li>\n<li>TF-IDF as in many public kernels</li>\n<li>Multimodal (image + text) similarity</li>\n<li>Trained model with NFNet-F0 and Indonesian BERT (concatenated at final feature layers)</li>\n<li>Graph features</li>\n<li>Avg and std of top-K cosine similarities of each item</li>\n<li>K=5, 10, 15, 30, etc</li>\n<li>Pagerank</li>\n<li>Text length</li>\n<li># of word</li>\n<li>Levenshtein distance</li>\n<li>Image file size</li>\n<li>Width and height of image</li>\n<li>Query Expansion</li>\n<li>Obtain an augmented embedding which weighted average neighbors.</li>\n<li>Concatenate the original and augmented embeddings like F.normalize(torch.cat([F.normalize(orig_emb), F.normalize(qe_emb)], axis=1))</li>\n</ul>\n<p><strong>Graph Attention Networks</strong></p>\n<p>They then constructed a model that based on Graph Attention Networks (GAT)</p>\n<p>They chose GAT due to it's simplicity and customizability</p>\n<p>At the end, they used only 4 features: image/bert/multi-modal/tf-idf similarities<br>\nApply graph attention to \"other edges connected to the node connected to the target edge\" as a neighborhood</p>\n<p><img src=\"https://user-images.githubusercontent.com/27487010/123567967-47ba7880-d7fe-11eb-9624-19fd8fd5d037.png\" alt=\"\"></p>\n<p><strong>Graph-Based Post-processing</strong></p>\n<p>They went further on and recursively removed edges that had the highest betweenness centrality</p>\n<ul>\n<li>The Intuition behind this is that Label groups should form a clique If there is an abundant edge, it should bridge two clique such an edge should have a higher betweenness centrality.</li>\n<li>Recursively removing such an edge until there is no such an edge or connected components become smaller than some threshold pushed their score by a large margin.</li>\n</ul>\n<h4><a href=\"https://www.kaggle.com/competitions/shopee-product-matching/discussion/237972\" target=\"_blank\">18th place solution</a></h4>\n<p><strong>Combining model outputs</strong></p>\n<p>They predicted with all their 1st stage models and then they replaced the predictions of each row with an average of predictions of its nearest neighbors. <br>\nThey set threshold for “nearest” so that 4 neighbors are used on average (this was found experimentally).</p>\n<p><strong>Forcing groups into the desired distribution</strong></p>\n<p>This was the largest single trick that greatly improved their score. <br>\nThey first make an educated guess that the distribution of group targets in the test data is similar to the one in train. Then they try to make their predictions have the same shape.<br>\nThey first decide that groups with 2 elements are going to be those where the third largest element is the lowest. They follow the same logic for all sizes up to 50 and submit the best.</p>\n<p><strong>Cross-mean embedding</strong></p>\n<p>They then applied some simple huristics to boost the score a bit further: </p>\n<ul>\n<li>If there are products(e.g. A,B and C) which have identical image, they are same product. (precision &gt; 99.9%)</li>\n<li>They replace each <em>title</em> embedding to mean of title embedding: (A+B+C)/3</li>\n<li>Same procedure for image embedding, mean by same title. [the competition was about images and texts matching of the same product. </li>\n</ul>\n<p><strong>DBA/QE</strong><br>\nFor image X, I got 3 nearest neighbor(e.g. X,Y,Z) by image embedding, then replace X's image embedding with weighted(logspace) sum of (X,Y,Z). same goes for title.</p>\n<p><strong>Graph merge</strong><br>\nNow we have two embeddings: image and title. They Merge them as graph.</p>\n<p>For each image or title embedding, they got 50 nearest neighbors and euclidean distance.<br>\nOne node represents one embedding. Edge weight is their euclidean distance.<br>\nThen We can merge image and title graph. If there is duplicated edge, pick small(near) one.</p>\n<p>They then employed some kind of <a href=\"https://en.wikipedia.org/wiki/Dijkstra%27s_algorithm\" target=\"_blank\">Dijkstra's algorithm</a> to form clusters.<br>\nAnd chose final threshold from LB score.</p>\n<p><img src=\"https://i.imgur.com/4OLmNIc.png\" alt=\"\"></p>\n<p><img src=\"https://i.imgur.com/Fp2gsrD.png\" alt=\"\"></p>\n<h4><a href=\"https://www.kaggle.com/competitions/shopee-product-matching/discussion/238078\" target=\"_blank\">5th Place Solution</a></h4>\n<p><strong>Matching</strong></p>\n<p><strong>2nd Stage Model</strong><br>\nThey have extracted some features for matched unique pairs and then fed them to XGB 2nd stage model which runs on GPU. The features where \"img_dist\", \"text_dist\", \"dist\", \"dist_rank\", \"cos_sim\", \"cos_sim2\". Basically, vector distances, their ranking within each posting id, 2 different tfidf cosine similarity with different parameters.<br>\nSince test set size was larger than their one fold size, some features were going to have different distribution on the test set due to higher possibility of False Positive matches. <br>\nTherefore they used another XGB model that was trained on percentage rank features and ensembled with the previous.</p>\n<p><strong>FP Features</strong><br>\nSince test set and train set have no maching between them, one can use training set for determining how easy it is for a posting to have FPs. They used closest match distances from training set as features. This method improved our CV score significantly but improved LB relatively less for them. They claim this is again due to the size difference between the train and test.</p>\n<p><strong>Agglomerative Clustering</strong></p>\n<p>Once they have match probability predictions from XGB models, they could then use a threshold and match the postings. But they did something a bit more complex. They sorted all pairs by their match probabilities and started matching from the most likely match. Each match above 0.8 probability, merges their clusters. Each posting with no cluster can be matched to a cluster if the probability is above 0.7. Each posting with no cluster can be matched with each other if the probability is above 0.3. This method allows to have confident clusters and not-so-confident many pairs.</p>\n<h4><a href=\"https://www.kaggle.com/competitions/shopee-product-matching/discussion/238126\" target=\"_blank\">26th Place Solution - Effective Cluster Separation and Neighbour Search</a></h4>\n<p><strong>Balancing recall by the neighbours</strong></p>\n<p>They realized that we they not doing good on lb because simple concatenation of text and image predictions increases the number of predicted values and hence decreased the recall. Thus they came up with this idea to balance the prediction and recall.</p>\n<p>They take the embeddings from each image model and pass it through the pre-processing layer, in the pre-processing layer they take out top 3 neighbours for each row using knns and multiply it with decreasing weights on log space so as to shift these top 3 embeddings into another cluster on the embedding space, while the other embeddings remain the same, (they tried many numbers instead of 3 but taking 3 best neighbours seem to work the best).</p>\n<p>They do the same with text embeddings and at the end they have 4 modified embeddings from four models.</p>\n<p><strong>Neighbourhood Search</strong></p>\n<p>They had five different models and five different embeddings, they tried a lot of different ensemble strategies that didn't help, they came up with this idea:</p>\n<p>They take each model and find 100 neighbours for each posting id and for each model, they make separate dataframes for each. For eg:- they take the nfnet model get 100 neighbours for each posting_id and form a dataframe, so now their dataframe has posting_id , indexes of neighbours and their distances.</p>\n<p>Now they take all the 5 dataframes, they are not the same as different models have different neigbours for different posting_ids, thus they do outer join, so instead of concatenating the embeddings we concat the distances. Now they have our final dataframe, in this they search for the final predictions using some threshold.</p>\n<h4><a href=\"https://www.kaggle.com/competitions/shopee-product-matching/discussion/238149\" target=\"_blank\">62th Place Solution (Stacking Logistic Regression)</a></h4>\n<p><strong>Ensembling</strong></p>\n<p>Concat all image models' outputs to get final image embedding and calculate cosine similarity (X) of all pairs (total N x N pairs). And do the same thing to get text cosine similarity (Y).<br>\nX:  image cosine similarity<br>\nY:  text cosine similarity</p>\n<p><strong>Downsampling</strong></p>\n<p>For each product, they used all products with the same label_group and sampled an equal amount of negative data from top products sorted by X+Y.<br>\nThey they use polynomial transformation (1,X,X^2,Y,Y^2,XY)<br>\nAnd then use Logistic regression to get final decision boundary.</p>\n<h4><a href=\"https://www.kaggle.com/competitions/shopee-product-matching/discussion/238295\" target=\"_blank\">4th Place Solution</a></h4>\n<p><strong>Ensemble</strong><br>\nOn their ensemble, they concatenated and normalized image embeddings, text embeddings (bert based) and tfdif embeddings. On each of those vectors, they calculated pair wise cosine similarity and received three matrices (cossim image, cossim bert, cossim tfidf). They combined those three matrices by first squareing them and then taking a weighted average.</p>\n<p><strong>Post Processing</strong><br>\nWhile the averaged cossin matrix already gave them a very competitive score when using a proper threshold, they still applied several post processing steps to squeeze out a few more points. </p>\n<p>The biggest contributors were:</p>\n<ul>\n<li>Thresholding</li>\n<li>Rank2 matching (if A has B on rank2 and B has A on rank2, add them to each other)</li>\n<li>Rank2 and rank3 difference is large -&gt; add rank2 id</li>\n<li>If there is a group of X members, and we have another row with the X preds on rank 1-X, make one large group.</li>\n<li>At least one other match (except the cosine similarity of rank2 is extremly low)</li>\n<li>Query Expansion</li>\n<li>Rematching unmatched rows</li>\n</ul>\n<h4><a href=\"https://www.kaggle.com/competitions/shopee-product-matching/discussion/238515\" target=\"_blank\">3rd Place Solution (Triplet loss, Boosting, Clustering)</a></h4>\n<p><strong>Postprocessing (clustering)</strong></p>\n<p>They perform clustering, not over embeddings but over pairwise distances. Distance (better say - similarity) for them was GBM probability of duplicate estimation. If there was no forecast for product pair, assume probability of duplicate eq 0. They took the idea of agglomerative clustering and slightly modify it for current task. They start from each single point as a cluster and merge them until average cluster size become equal threshold. If after clustering point stay single - merge it to nearest cluster. This algorithm boost them up the leaderboard after a lot of parameter search</p>",
      "rawMarkdown": "# Candidates evaluation tricks from previous winning solutions\n\nSince this competition many approaches involve generating candidates and ranking them, It can be useful to take a step back and look at some solutions from past competitions that are based on the same general Idea:\n\n> **Candidates Generation** -> **Rank / Match** -> **Final Candidates**\n\nI am crossposting a previous short summary I made a couple of month ago about this topic since it can be interesting in this context.\n\nEnjoy! \n\n_____\n\n\nA while ago there had been a competition where the participants needed to match up images/text of the same product. \nOn that competition, there were many ensemble / matching tricks deplyed by the winners, mostly done on embeddings extracted either from the images or from the text description.\n\nThis topic summarize all the matching tricks from the solution writeups.\n\n\n#### [Shopee's 1st Place Solution](https://www.kaggle.com/competitions/shopee-product-matching/discussion/238136)\n\n\n**Iterative Neighborhood Blending (INB)**\n\nApart from combining image and text matches, it was also crucial to properly utilize embeddings to produce matches. The author made a nontrivial pipeline for searching matches from embeddings.\nBased on QE(Query Expansion) and DBA(DataBase-side feature Augmentation), they created a pipeline called INB(Iterative Neighborhood Blending). \nMost of the ideas are shared with QE and DBA, but some details are different. INB pipeline consists of these components.\n\n**K Nearest Neighbor Search**\nThey use [faissed](https://github.com/facebookresearch/faiss) for knn search, and set k=51 (maximum 50 non-self matches + 1 self). They used inner product as similarity metric (the embedding is normalized so it is equivalent to cosine similarity)\n\n\n**To decide about the threshold**\n\nThey converted cosine similarity to cosine distance (= 1-cosine similarity) for some convenience in implementation, and obtained (matches, distances) pair that satisfies distance < threshold. For each item x, we call this (matches, distances) pair as \"neighborhood of x\".\n\n\n**Neighborhood Blending**\n\nThe intuition for neighborhood blending is straightforward. After knn and thresholding, they obtain the (matches, similarities) pair for each item, we have a graph, where each node is an item, and the edge weight is the similarity between two nodes. Only the neighborhoods are connected. That is, nodes that didn't pass the threshold condition and min2 condition from the query node, are disconnected.\n\nWe want to use the neighborhood items' information to refine the query item's embedding and make the cluster clearer. In order to do that, they simply weighted-sum the neighborhood embeddings with similarity as weights and add it to the query embedding. So they blend neighborhood embeddings. They call it NB(Neighborhood Blending).\n\n![](https://i.ibb.co/F4BCKjX/3.jpg)\n\n\nThe image illustrates how one step of neighborhood blending is performed on a toy example. Let's look at node A. Its embedding is [-0.588, 0.784, 0.196] and its similarity to node B, C, D is 0.94, 0.93, 0.52 respectively. Red line means two nodes are neighbors, so they are connected. Dashed line means two nodes didn't pass the threshold, so are disconnected. \n\nThey can apply NB iteratively. After blending neighborhood for stage1 embeddings, we do knn search & 'thresholding with min2' to get stage2 (matches, similarities). \nThey apply NB again, to further refine the embeddings. We can iterate until the evaluation metric stops improving. This is where Iterative comes from.\n\n\nCode:\n\n\n```\n\ndef blend_neighborhood(emb, match_index_lst, similarities_lst):\n    new_emb = emb.copy()\n    for i in range(emb.shape[0]):\n        cur_emb = emb[match_index_lst[i]]\n        weights = np.expand_dims(similarities_lst[i], 1)\n        new_emb[i] = (cur_emb * weights).sum(axis=0)\n    new_emb = normalize(new_emb, axis=1)\n    return new_emb\n\ndef iterative_neighborhood_blending(emb, threshes):\n    for thresh in threshes:\n        match_index_lst, similarities_lst = neighborhood_search(emb, thresh)\n        emb = blend_neighborhood(emb, match_index_lst, similarities_lst)\n    return match_index_lst\n```\n    \n    \n    \n#### [2nd place solution (matching prediction by GAT & LGB)](https://www.kaggle.com/competitions/shopee-product-matching/discussion/238022)\n\n\nThe second place of this competition used some extra features for their 2nd level models such as:\n\n- Image similarity\n- Cosine similarities of NFNet-F0, ViT embeddings\n- Loss: CurricularFace (better than ArcFace and others)\n- Optimizer: SAM (better than Adam, SGD, and others)\n- Concatenate the similarities like F.normalize(torch.cat([F.normalize(emb1), F.normalize(emb2)], axis=1))\n- Text similarity\n- Cosine similarities of Indonesian-BERT, Multilingual-BERT, and Paraphrase-XLM embeddings\n- TF-IDF as in many public kernels\n- Multimodal (image + text) similarity\n- Trained model with NFNet-F0 and Indonesian BERT (concatenated at final feature layers)\n- Graph features\n- Avg and std of top-K cosine similarities of each item\n- K=5, 10, 15, 30, etc\n- Pagerank\n- Text length\n- # of word\n- Levenshtein distance\n- Image file size\n- Width and height of image\n- Query Expansion\n- Obtain an augmented embedding which weighted average neighbors.\n- Concatenate the original and augmented embeddings like F.normalize(torch.cat([F.normalize(orig_emb), F.normalize(qe_emb)], axis=1))\n\n\n**Graph Attention Networks**\n\nThey then constructed a model that based on Graph Attention Networks (GAT)\n\nThey chose GAT due to it's simplicity and customizability\n\nAt the end, they used only 4 features: image/bert/multi-modal/tf-idf similarities\nApply graph attention to \"other edges connected to the node connected to the target edge\" as a neighborhood\n\n![](https://user-images.githubusercontent.com/27487010/123567967-47ba7880-d7fe-11eb-9624-19fd8fd5d037.png)\n\n\n**Graph-Based Post-processing**\n\nThey went further on and recursively removed edges that had the highest betweenness centrality\n\n- The Intuition behind this is that Label groups should form a clique If there is an abundant edge, it should bridge two clique such an edge should have a higher betweenness centrality.\n- Recursively removing such an edge until there is no such an edge or connected components become smaller than some threshold pushed their score by a large margin.\n\n\n#### [18th place solution](https://www.kaggle.com/competitions/shopee-product-matching/discussion/237972)\n\n**Combining model outputs**\n\nThey predicted with all their 1st stage models and then they replaced the predictions of each row with an average of predictions of its nearest neighbors. \nThey set threshold for “nearest” so that 4 neighbors are used on average (this was found experimentally).\n\n**Forcing groups into the desired distribution**\n\nThis was the largest single trick that greatly improved their score. \nThey first make an educated guess that the distribution of group targets in the test data is similar to the one in train. Then they try to make their predictions have the same shape.\nThey first decide that groups with 2 elements are going to be those where the third largest element is the lowest. They follow the same logic for all sizes up to 50 and submit the best.\n\n\n**Cross-mean embedding**\n\nThey then applied some simple huristics to boost the score a bit further: \n- If there are products(e.g. A,B and C) which have identical image, they are same product. (precision > 99.9%)\n- They replace each *title* embedding to mean of title embedding: (A+B+C)/3\n- Same procedure for image embedding, mean by same title. [the competition was about images and texts matching of the same product. \n\n**DBA/QE**\nFor image X, I got 3 nearest neighbor(e.g. X,Y,Z) by image embedding, then replace X's image embedding with weighted(logspace) sum of (X,Y,Z). same goes for title.\n\n**Graph merge**\nNow we have two embeddings: image and title. They Merge them as graph.\n\nFor each image or title embedding, they got 50 nearest neighbors and euclidean distance.\nOne node represents one embedding. Edge weight is their euclidean distance.\nThen We can merge image and title graph. If there is duplicated edge, pick small(near) one.\n\nThey then employed some kind of [Dijkstra's algorithm](https://en.wikipedia.org/wiki/Dijkstra%27s_algorithm) to form clusters.\nAnd chose final threshold from LB score.\n\n\n![](https://i.imgur.com/4OLmNIc.png)\n\n![](https://i.imgur.com/Fp2gsrD.png)\n\n\n\n#### [5th Place Solution](https://www.kaggle.com/competitions/shopee-product-matching/discussion/238078)\n\n\n**Matching**\n\n\n**2nd Stage Model**\nThey have extracted some features for matched unique pairs and then fed them to XGB 2nd stage model which runs on GPU. The features where \"img_dist\", \"text_dist\", \"dist\", \"dist_rank\", \"cos_sim\", \"cos_sim2\". Basically, vector distances, their ranking within each posting id, 2 different tfidf cosine similarity with different parameters.\nSince test set size was larger than their one fold size, some features were going to have different distribution on the test set due to higher possibility of False Positive matches. \nTherefore they used another XGB model that was trained on percentage rank features and ensembled with the previous.\n\n**FP Features**\nSince test set and train set have no maching between them, one can use training set for determining how easy it is for a posting to have FPs. They used closest match distances from training set as features. This method improved our CV score significantly but improved LB relatively less for them. They claim this is again due to the size difference between the train and test.\n\n**Agglomerative Clustering**\n\nOnce they have match probability predictions from XGB models, they could then use a threshold and match the postings. But they did something a bit more complex. They sorted all pairs by their match probabilities and started matching from the most likely match. Each match above 0.8 probability, merges their clusters. Each posting with no cluster can be matched to a cluster if the probability is above 0.7. Each posting with no cluster can be matched with each other if the probability is above 0.3. This method allows to have confident clusters and not-so-confident many pairs.\n\n\n\n#### [26th Place Solution - Effective Cluster Separation and Neighbour Search](https://www.kaggle.com/competitions/shopee-product-matching/discussion/238126)\n\n\n**Balancing recall by the neighbours**\n\n\nThey realized that we they not doing good on lb because simple concatenation of text and image predictions increases the number of predicted values and hence decreased the recall. Thus they came up with this idea to balance the prediction and recall.\n\nThey take the embeddings from each image model and pass it through the pre-processing layer, in the pre-processing layer they take out top 3 neighbours for each row using knns and multiply it with decreasing weights on log space so as to shift these top 3 embeddings into another cluster on the embedding space, while the other embeddings remain the same, (they tried many numbers instead of 3 but taking 3 best neighbours seem to work the best).\n\nThey do the same with text embeddings and at the end they have 4 modified embeddings from four models.\n\n\n**Neighbourhood Search**\n\nThey had five different models and five different embeddings, they tried a lot of different ensemble strategies that didn't help, they came up with this idea:\n\nThey take each model and find 100 neighbours for each posting id and for each model, they make separate dataframes for each. For eg:- they take the nfnet model get 100 neighbours for each posting_id and form a dataframe, so now their dataframe has posting_id , indexes of neighbours and their distances.\n\nNow they take all the 5 dataframes, they are not the same as different models have different neigbours for different posting_ids, thus they do outer join, so instead of concatenating the embeddings we concat the distances. Now they have our final dataframe, in this they search for the final predictions using some threshold.\n\n\n#### [62th Place Solution (Stacking Logistic Regression)](https://www.kaggle.com/competitions/shopee-product-matching/discussion/238149)\n\n\n**Ensembling**\n\nConcat all image models' outputs to get final image embedding and calculate cosine similarity (X) of all pairs (total N x N pairs). And do the same thing to get text cosine similarity (Y).\nX:  image cosine similarity\nY:  text cosine similarity\n\n**Downsampling**\n\nFor each product, they used all products with the same label_group and sampled an equal amount of negative data from top products sorted by X+Y.\nThey they use polynomial transformation (1,X,X^2,Y,Y^2,XY)\nAnd then use Logistic regression to get final decision boundary.\n\n\n\n#### [4th Place Solution](https://www.kaggle.com/competitions/shopee-product-matching/discussion/238295)\n\n**Ensemble**\nOn their ensemble, they concatenated and normalized image embeddings, text embeddings (bert based) and tfdif embeddings. On each of those vectors, they calculated pair wise cosine similarity and received three matrices (cossim image, cossim bert, cossim tfidf). They combined those three matrices by first squareing them and then taking a weighted average.\n\n**Post Processing**\nWhile the averaged cossin matrix already gave them a very competitive score when using a proper threshold, they still applied several post processing steps to squeeze out a few more points. \n\nThe biggest contributors were:\n\n- Thresholding\n- Rank2 matching (if A has B on rank2 and B has A on rank2, add them to each other)\n- Rank2 and rank3 difference is large -> add rank2 id\n- If there is a group of X members, and we have another row with the X preds on rank 1-X, make one large group.\n- At least one other match (except the cosine similarity of rank2 is extremly low)\n- Query Expansion\n- Rematching unmatched rows\n\n\n#### [3rd Place Solution (Triplet loss, Boosting, Clustering)](https://www.kaggle.com/competitions/shopee-product-matching/discussion/238515)\n\n**Postprocessing (clustering)**\n\nThey perform clustering, not over embeddings but over pairwise distances. Distance (better say - similarity) for them was GBM probability of duplicate estimation. If there was no forecast for product pair, assume probability of duplicate eq 0. They took the idea of agglomerative clustering and slightly modify it for current task. They start from each single point as a cluster and merge them until average cluster size become equal threshold. If after clustering point stay single - merge it to nearest cluster. This algorithm boost them up the leaderboard after a lot of parameter search\n\n\n",
      "votes": 13
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2090262": "# Candidates evaluation tricks from previous winning solutions\n\nSince this competition many approaches involve generating candidates and ranking them, It can be useful to take a step back and look at some solutions from past competitions that are based on the same general Idea:\n\n> **Candidates Generation** -> **Rank / Match** -> **Final Candidates**\n\nI am crossposting a previous short summary I made a couple of month ago about this topic since it can be interesting in this context.\n\nEnjoy! \n\n_____\n\n\nA while ago there had been a competition where the participants needed to match up images/text of the same product. \nOn that competition, there were many ensemble / matching tricks deplyed by the winners, mostly done on embeddings extracted either from the images or from the text description.\n\nThis topic summarize all the matching tricks from the solution writeups.\n\n\n#### [Shopee's 1st Place Solution](https://www.kaggle.com/competitions/shopee-product-matching/discussion/238136)\n\n\n**Iterative Neighborhood Blending (INB)**\n\nApart from combining image and text matches, it was also crucial to properly utilize embeddings to produce matches. The author made a nontrivial pipeline for searching matches from embeddings.\nBased on QE(Query Expansion) and DBA(DataBase-side feature Augmentation), they created a pipeline called INB(Iterative Neighborhood Blending). \nMost of the ideas are shared with QE and DBA, but some details are different. INB pipeline consists of these components.\n\n**K Nearest Neighbor Search**\nThey use [faissed](https://github.com/facebookresearch/faiss) for knn search, and set k=51 (maximum 50 non-self matches + 1 self). They used inner product as similarity metric (the embedding is normalized so it is equivalent to cosine similarity)\n\n\n**To decide about the threshold**\n\nThey converted cosine similarity to cosine distance (= 1-cosine similarity) for some convenience in implementation, and obtained (matches, distances) pair that satisfies distance < threshold. For each item x, we call this (matches, distances) pair as \"neighborhood of x\".\n\n\n**Neighborhood Blending**\n\nThe intuition for neighborhood blending is straightforward. After knn and thresholding, they obtain the (matches, similarities) pair for each item, we have a graph, where each node is an item, and the edge weight is the similarity between two nodes. Only the neighborhoods are connected. That is, nodes that didn't pass the threshold condition and min2 condition from the query node, are disconnected.\n\nWe want to use the neighborhood items' information to refine the query item's embedding and make the cluster clearer. In order to do that, they simply weighted-sum the neighborhood embeddings with similarity as weights and add it to the query embedding. So they blend neighborhood embeddings. They call it NB(Neighborhood Blending).\n\n![](https://i.ibb.co/F4BCKjX/3.jpg)\n\n\nThe image illustrates how one step of neighborhood blending is performed on a toy example. Let's look at node A. Its embedding is [-0.588, 0.784, 0.196] and its similarity to node B, C, D is 0.94, 0.93, 0.52 respectively. Red line means two nodes are neighbors, so they are connected. Dashed line means two nodes didn't pass the threshold, so are disconnected. \n\nThey can apply NB iteratively. After blending neighborhood for stage1 embeddings, we do knn search & 'thresholding with min2' to get stage2 (matches, similarities). \nThey apply NB again, to further refine the embeddings. We can iterate until the evaluation metric stops improving. This is where Iterative comes from.\n\n\nCode:\n\n\n```\n\ndef blend_neighborhood(emb, match_index_lst, similarities_lst):\n    new_emb = emb.copy()\n    for i in range(emb.shape[0]):\n        cur_emb = emb[match_index_lst[i]]\n        weights = np.expand_dims(similarities_lst[i], 1)\n        new_emb[i] = (cur_emb * weights).sum(axis=0)\n    new_emb = normalize(new_emb, axis=1)\n    return new_emb\n\ndef iterative_neighborhood_blending(emb, threshes):\n    for thresh in threshes:\n        match_index_lst, similarities_lst = neighborhood_search(emb, thresh)\n        emb = blend_neighborhood(emb, match_index_lst, similarities_lst)\n    return match_index_lst\n```\n    \n    \n    \n#### [2nd place solution (matching prediction by GAT & LGB)](https://www.kaggle.com/competitions/shopee-product-matching/discussion/238022)\n\n\nThe second place of this competition used some extra features for their 2nd level models such as:\n\n- Image similarity\n- Cosine similarities of NFNet-F0, ViT embeddings\n- Loss: CurricularFace (better than ArcFace and others)\n- Optimizer: SAM (better than Adam, SGD, and others)\n- Concatenate the similarities like F.normalize(torch.cat([F.normalize(emb1), F.normalize(emb2)], axis=1))\n- Text similarity\n- Cosine similarities of Indonesian-BERT, Multilingual-BERT, and Paraphrase-XLM embeddings\n- TF-IDF as in many public kernels\n- Multimodal (image + text) similarity\n- Trained model with NFNet-F0 and Indonesian BERT (concatenated at final feature layers)\n- Graph features\n- Avg and std of top-K cosine similarities of each item\n- K=5, 10, 15, 30, etc\n- Pagerank\n- Text length\n- # of word\n- Levenshtein distance\n- Image file size\n- Width and height of image\n- Query Expansion\n- Obtain an augmented embedding which weighted average neighbors.\n- Concatenate the original and augmented embeddings like F.normalize(torch.cat([F.normalize(orig_emb), F.normalize(qe_emb)], axis=1))\n\n\n**Graph Attention Networks**\n\nThey then constructed a model that based on Graph Attention Networks (GAT)\n\nThey chose GAT due to it's simplicity and customizability\n\nAt the end, they used only 4 features: image/bert/multi-modal/tf-idf similarities\nApply graph attention to \"other edges connected to the node connected to the target edge\" as a neighborhood\n\n![](https://user-images.githubusercontent.com/27487010/123567967-47ba7880-d7fe-11eb-9624-19fd8fd5d037.png)\n\n\n**Graph-Based Post-processing**\n\nThey went further on and recursively removed edges that had the highest betweenness centrality\n\n- The Intuition behind this is that Label groups should form a clique If there is an abundant edge, it should bridge two clique such an edge should have a higher betweenness centrality.\n- Recursively removing such an edge until there is no such an edge or connected components become smaller than some threshold pushed their score by a large margin.\n\n\n#### [18th place solution](https://www.kaggle.com/competitions/shopee-product-matching/discussion/237972)\n\n**Combining model outputs**\n\nThey predicted with all their 1st stage models and then they replaced the predictions of each row with an average of predictions of its nearest neighbors. \nThey set threshold for “nearest” so that 4 neighbors are used on average (this was found experimentally).\n\n**Forcing groups into the desired distribution**\n\nThis was the largest single trick that greatly improved their score. \nThey first make an educated guess that the distribution of group targets in the test data is similar to the one in train. Then they try to make their predictions have the same shape.\nThey first decide that groups with 2 elements are going to be those where the third largest element is the lowest. They follow the same logic for all sizes up to 50 and submit the best.\n\n\n**Cross-mean embedding**\n\nThey then applied some simple huristics to boost the score a bit further: \n- If there are products(e.g. A,B and C) which have identical image, they are same product. (precision > 99.9%)\n- They replace each *title* embedding to mean of title embedding: (A+B+C)/3\n- Same procedure for image embedding, mean by same title. [the competition was about images and texts matching of the same product. \n\n**DBA/QE**\nFor image X, I got 3 nearest neighbor(e.g. X,Y,Z) by image embedding, then replace X's image embedding with weighted(logspace) sum of (X,Y,Z). same goes for title.\n\n**Graph merge**\nNow we have two embeddings: image and title. They Merge them as graph.\n\nFor each image or title embedding, they got 50 nearest neighbors and euclidean distance.\nOne node represents one embedding. Edge weight is their euclidean distance.\nThen We can merge image and title graph. If there is duplicated edge, pick small(near) one.\n\nThey then employed some kind of [Dijkstra's algorithm](https://en.wikipedia.org/wiki/Dijkstra%27s_algorithm) to form clusters.\nAnd chose final threshold from LB score.\n\n\n![](https://i.imgur.com/4OLmNIc.png)\n\n![](https://i.imgur.com/Fp2gsrD.png)\n\n\n\n#### [5th Place Solution](https://www.kaggle.com/competitions/shopee-product-matching/discussion/238078)\n\n\n**Matching**\n\n\n**2nd Stage Model**\nThey have extracted some features for matched unique pairs and then fed them to XGB 2nd stage model which runs on GPU. The features where \"img_dist\", \"text_dist\", \"dist\", \"dist_rank\", \"cos_sim\", \"cos_sim2\". Basically, vector distances, their ranking within each posting id, 2 different tfidf cosine similarity with different parameters.\nSince test set size was larger than their one fold size, some features were going to have different distribution on the test set due to higher possibility of False Positive matches. \nTherefore they used another XGB model that was trained on percentage rank features and ensembled with the previous.\n\n**FP Features**\nSince test set and train set have no maching between them, one can use training set for determining how easy it is for a posting to have FPs. They used closest match distances from training set as features. This method improved our CV score significantly but improved LB relatively less for them. They claim this is again due to the size difference between the train and test.\n\n**Agglomerative Clustering**\n\nOnce they have match probability predictions from XGB models, they could then use a threshold and match the postings. But they did something a bit more complex. They sorted all pairs by their match probabilities and started matching from the most likely match. Each match above 0.8 probability, merges their clusters. Each posting with no cluster can be matched to a cluster if the probability is above 0.7. Each posting with no cluster can be matched with each other if the probability is above 0.3. This method allows to have confident clusters and not-so-confident many pairs.\n\n\n\n#### [26th Place Solution - Effective Cluster Separation and Neighbour Search](https://www.kaggle.com/competitions/shopee-product-matching/discussion/238126)\n\n\n**Balancing recall by the neighbours**\n\n\nThey realized that we they not doing good on lb because simple concatenation of text and image predictions increases the number of predicted values and hence decreased the recall. Thus they came up with this idea to balance the prediction and recall.\n\nThey take the embeddings from each image model and pass it through the pre-processing layer, in the pre-processing layer they take out top 3 neighbours for each row using knns and multiply it with decreasing weights on log space so as to shift these top 3 embeddings into another cluster on the embedding space, while the other embeddings remain the same, (they tried many numbers instead of 3 but taking 3 best neighbours seem to work the best).\n\nThey do the same with text embeddings and at the end they have 4 modified embeddings from four models.\n\n\n**Neighbourhood Search**\n\nThey had five different models and five different embeddings, they tried a lot of different ensemble strategies that didn't help, they came up with this idea:\n\nThey take each model and find 100 neighbours for each posting id and for each model, they make separate dataframes for each. For eg:- they take the nfnet model get 100 neighbours for each posting_id and form a dataframe, so now their dataframe has posting_id , indexes of neighbours and their distances.\n\nNow they take all the 5 dataframes, they are not the same as different models have different neigbours for different posting_ids, thus they do outer join, so instead of concatenating the embeddings we concat the distances. Now they have our final dataframe, in this they search for the final predictions using some threshold.\n\n\n#### [62th Place Solution (Stacking Logistic Regression)](https://www.kaggle.com/competitions/shopee-product-matching/discussion/238149)\n\n\n**Ensembling**\n\nConcat all image models' outputs to get final image embedding and calculate cosine similarity (X) of all pairs (total N x N pairs). And do the same thing to get text cosine similarity (Y).\nX:  image cosine similarity\nY:  text cosine similarity\n\n**Downsampling**\n\nFor each product, they used all products with the same label_group and sampled an equal amount of negative data from top products sorted by X+Y.\nThey they use polynomial transformation (1,X,X^2,Y,Y^2,XY)\nAnd then use Logistic regression to get final decision boundary.\n\n\n\n#### [4th Place Solution](https://www.kaggle.com/competitions/shopee-product-matching/discussion/238295)\n\n**Ensemble**\nOn their ensemble, they concatenated and normalized image embeddings, text embeddings (bert based) and tfdif embeddings. On each of those vectors, they calculated pair wise cosine similarity and received three matrices (cossim image, cossim bert, cossim tfidf). They combined those three matrices by first squareing them and then taking a weighted average.\n\n**Post Processing**\nWhile the averaged cossin matrix already gave them a very competitive score when using a proper threshold, they still applied several post processing steps to squeeze out a few more points. \n\nThe biggest contributors were:\n\n- Thresholding\n- Rank2 matching (if A has B on rank2 and B has A on rank2, add them to each other)\n- Rank2 and rank3 difference is large -> add rank2 id\n- If there is a group of X members, and we have another row with the X preds on rank 1-X, make one large group.\n- At least one other match (except the cosine similarity of rank2 is extremly low)\n- Query Expansion\n- Rematching unmatched rows\n\n\n#### [3rd Place Solution (Triplet loss, Boosting, Clustering)](https://www.kaggle.com/competitions/shopee-product-matching/discussion/238515)\n\n**Postprocessing (clustering)**\n\nThey perform clustering, not over embeddings but over pairwise distances. Distance (better say - similarity) for them was GBM probability of duplicate estimation. If there was no forecast for product pair, assume probability of duplicate eq 0. They took the idea of agglomerative clustering and slightly modify it for current task. They start from each single point as a cluster and merge them until average cluster size become equal threshold. If after clustering point stay single - merge it to nearest cluster. This algorithm boost them up the leaderboard after a lot of parameter search\n\n\n"
  }
}