{
  "id": 311933,
  "title": "Recap of the Top Solutions from a Similar Competition (Shopee)",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/discussion/311933",
  "author_name": "",
  "post_date": "2022-03-09T14:46:42.241834Z",
  "votes": 9,
  "comment_count": 4,
  "views": 0,
  "content": "<p><a href=\"https://www.kaggle.com/c/shopee-product-matching/overview\" target=\"_blank\">https://www.kaggle.com/c/shopee-product-matching/overview</a></p>\n<p><em>There were Indonesian texts in this competition.</em></p>\n<blockquote>\n  <p>While the modeling part is somewhat similar to each other, every solution brings something valuable to the table in terms of post-processing. So highly recommended to study them.</p>\n</blockquote>\n<ul>\n<li><p><a href=\"https://www.kaggle.com/c/shopee-product-matching/discussion/238136\" target=\"_blank\">1st Place Solution</a></p>\n<ul>\n<li>Two eca_nfnet_l1s from timm for image encoders, and xlm-roberta-large, xlm-roberta-base, cahya/bert-base-indonesian-1.5G, indobenchmark/indobert-large-p1, bert-base-multilingual-uncased from huggingface for text encoders.</li>\n<li>Used ArcFace to train the model. After pooling from image/text encoder, applied batchnormalization and feature-wise normalization to output embedding.<ul>\n<li>increase margin gradually while training</li>\n<li>use large warmup steps</li>\n<li>use larger learning rate for cosinehead</li>\n<li>use gradient clipping</li></ul></li>\n<li>Adding extra fc layers after global average pooling hurts the model performance, while adding batchnorm before feature-wise normalization improved the score.</li>\n<li>Normalize image embedding and text embedding then concat them to calculate comb-similarities.</li>\n<li>Iterative Neighborhood Blending (an approach to refine the embeddings which includes QE(Query Expansion) and DBA(DataBase-side feature Augmentation))</li></ul></li>\n<li><p><a href=\"https://www.kaggle.com/c/shopee-product-matching/discussion/238022\" target=\"_blank\">2nd Place Solution</a></p>\n<ul>\n<li>1st stage<ul>\n<li>Image<ul>\n<li>Cosine similarities of NFNet-F0, ViT embeddings</li>\n<li>Loss: CurricularFace</li>\n<li>Optimizer: SAM</li>\n<li>Concatenate the similarities like F.normalize(torch.cat([F.normalize(emb1), F.normalize(emb2)], axis=1))</li></ul></li>\n<li>Text<ul>\n<li>Cosine similarities of Indonesian-BERT, Multilingual-BERT, and Paraphrase-XLM embeddings</li>\n<li>TF-IDF</li></ul></li>\n<li>Image + Text<ul>\n<li>Trained model with NFNet-F0 and Indonesian BERT (concatenated at final feature layers)</li></ul></li></ul></li>\n<li>2nd stage: Train “meta” models to classify whether a pair of items belong to the same label group or not.<br>\nUsed LightGBM and GAT (Graph Attention Networks)</li>\n<li>Graph-Based post-processing</li></ul></li>\n<li><p><a href=\"https://www.kaggle.com/c/shopee-product-matching/discussion/238515\" target=\"_blank\">3rd Place Solution</a></p>\n<ul>\n<li>Build different (not trainable) products representations: efficient net embeddings, BERTs, TfIdf and search for nearest neighbours on fixed radius.</li>\n<li>For each object, union candidate neighbours from all embeddings and build a sample of pair objects with a binary target - if pair is a real duplicate or not. </li>\n<li>On that sample, build simple GradientBoosting model (catboost in my case) using next features, calculates separately for each embedding: pairwise distances (cosine, euclidean and etc), density around both points (frequency of points on different radiuses), points ranks.</li>\n<li>After model is tranined, threshold candidate points by probability of duplicate, that is searched by CV.</li></ul></li>\n<li><p><a href=\"https://www.kaggle.com/c/shopee-product-matching/discussion/238295\" target=\"_blank\">4th Place Solution</a></p>\n<ul>\n<li>Used Bert based Text Encoder and Image Encoders with different backbones (mainly nfnet, effnet)</li>\n<li>After the backbone, GeM and Avg pooling was used. The neck is comprised of a linear layer that is used for dimension tuning, followed by batch normalization and PReLu activation. </li>\n<li>Trained the image and text models with an ArcFaceLoss. Tuned the ArcFace margin.</li>\n<li>Concatenated and normalized image embeddings, text embeddings (bert based) and tfdif embeddings.</li>\n<li>On each of those vectors, calculated pair wise cosine similarity and received three matrices (cossim image, cossim bert, cossim tfidf). Combined those three matrices by first squareing them and then taking a weighted average.</li>\n<li>Postprocessing: <ul>\n<li>Thresholding </li>\n<li>Rank2 matching (if A has B on rank2 and B has A on rank2, add them to each other) </li>\n<li>Rank2 and rank3 difference is large -&gt; add rank2 id</li>\n<li>If there is a group of X members, and we have another row with the X preds on rank 1-X, make one large group.</li>\n<li>At least one other match (except the cosine similarity of rank2 is extremly low)</li>\n<li>Query Expansion </li>\n<li>Rematching unmatched rows</li></ul></li></ul></li>\n<li><p><a href=\"https://www.kaggle.com/c/shopee-product-matching/discussion/238078\" target=\"_blank\">5th Place Solution</a></p>\n<ul>\n<li>English Distilbert and Indonesian Distilbert models are trained for obtaining the title vectors. </li>\n<li>For the image vectors, ViT and Swin Transformers are trained on size 384 with relatively heavy augmentations and also EffNet B4 on size 512. </li>\n<li>Ensemble these models with vector concatenation.</li>\n<li>Weighted Database Augmentation</li>\n<li>Used cuML's NearestNeighbors for matching the vectors and used a threshold for filtering.</li>\n<li>Extracted  features for matched unique pairs and fed them to XGB model. The features are \"img_dist\", \"text_dist\", \"dist\", \"dist_rank\", \"cos_sim\", \"cos_sim2\". Basically, vector ditances, their ranking within each posting id, 2 different tfidf cosine similarity with different parameters.</li>\n<li>FP Features:  used closest match distances from training set as features.</li>\n<li>Agglomerative Clustering: After having the match probability predictions from XGB models, sorted all pairs by their match probabilities and started matching from the most likely match. Each match above 0.8 probability, merges their clusters</li></ul></li>\n<li><p><a href=\"https://www.kaggle.com/c/shopee-product-matching/discussion/238010\" target=\"_blank\">6th Place Solution</a></p>\n<ul>\n<li>img_backbone: swin_large_patch4_window7_224, efficientnet_b3</li>\n<li>img_size: 224, 560</li>\n<li>text_backbone    xlm-roberta-base, bert-base-multilingual-uncased, bert-base-uncased</li>\n<li>Euclidean distance and cosine similarity to get nearest neighbors.</li>\n<li>Ensemble: 4 models * 3 output(concat, text, img) * 2 distances = 24 votes</li>\n<li>Different learning rate for cnn/bert/fc</li></ul></li>\n<li><p><a href=\"https://www.kaggle.com/c/shopee-product-matching/discussion/238174\" target=\"_blank\">7th Place Solution</a></p>\n<ul>\n<li><p>Image model</p>\n<ul>\n<li>Backbone : NFNet-F0 with GeM pooling (p=3(fixed) for training, p=4 for inference)</li>\n<li>Input size : 420x420</li>\n<li>Epochs : 30</li>\n<li>Optimizer : madgrad(momentum=0.9, weight_decay=1e-5) [1]</li>\n<li>Learning rate scheduler : cosine annealing (1e-5 --&gt; 1e-8)</li>\n<li>Distance : cosine similarity</li>\n<li>Loss : multi-similarity loss (alpha=2, beta=50, base=0.5) with XBM(memory_size=1024)</li>\n<li>Batch size : 32, mini-batch is generated by randomly sampling pairs of images in same label_group</li>\n<li>Data augmentation : Rotate, ColorJitter, RandomBrightnessContrast, RandomGamma, HorizontalFlip, CoarseDropout</li></ul></li>\n<li><p>Text model</p>\n<ul>\n<li>Backbone : DistilBERT (distilbert-base-indonesian)</li>\n<li>pochs : 30</li>\n<li>Optimizer : madgrad(momentum=0.9, weight_decay=1e-5)</li>\n<li>Learning rate scheduler : cosine annealing (1e-4 --&gt; 1e-8)</li>\n<li>Distance : cosine similarity</li>\n<li>Loss : multi-similarity loss (alpha=2, beta=50, base=0.5)</li>\n<li>Miner : multi-similarity miner (epsilon=0.1)</li>\n<li>Batch size : 256, mini-batch is generated by randomly sampling pairs of titles in same label_group</li>\n<li>Data augmentation : OneOf([random_delete, random_swap, random_swap+random_delete, random_swap*2]) at p=0.1<ul>\n<li>random_delete : randomly delete a word if len(title) &gt; 2</li>\n<li>random_swap : randomly swap two words if len(title) &gt; 2</li></ul></li></ul></li>\n<li><p>alpha query expansion</p></li></ul></li>\n<li><p><a href=\"https://www.kaggle.com/c/shopee-product-matching/discussion/238125\" target=\"_blank\">8th Place Solution</a></p>\n<ul>\n<li><p>Image Embedding</p>\n<ul>\n<li>ResNet152x2, ResNet101x3 with GeM Pooling (5fold)</li>\n<li>Loss: CosFace</li>\n<li>Optimizer SGD lr=1e-3 WarmupCosineAnnealing LR Scheduling</li>\n<li>Input size 512x512</li>\n<li>Embedding dimension 512(ResNet152), 768(ResNet101)</li></ul></li>\n<li><p>Text Embedding</p>\n<ul>\n<li>distilbert_base_indonesian (5fold)</li>\n<li>Concatenate mean of 4,5,6 Layer, CLS and mean of token embeddings (total 3840dim)</li>\n<li>Add FC and Tanh activation to reduce embedding dimension from 3840 to 1536</li>\n<li>Loss: ArcFace</li>\n<li>Optimizer: AdamW lr=1e-4 WarmupLinear LR Scheduling</li></ul></li>\n<li><p>Brute-force kNN by Faiss on embeddings converted to fp16. </p></li>\n<li><p>Use αQE + DBA: 0.59 Threshold for cosine similarity (local CV best threshold +0.15)</p></li></ul></li>\n<li><p><a href=\"https://www.kaggle.com/c/shopee-product-matching/discussion/238039\" target=\"_blank\">10th Place Solution</a></p>\n<ul>\n<li>Trained each models by ArcFace.</li>\n<li>Image: swin-base-224, effnet-b6, resnet101d  </li>\n<li>Text: indobenchmark/indobert-base-p2, indobenchmark/indobert-large-p2 and TFIDF.</li>\n<li>Query Expansion: αQE with n=2 and normalized similarity</li></ul></li>\n<li><p><a href=\"https://www.kaggle.com/c/shopee-product-matching/discussion/238181\" target=\"_blank\">11th Place Solution</a></p>\n<ul>\n<li><p>Image model</p>\n<ul>\n<li>backbone: swin_base_patch4_window12_384</li>\n<li>input_size : 384</li>\n<li>loss : ArcFace (scale=34, margine=0.5)</li>\n<li>fc layer (embedding) dimensions : 768</li>\n<li>optimizer : Ranger</li>\n<li>augmentations : RandAugment</li></ul></li>\n<li><p>BERT models</p>\n<ul>\n<li>backbone_model1 : sentence-transformers/paraphrase-xlm-r-multilingual-v1</li>\n<li>backbone_model2 : cahya/distilbert-base-indonesian</li>\n<li>loss : ArcFace (scale=30, margine=0.5)</li>\n<li>fc layer (embedding) dimensions : 768</li>\n<li>optimizer : SAM with AdamW</li></ul></li></ul></li>\n<li><p><a href=\"https://www.kaggle.com/c/shopee-product-matching/discussion/238033\" target=\"_blank\">14th Place Solution</a></p>\n<ul>\n<li>RAPIDS TfidfVectorizer and EfficientNetB0 384x384 images ensembled with XLM-RoBERTa which has multilingual pretraining and ensembled with EfficientNetB3 512x512 image.</li>\n<li>Make a decision boundary using piecewise linear functions</li>\n<li>Do five additional techniques to remove false negatives and false positives (definitely worth to study)</li></ul></li>\n<li><p><a href=\"https://www.kaggle.com/c/shopee-product-matching/discussion/238029\" target=\"_blank\">15th Place Solution</a></p>\n<ul>\n<li>Image:<ul>\n<li>resnest101, resnest200, eca_nfnet_l1</li>\n<li>Trained with Arcface and GeM. Input size is 512x512.</li>\n<li>Usual augmentation(LR flip, cutout, bright, shiftscalerotate, etc) by albumentations.</li>\n<li>512 dimension for each model, concatenate them to produce 1536 dimensional vector.</li></ul></li>\n<li>Text<ul>\n<li>TF-IDF: Fit and transform to test dataset.</li>\n<li>LaBSE: Finetuned with Arcface and GeM. Augmentation by EDA(Easy Data Augmentation)-like method.</li></ul></li>\n<li>Post-processing<ul>\n<li>Cross-mean embedding: If there are products(e.g. A,B and C) which have identical image, they are same product. (precision &gt; 99.9%). Replace each <em>title</em> embedding to mean of title embedding: (A+B+C)/3. Same procedure for image embedding, mean by same title.</li>\n<li>DBA/QE: For image X, got 3 nearest neighbor(e.g. X,Y,Z) by image embedding, then replace X's image embedding with weighted(logspace) sum of (X,Y,Z). Same for title.</li></ul></li></ul></li>\n</ul>",
  "messages": [
    {
      "id": "1716970",
      "postDate": "03/09/2022 14:46:42",
      "content": "<p><a href=\"https://www.kaggle.com/c/shopee-product-matching/overview\" target=\"_blank\">https://www.kaggle.com/c/shopee-product-matching/overview</a></p>\n<p><em>There were Indonesian texts in this competition.</em></p>\n<blockquote>\n  <p>While the modeling part is somewhat similar to each other, every solution brings something valuable to the table in terms of post-processing. So highly recommended to study them.</p>\n</blockquote>\n<ul>\n<li><p><a href=\"https://www.kaggle.com/c/shopee-product-matching/discussion/238136\" target=\"_blank\">1st Place Solution</a></p>\n<ul>\n<li>Two eca_nfnet_l1s from timm for image encoders, and xlm-roberta-large, xlm-roberta-base, cahya/bert-base-indonesian-1.5G, indobenchmark/indobert-large-p1, bert-base-multilingual-uncased from huggingface for text encoders.</li>\n<li>Used ArcFace to train the model. After pooling from image/text encoder, applied batchnormalization and feature-wise normalization to output embedding.<ul>\n<li>increase margin gradually while training</li>\n<li>use large warmup steps</li>\n<li>use larger learning rate for cosinehead</li>\n<li>use gradient clipping</li></ul></li>\n<li>Adding extra fc layers after global average pooling hurts the model performance, while adding batchnorm before feature-wise normalization improved the score.</li>\n<li>Normalize image embedding and text embedding then concat them to calculate comb-similarities.</li>\n<li>Iterative Neighborhood Blending (an approach to refine the embeddings which includes QE(Query Expansion) and DBA(DataBase-side feature Augmentation))</li></ul></li>\n<li><p><a href=\"https://www.kaggle.com/c/shopee-product-matching/discussion/238022\" target=\"_blank\">2nd Place Solution</a></p>\n<ul>\n<li>1st stage<ul>\n<li>Image<ul>\n<li>Cosine similarities of NFNet-F0, ViT embeddings</li>\n<li>Loss: CurricularFace</li>\n<li>Optimizer: SAM</li>\n<li>Concatenate the similarities like F.normalize(torch.cat([F.normalize(emb1), F.normalize(emb2)], axis=1))</li></ul></li>\n<li>Text<ul>\n<li>Cosine similarities of Indonesian-BERT, Multilingual-BERT, and Paraphrase-XLM embeddings</li>\n<li>TF-IDF</li></ul></li>\n<li>Image + Text<ul>\n<li>Trained model with NFNet-F0 and Indonesian BERT (concatenated at final feature layers)</li></ul></li></ul></li>\n<li>2nd stage: Train “meta” models to classify whether a pair of items belong to the same label group or not.<br>\nUsed LightGBM and GAT (Graph Attention Networks)</li>\n<li>Graph-Based post-processing</li></ul></li>\n<li><p><a href=\"https://www.kaggle.com/c/shopee-product-matching/discussion/238515\" target=\"_blank\">3rd Place Solution</a></p>\n<ul>\n<li>Build different (not trainable) products representations: efficient net embeddings, BERTs, TfIdf and search for nearest neighbours on fixed radius.</li>\n<li>For each object, union candidate neighbours from all embeddings and build a sample of pair objects with a binary target - if pair is a real duplicate or not. </li>\n<li>On that sample, build simple GradientBoosting model (catboost in my case) using next features, calculates separately for each embedding: pairwise distances (cosine, euclidean and etc), density around both points (frequency of points on different radiuses), points ranks.</li>\n<li>After model is tranined, threshold candidate points by probability of duplicate, that is searched by CV.</li></ul></li>\n<li><p><a href=\"https://www.kaggle.com/c/shopee-product-matching/discussion/238295\" target=\"_blank\">4th Place Solution</a></p>\n<ul>\n<li>Used Bert based Text Encoder and Image Encoders with different backbones (mainly nfnet, effnet)</li>\n<li>After the backbone, GeM and Avg pooling was used. The neck is comprised of a linear layer that is used for dimension tuning, followed by batch normalization and PReLu activation. </li>\n<li>Trained the image and text models with an ArcFaceLoss. Tuned the ArcFace margin.</li>\n<li>Concatenated and normalized image embeddings, text embeddings (bert based) and tfdif embeddings.</li>\n<li>On each of those vectors, calculated pair wise cosine similarity and received three matrices (cossim image, cossim bert, cossim tfidf). Combined those three matrices by first squareing them and then taking a weighted average.</li>\n<li>Postprocessing: <ul>\n<li>Thresholding </li>\n<li>Rank2 matching (if A has B on rank2 and B has A on rank2, add them to each other) </li>\n<li>Rank2 and rank3 difference is large -&gt; add rank2 id</li>\n<li>If there is a group of X members, and we have another row with the X preds on rank 1-X, make one large group.</li>\n<li>At least one other match (except the cosine similarity of rank2 is extremly low)</li>\n<li>Query Expansion </li>\n<li>Rematching unmatched rows</li></ul></li></ul></li>\n<li><p><a href=\"https://www.kaggle.com/c/shopee-product-matching/discussion/238078\" target=\"_blank\">5th Place Solution</a></p>\n<ul>\n<li>English Distilbert and Indonesian Distilbert models are trained for obtaining the title vectors. </li>\n<li>For the image vectors, ViT and Swin Transformers are trained on size 384 with relatively heavy augmentations and also EffNet B4 on size 512. </li>\n<li>Ensemble these models with vector concatenation.</li>\n<li>Weighted Database Augmentation</li>\n<li>Used cuML's NearestNeighbors for matching the vectors and used a threshold for filtering.</li>\n<li>Extracted  features for matched unique pairs and fed them to XGB model. The features are \"img_dist\", \"text_dist\", \"dist\", \"dist_rank\", \"cos_sim\", \"cos_sim2\". Basically, vector ditances, their ranking within each posting id, 2 different tfidf cosine similarity with different parameters.</li>\n<li>FP Features:  used closest match distances from training set as features.</li>\n<li>Agglomerative Clustering: After having the match probability predictions from XGB models, sorted all pairs by their match probabilities and started matching from the most likely match. Each match above 0.8 probability, merges their clusters</li></ul></li>\n<li><p><a href=\"https://www.kaggle.com/c/shopee-product-matching/discussion/238010\" target=\"_blank\">6th Place Solution</a></p>\n<ul>\n<li>img_backbone: swin_large_patch4_window7_224, efficientnet_b3</li>\n<li>img_size: 224, 560</li>\n<li>text_backbone    xlm-roberta-base, bert-base-multilingual-uncased, bert-base-uncased</li>\n<li>Euclidean distance and cosine similarity to get nearest neighbors.</li>\n<li>Ensemble: 4 models * 3 output(concat, text, img) * 2 distances = 24 votes</li>\n<li>Different learning rate for cnn/bert/fc</li></ul></li>\n<li><p><a href=\"https://www.kaggle.com/c/shopee-product-matching/discussion/238174\" target=\"_blank\">7th Place Solution</a></p>\n<ul>\n<li><p>Image model</p>\n<ul>\n<li>Backbone : NFNet-F0 with GeM pooling (p=3(fixed) for training, p=4 for inference)</li>\n<li>Input size : 420x420</li>\n<li>Epochs : 30</li>\n<li>Optimizer : madgrad(momentum=0.9, weight_decay=1e-5) [1]</li>\n<li>Learning rate scheduler : cosine annealing (1e-5 --&gt; 1e-8)</li>\n<li>Distance : cosine similarity</li>\n<li>Loss : multi-similarity loss (alpha=2, beta=50, base=0.5) with XBM(memory_size=1024)</li>\n<li>Batch size : 32, mini-batch is generated by randomly sampling pairs of images in same label_group</li>\n<li>Data augmentation : Rotate, ColorJitter, RandomBrightnessContrast, RandomGamma, HorizontalFlip, CoarseDropout</li></ul></li>\n<li><p>Text model</p>\n<ul>\n<li>Backbone : DistilBERT (distilbert-base-indonesian)</li>\n<li>pochs : 30</li>\n<li>Optimizer : madgrad(momentum=0.9, weight_decay=1e-5)</li>\n<li>Learning rate scheduler : cosine annealing (1e-4 --&gt; 1e-8)</li>\n<li>Distance : cosine similarity</li>\n<li>Loss : multi-similarity loss (alpha=2, beta=50, base=0.5)</li>\n<li>Miner : multi-similarity miner (epsilon=0.1)</li>\n<li>Batch size : 256, mini-batch is generated by randomly sampling pairs of titles in same label_group</li>\n<li>Data augmentation : OneOf([random_delete, random_swap, random_swap+random_delete, random_swap*2]) at p=0.1<ul>\n<li>random_delete : randomly delete a word if len(title) &gt; 2</li>\n<li>random_swap : randomly swap two words if len(title) &gt; 2</li></ul></li></ul></li>\n<li><p>alpha query expansion</p></li></ul></li>\n<li><p><a href=\"https://www.kaggle.com/c/shopee-product-matching/discussion/238125\" target=\"_blank\">8th Place Solution</a></p>\n<ul>\n<li><p>Image Embedding</p>\n<ul>\n<li>ResNet152x2, ResNet101x3 with GeM Pooling (5fold)</li>\n<li>Loss: CosFace</li>\n<li>Optimizer SGD lr=1e-3 WarmupCosineAnnealing LR Scheduling</li>\n<li>Input size 512x512</li>\n<li>Embedding dimension 512(ResNet152), 768(ResNet101)</li></ul></li>\n<li><p>Text Embedding</p>\n<ul>\n<li>distilbert_base_indonesian (5fold)</li>\n<li>Concatenate mean of 4,5,6 Layer, CLS and mean of token embeddings (total 3840dim)</li>\n<li>Add FC and Tanh activation to reduce embedding dimension from 3840 to 1536</li>\n<li>Loss: ArcFace</li>\n<li>Optimizer: AdamW lr=1e-4 WarmupLinear LR Scheduling</li></ul></li>\n<li><p>Brute-force kNN by Faiss on embeddings converted to fp16. </p></li>\n<li><p>Use αQE + DBA: 0.59 Threshold for cosine similarity (local CV best threshold +0.15)</p></li></ul></li>\n<li><p><a href=\"https://www.kaggle.com/c/shopee-product-matching/discussion/238039\" target=\"_blank\">10th Place Solution</a></p>\n<ul>\n<li>Trained each models by ArcFace.</li>\n<li>Image: swin-base-224, effnet-b6, resnet101d  </li>\n<li>Text: indobenchmark/indobert-base-p2, indobenchmark/indobert-large-p2 and TFIDF.</li>\n<li>Query Expansion: αQE with n=2 and normalized similarity</li></ul></li>\n<li><p><a href=\"https://www.kaggle.com/c/shopee-product-matching/discussion/238181\" target=\"_blank\">11th Place Solution</a></p>\n<ul>\n<li><p>Image model</p>\n<ul>\n<li>backbone: swin_base_patch4_window12_384</li>\n<li>input_size : 384</li>\n<li>loss : ArcFace (scale=34, margine=0.5)</li>\n<li>fc layer (embedding) dimensions : 768</li>\n<li>optimizer : Ranger</li>\n<li>augmentations : RandAugment</li></ul></li>\n<li><p>BERT models</p>\n<ul>\n<li>backbone_model1 : sentence-transformers/paraphrase-xlm-r-multilingual-v1</li>\n<li>backbone_model2 : cahya/distilbert-base-indonesian</li>\n<li>loss : ArcFace (scale=30, margine=0.5)</li>\n<li>fc layer (embedding) dimensions : 768</li>\n<li>optimizer : SAM with AdamW</li></ul></li></ul></li>\n<li><p><a href=\"https://www.kaggle.com/c/shopee-product-matching/discussion/238033\" target=\"_blank\">14th Place Solution</a></p>\n<ul>\n<li>RAPIDS TfidfVectorizer and EfficientNetB0 384x384 images ensembled with XLM-RoBERTa which has multilingual pretraining and ensembled with EfficientNetB3 512x512 image.</li>\n<li>Make a decision boundary using piecewise linear functions</li>\n<li>Do five additional techniques to remove false negatives and false positives (definitely worth to study)</li></ul></li>\n<li><p><a href=\"https://www.kaggle.com/c/shopee-product-matching/discussion/238029\" target=\"_blank\">15th Place Solution</a></p>\n<ul>\n<li>Image:<ul>\n<li>resnest101, resnest200, eca_nfnet_l1</li>\n<li>Trained with Arcface and GeM. Input size is 512x512.</li>\n<li>Usual augmentation(LR flip, cutout, bright, shiftscalerotate, etc) by albumentations.</li>\n<li>512 dimension for each model, concatenate them to produce 1536 dimensional vector.</li></ul></li>\n<li>Text<ul>\n<li>TF-IDF: Fit and transform to test dataset.</li>\n<li>LaBSE: Finetuned with Arcface and GeM. Augmentation by EDA(Easy Data Augmentation)-like method.</li></ul></li>\n<li>Post-processing<ul>\n<li>Cross-mean embedding: If there are products(e.g. A,B and C) which have identical image, they are same product. (precision &gt; 99.9%). Replace each <em>title</em> embedding to mean of title embedding: (A+B+C)/3. Same procedure for image embedding, mean by same title.</li>\n<li>DBA/QE: For image X, got 3 nearest neighbor(e.g. X,Y,Z) by image embedding, then replace X's image embedding with weighted(logspace) sum of (X,Y,Z). Same for title.</li></ul></li></ul></li>\n</ul>",
      "rawMarkdown": "https://www.kaggle.com/c/shopee-product-matching/overview\n\n*There were Indonesian texts in this competition.*\n> While the modeling part is somewhat similar to each other, every solution brings something valuable to the table in terms of post-processing. So highly recommended to study them.\n\n- [1st Place Solution](https://www.kaggle.com/c/shopee-product-matching/discussion/238136)\n    -  Two eca_nfnet_l1s from timm for image encoders, and xlm-roberta-large, xlm-roberta-base, cahya/bert-base-indonesian-1.5G, indobenchmark/indobert-large-p1, bert-base-multilingual-uncased from huggingface for text encoders.\n    -  Used ArcFace to train the model. After pooling from image/text encoder, applied batchnormalization and feature-wise normalization to output embedding.\n        - increase margin gradually while training\n        - use large warmup steps\n        - use larger learning rate for cosinehead\n        - use gradient clipping\n    - Adding extra fc layers after global average pooling hurts the model performance, while adding batchnorm before feature-wise normalization improved the score.\n    - Normalize image embedding and text embedding then concat them to calculate comb-similarities.\n    - Iterative Neighborhood Blending (an approach to refine the embeddings which includes QE(Query Expansion) and DBA(DataBase-side feature Augmentation))\n    \n- [2nd Place Solution](https://www.kaggle.com/c/shopee-product-matching/discussion/238022)\n    - 1st stage\n        - Image\n            - Cosine similarities of NFNet-F0, ViT embeddings\n            - Loss: CurricularFace\n            - Optimizer: SAM\n            - Concatenate the similarities like F.normalize(torch.cat([F.normalize(emb1), F.normalize(emb2)], axis=1))\n        - Text\n            - Cosine similarities of Indonesian-BERT, Multilingual-BERT, and Paraphrase-XLM embeddings\n            - TF-IDF\n        - Image + Text\n            - Trained model with NFNet-F0 and Indonesian BERT (concatenated at final feature layers)\n    - 2nd stage: Train “meta” models to classify whether a pair of items belong to the same label group or not.\nUsed LightGBM and GAT (Graph Attention Networks)\n    - Graph-Based post-processing\n\n- [3rd Place Solution](https://www.kaggle.com/c/shopee-product-matching/discussion/238515)\n    - Build different (not trainable) products representations: efficient net embeddings, BERTs, TfIdf and search for nearest neighbours on fixed radius.\n    - For each object, union candidate neighbours from all embeddings and build a sample of pair objects with a binary target - if pair is a real duplicate or not. \n    - On that sample, build simple GradientBoosting model (catboost in my case) using next features, calculates separately for each embedding: pairwise distances (cosine, euclidean and etc), density around both points (frequency of points on different radiuses), points ranks.\n    -  After model is tranined, threshold candidate points by probability of duplicate, that is searched by CV.\n\n- [4th Place Solution](https://www.kaggle.com/c/shopee-product-matching/discussion/238295)\n    - Used Bert based Text Encoder and Image Encoders with different backbones (mainly nfnet, effnet)\n    - After the backbone, GeM and Avg pooling was used. The neck is comprised of a linear layer that is used for dimension tuning, followed by batch normalization and PReLu activation. \n    - Trained the image and text models with an ArcFaceLoss. Tuned the ArcFace margin.\n    - Concatenated and normalized image embeddings, text embeddings (bert based) and tfdif embeddings.\n    - On each of those vectors, calculated pair wise cosine similarity and received three matrices (cossim image, cossim bert, cossim tfidf). Combined those three matrices by first squareing them and then taking a weighted average.\n    - Postprocessing: \n        - Thresholding \n        - Rank2 matching (if A has B on rank2 and B has A on rank2, add them to each other) \n        - Rank2 and rank3 difference is large -> add rank2 id\n        - If there is a group of X members, and we have another row with the X preds on rank 1-X, make one large group.\n        - At least one other match (except the cosine similarity of rank2 is extremly low)\n        - Query Expansion \n        - Rematching unmatched rows\n\n\n- [5th Place Solution](https://www.kaggle.com/c/shopee-product-matching/discussion/238078)\n\n    - English Distilbert and Indonesian Distilbert models are trained for obtaining the title vectors. \n    - For the image vectors, ViT and Swin Transformers are trained on size 384 with relatively heavy augmentations and also EffNet B4 on size 512. \n    - Ensemble these models with vector concatenation.\n    - Weighted Database Augmentation\n    - Used cuML's NearestNeighbors for matching the vectors and used a threshold for filtering.\n    - Extracted  features for matched unique pairs and fed them to XGB model. The features are \"img_dist\", \"text_dist\", \"dist\", \"dist_rank\", \"cos_sim\", \"cos_sim2\". Basically, vector ditances, their ranking within each posting id, 2 different tfidf cosine similarity with different parameters.\n    - FP Features:  used closest match distances from training set as features.\n    - Agglomerative Clustering: After having the match probability predictions from XGB models, sorted all pairs by their match probabilities and started matching from the most likely match. Each match above 0.8 probability, merges their clusters\n\n- [6th Place Solution](https://www.kaggle.com/c/shopee-product-matching/discussion/238010)\n    - img_backbone: swin_large_patch4_window7_224, efficientnet_b3\n    - img_size: 224, 560\n    - text_backbone\txlm-roberta-base, bert-base-multilingual-uncased, bert-base-uncased\n    - Euclidean distance and cosine similarity to get nearest neighbors.\n    - Ensemble: 4 models * 3 output(concat, text, img) * 2 distances = 24 votes\n    - Different learning rate for cnn/bert/fc\n\n- [7th Place Solution](https://www.kaggle.com/c/shopee-product-matching/discussion/238174)\n\n    - Image model\n        - Backbone : NFNet-F0 with GeM pooling (p=3(fixed) for training, p=4 for inference)\n        - Input size : 420x420\n        - Epochs : 30\n        - Optimizer : madgrad(momentum=0.9, weight_decay=1e-5) [1]\n        - Learning rate scheduler : cosine annealing (1e-5 --> 1e-8)\n        - Distance : cosine similarity\n        - Loss : multi-similarity loss (alpha=2, beta=50, base=0.5) with XBM(memory_size=1024)\n        - Batch size : 32, mini-batch is generated by randomly sampling pairs of images in same label_group\n        - Data augmentation : Rotate, ColorJitter, RandomBrightnessContrast, RandomGamma, HorizontalFlip, CoarseDropout\n\n    - Text model\n        - Backbone : DistilBERT (distilbert-base-indonesian)\n        - pochs : 30\n        - Optimizer : madgrad(momentum=0.9, weight_decay=1e-5)\n        - Learning rate scheduler : cosine annealing (1e-4 --> 1e-8)\n        - Distance : cosine similarity\n        - Loss : multi-similarity loss (alpha=2, beta=50, base=0.5)\n        - Miner : multi-similarity miner (epsilon=0.1)\n        - Batch size : 256, mini-batch is generated by randomly sampling pairs of titles in same label_group\n        - Data augmentation : OneOf([random_delete, random_swap, random_swap+random_delete, random_swap*2]) at p=0.1\n            - random_delete : randomly delete a word if len(title) > 2\n            - random_swap : randomly swap two words if len(title) > 2\n    \n    - alpha query expansion\n\n- [8th Place Solution](https://www.kaggle.com/c/shopee-product-matching/discussion/238125)\n\n    - Image Embedding\n        - ResNet152x2, ResNet101x3 with GeM Pooling (5fold)\n        - Loss: CosFace\n        - Optimizer SGD lr=1e-3 WarmupCosineAnnealing LR Scheduling\n        - Input size 512x512\n        - Embedding dimension 512(ResNet152), 768(ResNet101)\n    - Text Embedding\n\n        - distilbert_base_indonesian (5fold)\n        - Concatenate mean of 4,5,6 Layer, CLS and mean of token embeddings (total 3840dim)\n        - Add FC and Tanh activation to reduce embedding dimension from 3840 to 1536\n        - Loss: ArcFace\n        - Optimizer: AdamW lr=1e-4 WarmupLinear LR Scheduling\n    \n    - Brute-force kNN by Faiss on embeddings converted to fp16. \n    - Use αQE + DBA: 0.59 Threshold for cosine similarity (local CV best threshold +0.15)\n    \n- [10th Place Solution](https://www.kaggle.com/c/shopee-product-matching/discussion/238039)\n\n    - Trained each models by ArcFace.\n    - Image: swin-base-224, effnet-b6, resnet101d  \n    - Text: indobenchmark/indobert-base-p2, indobenchmark/indobert-large-p2 and TFIDF.\n    - Query Expansion: αQE with n=2 and normalized similarity\n    \n- [11th Place Solution](https://www.kaggle.com/c/shopee-product-matching/discussion/238181)\n    - Image model\n        - backbone: swin_base_patch4_window12_384\n        - input_size : 384\n        - loss : ArcFace (scale=34, margine=0.5)\n        - fc layer (embedding) dimensions : 768\n        - optimizer : Ranger\n        - augmentations : RandAugment\n\n    - BERT models\n        - backbone_model1 : sentence-transformers/paraphrase-xlm-r-multilingual-v1\n        - backbone_model2 : cahya/distilbert-base-indonesian\n        - loss : ArcFace (scale=30, margine=0.5)\n        - fc layer (embedding) dimensions : 768\n        - optimizer : SAM with AdamW\n        \n- [14th Place Solution](https://www.kaggle.com/c/shopee-product-matching/discussion/238033)\n    - RAPIDS TfidfVectorizer and EfficientNetB0 384x384 images ensembled with XLM-RoBERTa which has multilingual pretraining and ensembled with EfficientNetB3 512x512 image.\n    - Make a decision boundary using piecewise linear functions\n    - Do five additional techniques to remove false negatives and false positives (definitely worth to study)\n\n- [15th Place Solution](https://www.kaggle.com/c/shopee-product-matching/discussion/238029)\n    - Image:\n        - resnest101, resnest200, eca_nfnet_l1\n        - Trained with Arcface and GeM. Input size is 512x512.\n        - Usual augmentation(LR flip, cutout, bright, shiftscalerotate, etc) by albumentations.\n        - 512 dimension for each model, concatenate them to produce 1536 dimensional vector.\n    - Text\n        - TF-IDF: Fit and transform to test dataset.\n        - LaBSE: Finetuned with Arcface and GeM. Augmentation by EDA(Easy Data Augmentation)-like method.\n    - Post-processing\n        - Cross-mean embedding: If there are products(e.g. A,B and C) which have identical image, they are same product. (precision > 99.9%). Replace each *title* embedding to mean of title embedding: (A+B+C)/3. Same procedure for image embedding, mean by same title.\n        - DBA/QE: For image X, got 3 nearest neighbor(e.g. X,Y,Z) by image embedding, then replace X's image embedding with weighted(logspace) sum of (X,Y,Z). Same for title.",
      "votes": null
    },
    {
      "id": "1716997",
      "postDate": "03/09/2022 15:15:09",
      "content": "<p>Fascinating !!  Thanks for sharing  💯</p>",
      "rawMarkdown": "Fascinating !!  Thanks for sharing  💯",
      "votes": null
    },
    {
      "id": "1717140",
      "postDate": "03/09/2022 17:22:09",
      "content": "<p>Glad you like it! Thanks!! </p>",
      "rawMarkdown": "Glad you like it! Thanks!!",
      "votes": null
    },
    {
      "id": "1718034",
      "postDate": "03/10/2022 13:03:39",
      "content": "<p>These are interesting ideas for a multi-modal model that could include product descriptions and images. I think they are missing the ranking part of this competition. Quoting from the Shopee competition:</p>\n<blockquote>\n  <p>In this competition, you’ll apply your machine learning skills to build a model that <strong>predicts which items are the same products</strong>.</p>\n</blockquote>",
      "rawMarkdown": "These are interesting ideas for a multi-modal model that could include product descriptions and images. I think they are missing the ranking part of this competition. Quoting from the Shopee competition:\n\n> In this competition, you’ll apply your machine learning skills to build a model that **predicts which items are the same products**.",
      "votes": null
    },
    {
      "id": "1721116",
      "postDate": "03/13/2022 12:28:29",
      "content": "<p>Yes, but the ideas to represent the images/texts would be directly applicable to this problem I believe. So I wanted to study &amp; share. </p>\n<p>Hope it can be useful. </p>",
      "rawMarkdown": "Yes, but the ideas to represent the images/texts would be directly applicable to this problem I believe. So I wanted to study & share. \n\nHope it can be useful.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1716997,
      "author_name": "emirkocak",
      "author_url": "",
      "post_date": "03/09/2022 15:15:09",
      "content": "<p>Fascinating !!  Thanks for sharing  💯</p>",
      "votes": null,
      "replies": [
        {
          "id": 1717140,
          "author_name": "snnclsr",
          "author_url": "",
          "post_date": "03/09/2022 17:22:09",
          "content": "<p>Glad you like it! Thanks!! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1718034,
      "author_name": "tbierhance",
      "author_url": "",
      "post_date": "03/10/2022 13:03:39",
      "content": "<p>These are interesting ideas for a multi-modal model that could include product descriptions and images. I think they are missing the ranking part of this competition. Quoting from the Shopee competition:</p>\n<blockquote>\n  <p>In this competition, you’ll apply your machine learning skills to build a model that <strong>predicts which items are the same products</strong>.</p>\n</blockquote>",
      "votes": null,
      "replies": [
        {
          "id": 1721116,
          "author_name": "snnclsr",
          "author_url": "",
          "post_date": "03/13/2022 12:28:29",
          "content": "<p>Yes, but the ideas to represent the images/texts would be directly applicable to this problem I believe. So I wanted to study &amp; share. </p>\n<p>Hope it can be useful. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1716970": "https://www.kaggle.com/c/shopee-product-matching/overview\n\n*There were Indonesian texts in this competition.*\n> While the modeling part is somewhat similar to each other, every solution brings something valuable to the table in terms of post-processing. So highly recommended to study them.\n\n- [1st Place Solution](https://www.kaggle.com/c/shopee-product-matching/discussion/238136)\n    -  Two eca_nfnet_l1s from timm for image encoders, and xlm-roberta-large, xlm-roberta-base, cahya/bert-base-indonesian-1.5G, indobenchmark/indobert-large-p1, bert-base-multilingual-uncased from huggingface for text encoders.\n    -  Used ArcFace to train the model. After pooling from image/text encoder, applied batchnormalization and feature-wise normalization to output embedding.\n        - increase margin gradually while training\n        - use large warmup steps\n        - use larger learning rate for cosinehead\n        - use gradient clipping\n    - Adding extra fc layers after global average pooling hurts the model performance, while adding batchnorm before feature-wise normalization improved the score.\n    - Normalize image embedding and text embedding then concat them to calculate comb-similarities.\n    - Iterative Neighborhood Blending (an approach to refine the embeddings which includes QE(Query Expansion) and DBA(DataBase-side feature Augmentation))\n    \n- [2nd Place Solution](https://www.kaggle.com/c/shopee-product-matching/discussion/238022)\n    - 1st stage\n        - Image\n            - Cosine similarities of NFNet-F0, ViT embeddings\n            - Loss: CurricularFace\n            - Optimizer: SAM\n            - Concatenate the similarities like F.normalize(torch.cat([F.normalize(emb1), F.normalize(emb2)], axis=1))\n        - Text\n            - Cosine similarities of Indonesian-BERT, Multilingual-BERT, and Paraphrase-XLM embeddings\n            - TF-IDF\n        - Image + Text\n            - Trained model with NFNet-F0 and Indonesian BERT (concatenated at final feature layers)\n    - 2nd stage: Train “meta” models to classify whether a pair of items belong to the same label group or not.\nUsed LightGBM and GAT (Graph Attention Networks)\n    - Graph-Based post-processing\n\n- [3rd Place Solution](https://www.kaggle.com/c/shopee-product-matching/discussion/238515)\n    - Build different (not trainable) products representations: efficient net embeddings, BERTs, TfIdf and search for nearest neighbours on fixed radius.\n    - For each object, union candidate neighbours from all embeddings and build a sample of pair objects with a binary target - if pair is a real duplicate or not. \n    - On that sample, build simple GradientBoosting model (catboost in my case) using next features, calculates separately for each embedding: pairwise distances (cosine, euclidean and etc), density around both points (frequency of points on different radiuses), points ranks.\n    -  After model is tranined, threshold candidate points by probability of duplicate, that is searched by CV.\n\n- [4th Place Solution](https://www.kaggle.com/c/shopee-product-matching/discussion/238295)\n    - Used Bert based Text Encoder and Image Encoders with different backbones (mainly nfnet, effnet)\n    - After the backbone, GeM and Avg pooling was used. The neck is comprised of a linear layer that is used for dimension tuning, followed by batch normalization and PReLu activation. \n    - Trained the image and text models with an ArcFaceLoss. Tuned the ArcFace margin.\n    - Concatenated and normalized image embeddings, text embeddings (bert based) and tfdif embeddings.\n    - On each of those vectors, calculated pair wise cosine similarity and received three matrices (cossim image, cossim bert, cossim tfidf). Combined those three matrices by first squareing them and then taking a weighted average.\n    - Postprocessing: \n        - Thresholding \n        - Rank2 matching (if A has B on rank2 and B has A on rank2, add them to each other) \n        - Rank2 and rank3 difference is large -> add rank2 id\n        - If there is a group of X members, and we have another row with the X preds on rank 1-X, make one large group.\n        - At least one other match (except the cosine similarity of rank2 is extremly low)\n        - Query Expansion \n        - Rematching unmatched rows\n\n\n- [5th Place Solution](https://www.kaggle.com/c/shopee-product-matching/discussion/238078)\n\n    - English Distilbert and Indonesian Distilbert models are trained for obtaining the title vectors. \n    - For the image vectors, ViT and Swin Transformers are trained on size 384 with relatively heavy augmentations and also EffNet B4 on size 512. \n    - Ensemble these models with vector concatenation.\n    - Weighted Database Augmentation\n    - Used cuML's NearestNeighbors for matching the vectors and used a threshold for filtering.\n    - Extracted  features for matched unique pairs and fed them to XGB model. The features are \"img_dist\", \"text_dist\", \"dist\", \"dist_rank\", \"cos_sim\", \"cos_sim2\". Basically, vector ditances, their ranking within each posting id, 2 different tfidf cosine similarity with different parameters.\n    - FP Features:  used closest match distances from training set as features.\n    - Agglomerative Clustering: After having the match probability predictions from XGB models, sorted all pairs by their match probabilities and started matching from the most likely match. Each match above 0.8 probability, merges their clusters\n\n- [6th Place Solution](https://www.kaggle.com/c/shopee-product-matching/discussion/238010)\n    - img_backbone: swin_large_patch4_window7_224, efficientnet_b3\n    - img_size: 224, 560\n    - text_backbone\txlm-roberta-base, bert-base-multilingual-uncased, bert-base-uncased\n    - Euclidean distance and cosine similarity to get nearest neighbors.\n    - Ensemble: 4 models * 3 output(concat, text, img) * 2 distances = 24 votes\n    - Different learning rate for cnn/bert/fc\n\n- [7th Place Solution](https://www.kaggle.com/c/shopee-product-matching/discussion/238174)\n\n    - Image model\n        - Backbone : NFNet-F0 with GeM pooling (p=3(fixed) for training, p=4 for inference)\n        - Input size : 420x420\n        - Epochs : 30\n        - Optimizer : madgrad(momentum=0.9, weight_decay=1e-5) [1]\n        - Learning rate scheduler : cosine annealing (1e-5 --> 1e-8)\n        - Distance : cosine similarity\n        - Loss : multi-similarity loss (alpha=2, beta=50, base=0.5) with XBM(memory_size=1024)\n        - Batch size : 32, mini-batch is generated by randomly sampling pairs of images in same label_group\n        - Data augmentation : Rotate, ColorJitter, RandomBrightnessContrast, RandomGamma, HorizontalFlip, CoarseDropout\n\n    - Text model\n        - Backbone : DistilBERT (distilbert-base-indonesian)\n        - pochs : 30\n        - Optimizer : madgrad(momentum=0.9, weight_decay=1e-5)\n        - Learning rate scheduler : cosine annealing (1e-4 --> 1e-8)\n        - Distance : cosine similarity\n        - Loss : multi-similarity loss (alpha=2, beta=50, base=0.5)\n        - Miner : multi-similarity miner (epsilon=0.1)\n        - Batch size : 256, mini-batch is generated by randomly sampling pairs of titles in same label_group\n        - Data augmentation : OneOf([random_delete, random_swap, random_swap+random_delete, random_swap*2]) at p=0.1\n            - random_delete : randomly delete a word if len(title) > 2\n            - random_swap : randomly swap two words if len(title) > 2\n    \n    - alpha query expansion\n\n- [8th Place Solution](https://www.kaggle.com/c/shopee-product-matching/discussion/238125)\n\n    - Image Embedding\n        - ResNet152x2, ResNet101x3 with GeM Pooling (5fold)\n        - Loss: CosFace\n        - Optimizer SGD lr=1e-3 WarmupCosineAnnealing LR Scheduling\n        - Input size 512x512\n        - Embedding dimension 512(ResNet152), 768(ResNet101)\n    - Text Embedding\n\n        - distilbert_base_indonesian (5fold)\n        - Concatenate mean of 4,5,6 Layer, CLS and mean of token embeddings (total 3840dim)\n        - Add FC and Tanh activation to reduce embedding dimension from 3840 to 1536\n        - Loss: ArcFace\n        - Optimizer: AdamW lr=1e-4 WarmupLinear LR Scheduling\n    \n    - Brute-force kNN by Faiss on embeddings converted to fp16. \n    - Use αQE + DBA: 0.59 Threshold for cosine similarity (local CV best threshold +0.15)\n    \n- [10th Place Solution](https://www.kaggle.com/c/shopee-product-matching/discussion/238039)\n\n    - Trained each models by ArcFace.\n    - Image: swin-base-224, effnet-b6, resnet101d  \n    - Text: indobenchmark/indobert-base-p2, indobenchmark/indobert-large-p2 and TFIDF.\n    - Query Expansion: αQE with n=2 and normalized similarity\n    \n- [11th Place Solution](https://www.kaggle.com/c/shopee-product-matching/discussion/238181)\n    - Image model\n        - backbone: swin_base_patch4_window12_384\n        - input_size : 384\n        - loss : ArcFace (scale=34, margine=0.5)\n        - fc layer (embedding) dimensions : 768\n        - optimizer : Ranger\n        - augmentations : RandAugment\n\n    - BERT models\n        - backbone_model1 : sentence-transformers/paraphrase-xlm-r-multilingual-v1\n        - backbone_model2 : cahya/distilbert-base-indonesian\n        - loss : ArcFace (scale=30, margine=0.5)\n        - fc layer (embedding) dimensions : 768\n        - optimizer : SAM with AdamW\n        \n- [14th Place Solution](https://www.kaggle.com/c/shopee-product-matching/discussion/238033)\n    - RAPIDS TfidfVectorizer and EfficientNetB0 384x384 images ensembled with XLM-RoBERTa which has multilingual pretraining and ensembled with EfficientNetB3 512x512 image.\n    - Make a decision boundary using piecewise linear functions\n    - Do five additional techniques to remove false negatives and false positives (definitely worth to study)\n\n- [15th Place Solution](https://www.kaggle.com/c/shopee-product-matching/discussion/238029)\n    - Image:\n        - resnest101, resnest200, eca_nfnet_l1\n        - Trained with Arcface and GeM. Input size is 512x512.\n        - Usual augmentation(LR flip, cutout, bright, shiftscalerotate, etc) by albumentations.\n        - 512 dimension for each model, concatenate them to produce 1536 dimensional vector.\n    - Text\n        - TF-IDF: Fit and transform to test dataset.\n        - LaBSE: Finetuned with Arcface and GeM. Augmentation by EDA(Easy Data Augmentation)-like method.\n    - Post-processing\n        - Cross-mean embedding: If there are products(e.g. A,B and C) which have identical image, they are same product. (precision > 99.9%). Replace each *title* embedding to mean of title embedding: (A+B+C)/3. Same procedure for image embedding, mean by same title.\n        - DBA/QE: For image X, got 3 nearest neighbor(e.g. X,Y,Z) by image embedding, then replace X's image embedding with weighted(logspace) sum of (X,Y,Z). Same for title.",
    "1716997": "Fascinating !!  Thanks for sharing  💯",
    "1717140": "Glad you like it! Thanks!!",
    "1718034": "These are interesting ideas for a multi-modal model that could include product descriptions and images. I think they are missing the ranking part of this competition. Quoting from the Shopee competition:\n\n> In this competition, you’ll apply your machine learning skills to build a model that **predicts which items are the same products**.",
    "1721116": "Yes, but the ideas to represent the images/texts would be directly applicable to this problem I believe. So I wanted to study & share. \n\nHope it can be useful."
  },
  "source": "meta"
}