{
  "id": 681465,
  "title": "4th Place Solution WriteUp",
  "url": "/competitions/jaguar-re-id/writeups/4th-place-solution-writeup",
  "author_name": "",
  "post_date": "2026-03-15T08:22:43.033Z",
  "votes": 6,
  "comment_count": 1,
  "views": 0,
  "content": "<h1>Solution Writeup — 4th Place, Jaguar Re-ID (0.939 private LB)</h1>\n<p>First of all, thank you for hosting this amazing competition, it's been a lot of fun ! I did not expect such a shake up at the end, but it's for once a very welcome one :)</p>\n<h2>Quick summary</h2>\n<p><strong>4th place / 348 teams</strong> — private LB 0.939.</p>\n<p>The core idea: train two EVA-02 Large models with different random seeds + one DINOv2-Large model, apply pseudo-labeling to expand the training set with test images, then ensemble all three with query expansion (QE) and k-reciprocal reranking at inference and no external data.</p>\n<p>All notebook versions (v01 through v07g) are public on the competition page.</p>\n<hr>\n<h2>Trials and tribulations</h2>\n<p><strong>v01 — Frozen EfficientNet-B0 (LB ~0.28)</strong>: The floor. ImageNet features + cosine similarity, no training. Useful mainly to confirm there is signal in pretrained features — jaguars do look different enough that off-the-shelf networks can partially distinguish them.</p>\n<p><strong>v02 — Fine-tuned 31-class classifier (LB 0.414)</strong>: Standard image classification. The embedding space isn't explicitly trained for retrieval, but fine-tuning on the actual 31 identities helped a lot. +73% over the frozen baseline.</p>\n<p><strong>v03 — ArcFace metric learning (LB 0.609)</strong>: This was the first real jump. ArcFace directly optimises the embedding space for re-identification by adding a margin penalty on the angle between an embedding and its class centre. Switching from CrossEntropy to ArcFace: +47%.</p>\n<p><strong>v04 — Backbone search (LB 0.797)</strong>: Compared tf_efficientnetv2_s, convnext_small, vit_small. ConvNeXt won cleanly. Added image caching + AMP to fit 3 backbones in under 25 minutes total.</p>\n<p><strong>v05 — Heavy augmentation + TTA (LB 0.795)</strong>: This one hurt. Heavy aug (random rotation, aggressive colour jitter, random erasing) + 5-view test-time augmentation. Regression on LB. With only 1,895 training images, strong augmentation over-regularises — the model becomes too robust to transformations it shouldn't be robust to.</p>\n<p><strong>v06 — Hybrid LightGBM (LB 0.799)</strong>: Built 571-dim pairwise features (cosine, L2, |emb_i − emb_j|, colour histogram diffs) and trained LightGBM on top of ArcFace embeddings. The LightGBM val mAP (0.491) was below raw cosine (0.510). ArcFace already encodes the right similarity — adding a learned layer on top adds noise.</p>\n<p><strong>v07a/b — EVA-02 Large (LB 0.918 / 0.915)</strong>: This was the big leap. Moved from ConvNeXt-Small to EVA-02 Large (300M parameters, patch14 at 448px). The private LB went from 0.799 to 0.918 — a +15% jump in a single step. Two seeds (42 and 123) to enable ensemble diversity later.</p>\n<p><strong>v07f — DINOv2-Large-reg4 (LB 0.886)</strong>: Added DINOv2 as a third model. The reasoning: EVA-02 is supervised (trained on ImageNet-21k labels), DINOv2 is self-supervised (DINO+iBOT on 142M curated images). Different pretraining paradigms should produce genuinely different feature representations, making them useful ensemble members even when all three are the same architecture class.</p>\n<p><strong>v07d — Pseudo-labeling (LB 0.935)</strong>: Used the three trained models to generate pseudo-labels for test images: only images where all three models agreed (confidence &gt; 0.90) were added to training. Then fine-tuned EVA-02a for 5 more epochs at lr=1e-5. ~100-150 pseudo-labels added.</p>\n<p><strong>v07e — Final ensemble (LB 0.939)</strong>: Simple 3-way average of EVA-A + EVA-B + DINOv2 embeddings (all 1024-dim), followed by query expansion and k-reciprocal reranking. This is the selected submission.</p>\n<p><strong>v07g — Round-2 pseudo-labeling (LB 0.937)</strong>: Tried a second round of pseudo-labeling on the already-fine-tuned model with a more aggressive strategy (single-model confidence using ArcFace head weights, ~300+ labels). Small regression. Stacking two PL rounds adds noise rather than signal — the first round already captured most of the easy assignments.</p>\n<hr>\n<h2>The best notebook in detail (v07e)</h2>\n<p>The final pipeline has five distinct stages. Here's what each does and why.</p>\n<h3>1. Backbone: EVA-02 Large @ 448px</h3>\n<p>EVA-02 is a ViT-based model (patch size 14, 448×448 input → 32×32 = 1024 patch tokens) pretrained with masked image modelling on ImageNet-21k then fine-tuned on ImageNet-1k. The key property for re-ID is the resolution: at 448px with 14px patches, each token covers a 14×14px region of the image. Jaguar spot patterns have detail at roughly that scale, so the model's attention can resolve individual spots rather than blurred regions.</p>\n<p>ConvNeXt at 224px with its effective receptive fields was already working well, but EVA-02 at 448px sees nearly 4× more pixel area per decision. That's the main reason for the +15% LB jump.</p>\n<p>Two seeds (42 and 123) are trained separately. They converge to slightly different regions of the embedding space — enough to reduce ensemble variance when averaged later.</p>\n<h3>2. Embedding head: GeM + BN + ArcFace</h3>\n<p><strong>GeM pooling</strong> (Generalised Mean Pooling, learnable p≈3): takes the 1024 patch tokens from the ViT, reshapes them into a 32×32 spatial grid, and pools with a learnable power mean. When p=1 it's average pooling; as p increases it increasingly emphasises the highest-activating (most discriminative) spatial locations. For re-ID this matters — you want the spot pattern regions to dominate the embedding, not the background.</p>\n<p><strong>Batch normalisation</strong> on the embedding before ArcFace stabilises training and makes the embedding norm uniform across samples, which is important for cosine similarity to be meaningful.</p>\n<p><strong>ArcFace</strong> (s=30, m=0.5): the angular margin loss. For each batch, it adds a penalty of 0.5 radians (≈29°) to the angle between each embedding and its correct class centre. This forces the model to place embeddings closer to their class centre than any other centre — directly optimising the geometry used at inference.</p>\n<h3>3. Training: EMA + gradient checkpointing</h3>\n<p><strong>EMA</strong> (exponential moving average, decay=0.999): maintains a shadow copy of the model weights that is a running average over the last ~1000 gradient steps. The shadow model is used at inference. This smooths out late-training oscillations and consistently produces better retrieval performance than the final raw checkpoint.</p>\n<p><strong>Gradient checkpointing</strong>: at 448px with a 300M parameter ViT and batch size 4, we're near the P100's VRAM limit. Gradient checkpointing recomputes activations during the backward pass rather than storing them, trading ~30% compute overhead for ~40% VRAM savings. Without it, training would require batch size 1-2 which hurts convergence.</p>\n<p><strong>20 epochs matters</strong>: v07a at 10 epochs scored 0.849. At 20 epochs it scored 0.918. The +0.069 gap is because EVA-02's large capacity takes time to converge on a 1,895-image dataset — it's learning slowly but surely.</p>\n<h3>4. Pseudo-labeling</h3>\n<p>After training all three models, we use them to label the 371 test images. Only images where all three models predicted the same identity with confidence &gt; 0.90 were accepted as pseudo-labels. This 3-way agreement filter is conservative (~100-150 samples) but high precision — we're adding mostly correct labels.</p>\n<p>The fine-tuning step (5 epochs at lr=1e-5 = 10× smaller than the original lr) is deliberately gentle. The goal is to reinforce known patterns with the new test-domain samples, not to overfit to potentially noisy pseudo-labels.</p>\n<h3>5. Ensemble + post-processing</h3>\n<p><strong>3-way average</strong>: simple arithmetic mean of the L2-normalised embeddings from EVA-A, EVA-B, and DINOv2 (all 1024-dim). Then re-normalise. This is the right fusion strategy for same-dimension embeddings: it's equivalent to taking the centroid in the embedding space. We tried concatenation (EVA 1024d + ConvNeXt 1536d) earlier and it regressed results — concatenation puts architecturally different features in the same vector without any calibration, the distance metric becomes meaningless.</p>\n<p><strong>Query Expansion</strong> (top-3 DBA): for each image, replace its embedding with the average of itself and its 3 nearest neighbours in the gallery. This uses the assumption that neighbours in embedding space are likely the same identity — a reasonable assumption for well-trained ArcFace models.</p>\n<p><strong>K-reciprocal reranking</strong> (k1=20, λ=0.3): the Zhong et al. 2017 reranking algorithm. The idea: if image A is in image B's top-k neighbourhood AND B is in A's top-k neighbourhood, they are \"k-reciprocal nearest neighbours\" — very likely the same identity. It computes a Jaccard similarity over these reciprocal sets and blends it with the original cosine similarity (λ controls the blend). This consistently boosted scores by 0.01-0.02 throughout the competition.</p>\n<hr>\n<h2>What didn't work</h2>\n<ul>\n<li><strong>Heavy augmentation</strong>: over-regularises at 1,895 images. Jaguars are already diverse enough.</li>\n<li><strong>LightGBM on pairwise features</strong>: ArcFace geometry is already optimal for cosine; adding a nonlinear layer on top hurts.</li>\n<li><strong>Cross-architecture concat</strong>: mixing 1024-dim and 1536-dim embeddings. The Euclidean distances become dominated by the larger-norm component.</li>\n<li><strong>Round-2 pseudo-labeling</strong>: stacking PL rounds. Second round adds noise the first round already filtered.</li>\n</ul>\n<h2>- ** Convnext underperformed compared to ViTs</h2>\n<h2>Project</h2>\n<p>All training and inference code is on GitHub: <a href=\"https://github.com/Smooth-Cactus0/jaguar-re-identification\" target=\"_blank\">https://github.com/Smooth-Cactus0/jaguar-re-identification</a>. The repo includes all notebook versions (v01–v07g), EDA notebooks, reusable <code>src/</code> modules, and benchmark results.</p>\n<p>Most of the notebooks in this project were built with the help of Claude Sonnet 4.6</p>\n<p>— Alexy </p>",
  "messages": [
    {
      "id": "3421324",
      "postDate": "03/15/2026 08:22:29",
      "content": "<h1>Solution Writeup — 4th Place, Jaguar Re-ID (0.939 private LB)</h1>\n<p>First of all, thank you for hosting this amazing competition, it's been a lot of fun ! I did not expect such a shake up at the end, but it's for once a very welcome one :)</p>\n<h2>Quick summary</h2>\n<p><strong>4th place / 348 teams</strong> — private LB 0.939.</p>\n<p>The core idea: train two EVA-02 Large models with different random seeds + one DINOv2-Large model, apply pseudo-labeling to expand the training set with test images, then ensemble all three with query expansion (QE) and k-reciprocal reranking at inference and no external data.</p>\n<p>All notebook versions (v01 through v07g) are public on the competition page.</p>\n<hr>\n<h2>Trials and tribulations</h2>\n<p><strong>v01 — Frozen EfficientNet-B0 (LB ~0.28)</strong>: The floor. ImageNet features + cosine similarity, no training. Useful mainly to confirm there is signal in pretrained features — jaguars do look different enough that off-the-shelf networks can partially distinguish them.</p>\n<p><strong>v02 — Fine-tuned 31-class classifier (LB 0.414)</strong>: Standard image classification. The embedding space isn't explicitly trained for retrieval, but fine-tuning on the actual 31 identities helped a lot. +73% over the frozen baseline.</p>\n<p><strong>v03 — ArcFace metric learning (LB 0.609)</strong>: This was the first real jump. ArcFace directly optimises the embedding space for re-identification by adding a margin penalty on the angle between an embedding and its class centre. Switching from CrossEntropy to ArcFace: +47%.</p>\n<p><strong>v04 — Backbone search (LB 0.797)</strong>: Compared tf_efficientnetv2_s, convnext_small, vit_small. ConvNeXt won cleanly. Added image caching + AMP to fit 3 backbones in under 25 minutes total.</p>\n<p><strong>v05 — Heavy augmentation + TTA (LB 0.795)</strong>: This one hurt. Heavy aug (random rotation, aggressive colour jitter, random erasing) + 5-view test-time augmentation. Regression on LB. With only 1,895 training images, strong augmentation over-regularises — the model becomes too robust to transformations it shouldn't be robust to.</p>\n<p><strong>v06 — Hybrid LightGBM (LB 0.799)</strong>: Built 571-dim pairwise features (cosine, L2, |emb_i − emb_j|, colour histogram diffs) and trained LightGBM on top of ArcFace embeddings. The LightGBM val mAP (0.491) was below raw cosine (0.510). ArcFace already encodes the right similarity — adding a learned layer on top adds noise.</p>\n<p><strong>v07a/b — EVA-02 Large (LB 0.918 / 0.915)</strong>: This was the big leap. Moved from ConvNeXt-Small to EVA-02 Large (300M parameters, patch14 at 448px). The private LB went from 0.799 to 0.918 — a +15% jump in a single step. Two seeds (42 and 123) to enable ensemble diversity later.</p>\n<p><strong>v07f — DINOv2-Large-reg4 (LB 0.886)</strong>: Added DINOv2 as a third model. The reasoning: EVA-02 is supervised (trained on ImageNet-21k labels), DINOv2 is self-supervised (DINO+iBOT on 142M curated images). Different pretraining paradigms should produce genuinely different feature representations, making them useful ensemble members even when all three are the same architecture class.</p>\n<p><strong>v07d — Pseudo-labeling (LB 0.935)</strong>: Used the three trained models to generate pseudo-labels for test images: only images where all three models agreed (confidence &gt; 0.90) were added to training. Then fine-tuned EVA-02a for 5 more epochs at lr=1e-5. ~100-150 pseudo-labels added.</p>\n<p><strong>v07e — Final ensemble (LB 0.939)</strong>: Simple 3-way average of EVA-A + EVA-B + DINOv2 embeddings (all 1024-dim), followed by query expansion and k-reciprocal reranking. This is the selected submission.</p>\n<p><strong>v07g — Round-2 pseudo-labeling (LB 0.937)</strong>: Tried a second round of pseudo-labeling on the already-fine-tuned model with a more aggressive strategy (single-model confidence using ArcFace head weights, ~300+ labels). Small regression. Stacking two PL rounds adds noise rather than signal — the first round already captured most of the easy assignments.</p>\n<hr>\n<h2>The best notebook in detail (v07e)</h2>\n<p>The final pipeline has five distinct stages. Here's what each does and why.</p>\n<h3>1. Backbone: EVA-02 Large @ 448px</h3>\n<p>EVA-02 is a ViT-based model (patch size 14, 448×448 input → 32×32 = 1024 patch tokens) pretrained with masked image modelling on ImageNet-21k then fine-tuned on ImageNet-1k. The key property for re-ID is the resolution: at 448px with 14px patches, each token covers a 14×14px region of the image. Jaguar spot patterns have detail at roughly that scale, so the model's attention can resolve individual spots rather than blurred regions.</p>\n<p>ConvNeXt at 224px with its effective receptive fields was already working well, but EVA-02 at 448px sees nearly 4× more pixel area per decision. That's the main reason for the +15% LB jump.</p>\n<p>Two seeds (42 and 123) are trained separately. They converge to slightly different regions of the embedding space — enough to reduce ensemble variance when averaged later.</p>\n<h3>2. Embedding head: GeM + BN + ArcFace</h3>\n<p><strong>GeM pooling</strong> (Generalised Mean Pooling, learnable p≈3): takes the 1024 patch tokens from the ViT, reshapes them into a 32×32 spatial grid, and pools with a learnable power mean. When p=1 it's average pooling; as p increases it increasingly emphasises the highest-activating (most discriminative) spatial locations. For re-ID this matters — you want the spot pattern regions to dominate the embedding, not the background.</p>\n<p><strong>Batch normalisation</strong> on the embedding before ArcFace stabilises training and makes the embedding norm uniform across samples, which is important for cosine similarity to be meaningful.</p>\n<p><strong>ArcFace</strong> (s=30, m=0.5): the angular margin loss. For each batch, it adds a penalty of 0.5 radians (≈29°) to the angle between each embedding and its correct class centre. This forces the model to place embeddings closer to their class centre than any other centre — directly optimising the geometry used at inference.</p>\n<h3>3. Training: EMA + gradient checkpointing</h3>\n<p><strong>EMA</strong> (exponential moving average, decay=0.999): maintains a shadow copy of the model weights that is a running average over the last ~1000 gradient steps. The shadow model is used at inference. This smooths out late-training oscillations and consistently produces better retrieval performance than the final raw checkpoint.</p>\n<p><strong>Gradient checkpointing</strong>: at 448px with a 300M parameter ViT and batch size 4, we're near the P100's VRAM limit. Gradient checkpointing recomputes activations during the backward pass rather than storing them, trading ~30% compute overhead for ~40% VRAM savings. Without it, training would require batch size 1-2 which hurts convergence.</p>\n<p><strong>20 epochs matters</strong>: v07a at 10 epochs scored 0.849. At 20 epochs it scored 0.918. The +0.069 gap is because EVA-02's large capacity takes time to converge on a 1,895-image dataset — it's learning slowly but surely.</p>\n<h3>4. Pseudo-labeling</h3>\n<p>After training all three models, we use them to label the 371 test images. Only images where all three models predicted the same identity with confidence &gt; 0.90 were accepted as pseudo-labels. This 3-way agreement filter is conservative (~100-150 samples) but high precision — we're adding mostly correct labels.</p>\n<p>The fine-tuning step (5 epochs at lr=1e-5 = 10× smaller than the original lr) is deliberately gentle. The goal is to reinforce known patterns with the new test-domain samples, not to overfit to potentially noisy pseudo-labels.</p>\n<h3>5. Ensemble + post-processing</h3>\n<p><strong>3-way average</strong>: simple arithmetic mean of the L2-normalised embeddings from EVA-A, EVA-B, and DINOv2 (all 1024-dim). Then re-normalise. This is the right fusion strategy for same-dimension embeddings: it's equivalent to taking the centroid in the embedding space. We tried concatenation (EVA 1024d + ConvNeXt 1536d) earlier and it regressed results — concatenation puts architecturally different features in the same vector without any calibration, the distance metric becomes meaningless.</p>\n<p><strong>Query Expansion</strong> (top-3 DBA): for each image, replace its embedding with the average of itself and its 3 nearest neighbours in the gallery. This uses the assumption that neighbours in embedding space are likely the same identity — a reasonable assumption for well-trained ArcFace models.</p>\n<p><strong>K-reciprocal reranking</strong> (k1=20, λ=0.3): the Zhong et al. 2017 reranking algorithm. The idea: if image A is in image B's top-k neighbourhood AND B is in A's top-k neighbourhood, they are \"k-reciprocal nearest neighbours\" — very likely the same identity. It computes a Jaccard similarity over these reciprocal sets and blends it with the original cosine similarity (λ controls the blend). This consistently boosted scores by 0.01-0.02 throughout the competition.</p>\n<hr>\n<h2>What didn't work</h2>\n<ul>\n<li><strong>Heavy augmentation</strong>: over-regularises at 1,895 images. Jaguars are already diverse enough.</li>\n<li><strong>LightGBM on pairwise features</strong>: ArcFace geometry is already optimal for cosine; adding a nonlinear layer on top hurts.</li>\n<li><strong>Cross-architecture concat</strong>: mixing 1024-dim and 1536-dim embeddings. The Euclidean distances become dominated by the larger-norm component.</li>\n<li><strong>Round-2 pseudo-labeling</strong>: stacking PL rounds. Second round adds noise the first round already filtered.</li>\n</ul>\n<h2>- ** Convnext underperformed compared to ViTs</h2>\n<h2>Project</h2>\n<p>All training and inference code is on GitHub: <a href=\"https://github.com/Smooth-Cactus0/jaguar-re-identification\" target=\"_blank\">https://github.com/Smooth-Cactus0/jaguar-re-identification</a>. The repo includes all notebook versions (v01–v07g), EDA notebooks, reusable <code>src/</code> modules, and benchmark results.</p>\n<p>Most of the notebooks in this project were built with the help of Claude Sonnet 4.6</p>\n<p>— Alexy </p>",
      "rawMarkdown": "# Solution Writeup — 4th Place, Jaguar Re-ID (0.939 private LB)\n\nFirst of all, thank you for hosting this amazing competition, it's been a lot of fun ! I did not expect such a shake up at the end, but it's for once a very welcome one :)\n\n\n## Quick summary\n\n**4th place / 348 teams** — private LB 0.939.\n\nThe core idea: train two EVA-02 Large models with different random seeds + one DINOv2-Large model, apply pseudo-labeling to expand the training set with test images, then ensemble all three with query expansion (QE) and k-reciprocal reranking at inference and no external data.\n\nAll notebook versions (v01 through v07g) are public on the competition page.\n\n---\n\n## Trials and tribulations\n\n\n\n**v01 — Frozen EfficientNet-B0 (LB ~0.28)**: The floor. ImageNet features + cosine similarity, no training. Useful mainly to confirm there is signal in pretrained features — jaguars do look different enough that off-the-shelf networks can partially distinguish them.\n\n**v02 — Fine-tuned 31-class classifier (LB 0.414)**: Standard image classification. The embedding space isn't explicitly trained for retrieval, but fine-tuning on the actual 31 identities helped a lot. +73% over the frozen baseline.\n\n**v03 — ArcFace metric learning (LB 0.609)**: This was the first real jump. ArcFace directly optimises the embedding space for re-identification by adding a margin penalty on the angle between an embedding and its class centre. Switching from CrossEntropy to ArcFace: +47%.\n\n**v04 — Backbone search (LB 0.797)**: Compared tf_efficientnetv2_s, convnext_small, vit_small. ConvNeXt won cleanly. Added image caching + AMP to fit 3 backbones in under 25 minutes total.\n\n**v05 — Heavy augmentation + TTA (LB 0.795)**: This one hurt. Heavy aug (random rotation, aggressive colour jitter, random erasing) + 5-view test-time augmentation. Regression on LB. With only 1,895 training images, strong augmentation over-regularises — the model becomes too robust to transformations it shouldn't be robust to.\n\n**v06 — Hybrid LightGBM (LB 0.799)**: Built 571-dim pairwise features (cosine, L2, |emb_i − emb_j|, colour histogram diffs) and trained LightGBM on top of ArcFace embeddings. The LightGBM val mAP (0.491) was below raw cosine (0.510). ArcFace already encodes the right similarity — adding a learned layer on top adds noise.\n\n**v07a/b — EVA-02 Large (LB 0.918 / 0.915)**: This was the big leap. Moved from ConvNeXt-Small to EVA-02 Large (300M parameters, patch14 at 448px). The private LB went from 0.799 to 0.918 — a +15% jump in a single step. Two seeds (42 and 123) to enable ensemble diversity later.\n\n**v07f — DINOv2-Large-reg4 (LB 0.886)**: Added DINOv2 as a third model. The reasoning: EVA-02 is supervised (trained on ImageNet-21k labels), DINOv2 is self-supervised (DINO+iBOT on 142M curated images). Different pretraining paradigms should produce genuinely different feature representations, making them useful ensemble members even when all three are the same architecture class.\n\n**v07d — Pseudo-labeling (LB 0.935)**: Used the three trained models to generate pseudo-labels for test images: only images where all three models agreed (confidence > 0.90) were added to training. Then fine-tuned EVA-02a for 5 more epochs at lr=1e-5. ~100-150 pseudo-labels added.\n\n**v07e — Final ensemble (LB 0.939)**: Simple 3-way average of EVA-A + EVA-B + DINOv2 embeddings (all 1024-dim), followed by query expansion and k-reciprocal reranking. This is the selected submission.\n\n**v07g — Round-2 pseudo-labeling (LB 0.937)**: Tried a second round of pseudo-labeling on the already-fine-tuned model with a more aggressive strategy (single-model confidence using ArcFace head weights, ~300+ labels). Small regression. Stacking two PL rounds adds noise rather than signal — the first round already captured most of the easy assignments.\n\n---\n\n## The best notebook in detail (v07e)\n\nThe final pipeline has five distinct stages. Here's what each does and why.\n\n### 1. Backbone: EVA-02 Large @ 448px\n\nEVA-02 is a ViT-based model (patch size 14, 448×448 input → 32×32 = 1024 patch tokens) pretrained with masked image modelling on ImageNet-21k then fine-tuned on ImageNet-1k. The key property for re-ID is the resolution: at 448px with 14px patches, each token covers a 14×14px region of the image. Jaguar spot patterns have detail at roughly that scale, so the model's attention can resolve individual spots rather than blurred regions.\n\nConvNeXt at 224px with its effective receptive fields was already working well, but EVA-02 at 448px sees nearly 4× more pixel area per decision. That's the main reason for the +15% LB jump.\n\nTwo seeds (42 and 123) are trained separately. They converge to slightly different regions of the embedding space — enough to reduce ensemble variance when averaged later.\n\n### 2. Embedding head: GeM + BN + ArcFace\n\n**GeM pooling** (Generalised Mean Pooling, learnable p≈3): takes the 1024 patch tokens from the ViT, reshapes them into a 32×32 spatial grid, and pools with a learnable power mean. When p=1 it's average pooling; as p increases it increasingly emphasises the highest-activating (most discriminative) spatial locations. For re-ID this matters — you want the spot pattern regions to dominate the embedding, not the background.\n\n**Batch normalisation** on the embedding before ArcFace stabilises training and makes the embedding norm uniform across samples, which is important for cosine similarity to be meaningful.\n\n**ArcFace** (s=30, m=0.5): the angular margin loss. For each batch, it adds a penalty of 0.5 radians (≈29°) to the angle between each embedding and its correct class centre. This forces the model to place embeddings closer to their class centre than any other centre — directly optimising the geometry used at inference.\n\n### 3. Training: EMA + gradient checkpointing\n\n**EMA** (exponential moving average, decay=0.999): maintains a shadow copy of the model weights that is a running average over the last ~1000 gradient steps. The shadow model is used at inference. This smooths out late-training oscillations and consistently produces better retrieval performance than the final raw checkpoint.\n\n**Gradient checkpointing**: at 448px with a 300M parameter ViT and batch size 4, we're near the P100's VRAM limit. Gradient checkpointing recomputes activations during the backward pass rather than storing them, trading ~30% compute overhead for ~40% VRAM savings. Without it, training would require batch size 1-2 which hurts convergence.\n\n**20 epochs matters**: v07a at 10 epochs scored 0.849. At 20 epochs it scored 0.918. The +0.069 gap is because EVA-02's large capacity takes time to converge on a 1,895-image dataset — it's learning slowly but surely.\n\n### 4. Pseudo-labeling\n\nAfter training all three models, we use them to label the 371 test images. Only images where all three models predicted the same identity with confidence > 0.90 were accepted as pseudo-labels. This 3-way agreement filter is conservative (~100-150 samples) but high precision — we're adding mostly correct labels.\n\nThe fine-tuning step (5 epochs at lr=1e-5 = 10× smaller than the original lr) is deliberately gentle. The goal is to reinforce known patterns with the new test-domain samples, not to overfit to potentially noisy pseudo-labels.\n\n### 5. Ensemble + post-processing\n\n**3-way average**: simple arithmetic mean of the L2-normalised embeddings from EVA-A, EVA-B, and DINOv2 (all 1024-dim). Then re-normalise. This is the right fusion strategy for same-dimension embeddings: it's equivalent to taking the centroid in the embedding space. We tried concatenation (EVA 1024d + ConvNeXt 1536d) earlier and it regressed results — concatenation puts architecturally different features in the same vector without any calibration, the distance metric becomes meaningless.\n\n**Query Expansion** (top-3 DBA): for each image, replace its embedding with the average of itself and its 3 nearest neighbours in the gallery. This uses the assumption that neighbours in embedding space are likely the same identity — a reasonable assumption for well-trained ArcFace models.\n\n**K-reciprocal reranking** (k1=20, λ=0.3): the Zhong et al. 2017 reranking algorithm. The idea: if image A is in image B's top-k neighbourhood AND B is in A's top-k neighbourhood, they are \"k-reciprocal nearest neighbours\" — very likely the same identity. It computes a Jaccard similarity over these reciprocal sets and blends it with the original cosine similarity (λ controls the blend). This consistently boosted scores by 0.01-0.02 throughout the competition.\n\n---\n\n## What didn't work\n\n- **Heavy augmentation**: over-regularises at 1,895 images. Jaguars are already diverse enough.\n- **LightGBM on pairwise features**: ArcFace geometry is already optimal for cosine; adding a nonlinear layer on top hurts.\n- **Cross-architecture concat**: mixing 1024-dim and 1536-dim embeddings. The Euclidean distances become dominated by the larger-norm component.\n- **Round-2 pseudo-labeling**: stacking PL rounds. Second round adds noise the first round already filtered.\n- ** Convnext underperformed compared to ViTs\n---\n\n## Project\n\nAll training and inference code is on GitHub: https://github.com/Smooth-Cactus0/jaguar-re-identification. The repo includes all notebook versions (v01–v07g), EDA notebooks, reusable `src/` modules, and benchmark results.\n\nMost of the notebooks in this project were built with the help of Claude Sonnet 4.6\n\n— Alexy",
      "votes": null
    },
    {
      "id": "3421400",
      "postDate": "03/15/2026 12:52:54",
      "content": "<p>Great writeup! Thanks for sharing your journey, <a href=\"https://www.kaggle.com/alexycactus\" target=\"_blank\">@alexycactus</a> </p>",
      "rawMarkdown": "Great writeup! Thanks for sharing your journey, @alexycactus",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3421400,
      "author_name": "andandand",
      "author_url": "",
      "post_date": "03/15/2026 12:52:54",
      "content": "<p>Great writeup! Thanks for sharing your journey, <a href=\"https://www.kaggle.com/alexycactus\" target=\"_blank\">@alexycactus</a> </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3421324": "# Solution Writeup — 4th Place, Jaguar Re-ID (0.939 private LB)\n\nFirst of all, thank you for hosting this amazing competition, it's been a lot of fun ! I did not expect such a shake up at the end, but it's for once a very welcome one :)\n\n\n## Quick summary\n\n**4th place / 348 teams** — private LB 0.939.\n\nThe core idea: train two EVA-02 Large models with different random seeds + one DINOv2-Large model, apply pseudo-labeling to expand the training set with test images, then ensemble all three with query expansion (QE) and k-reciprocal reranking at inference and no external data.\n\nAll notebook versions (v01 through v07g) are public on the competition page.\n\n---\n\n## Trials and tribulations\n\n\n\n**v01 — Frozen EfficientNet-B0 (LB ~0.28)**: The floor. ImageNet features + cosine similarity, no training. Useful mainly to confirm there is signal in pretrained features — jaguars do look different enough that off-the-shelf networks can partially distinguish them.\n\n**v02 — Fine-tuned 31-class classifier (LB 0.414)**: Standard image classification. The embedding space isn't explicitly trained for retrieval, but fine-tuning on the actual 31 identities helped a lot. +73% over the frozen baseline.\n\n**v03 — ArcFace metric learning (LB 0.609)**: This was the first real jump. ArcFace directly optimises the embedding space for re-identification by adding a margin penalty on the angle between an embedding and its class centre. Switching from CrossEntropy to ArcFace: +47%.\n\n**v04 — Backbone search (LB 0.797)**: Compared tf_efficientnetv2_s, convnext_small, vit_small. ConvNeXt won cleanly. Added image caching + AMP to fit 3 backbones in under 25 minutes total.\n\n**v05 — Heavy augmentation + TTA (LB 0.795)**: This one hurt. Heavy aug (random rotation, aggressive colour jitter, random erasing) + 5-view test-time augmentation. Regression on LB. With only 1,895 training images, strong augmentation over-regularises — the model becomes too robust to transformations it shouldn't be robust to.\n\n**v06 — Hybrid LightGBM (LB 0.799)**: Built 571-dim pairwise features (cosine, L2, |emb_i − emb_j|, colour histogram diffs) and trained LightGBM on top of ArcFace embeddings. The LightGBM val mAP (0.491) was below raw cosine (0.510). ArcFace already encodes the right similarity — adding a learned layer on top adds noise.\n\n**v07a/b — EVA-02 Large (LB 0.918 / 0.915)**: This was the big leap. Moved from ConvNeXt-Small to EVA-02 Large (300M parameters, patch14 at 448px). The private LB went from 0.799 to 0.918 — a +15% jump in a single step. Two seeds (42 and 123) to enable ensemble diversity later.\n\n**v07f — DINOv2-Large-reg4 (LB 0.886)**: Added DINOv2 as a third model. The reasoning: EVA-02 is supervised (trained on ImageNet-21k labels), DINOv2 is self-supervised (DINO+iBOT on 142M curated images). Different pretraining paradigms should produce genuinely different feature representations, making them useful ensemble members even when all three are the same architecture class.\n\n**v07d — Pseudo-labeling (LB 0.935)**: Used the three trained models to generate pseudo-labels for test images: only images where all three models agreed (confidence > 0.90) were added to training. Then fine-tuned EVA-02a for 5 more epochs at lr=1e-5. ~100-150 pseudo-labels added.\n\n**v07e — Final ensemble (LB 0.939)**: Simple 3-way average of EVA-A + EVA-B + DINOv2 embeddings (all 1024-dim), followed by query expansion and k-reciprocal reranking. This is the selected submission.\n\n**v07g — Round-2 pseudo-labeling (LB 0.937)**: Tried a second round of pseudo-labeling on the already-fine-tuned model with a more aggressive strategy (single-model confidence using ArcFace head weights, ~300+ labels). Small regression. Stacking two PL rounds adds noise rather than signal — the first round already captured most of the easy assignments.\n\n---\n\n## The best notebook in detail (v07e)\n\nThe final pipeline has five distinct stages. Here's what each does and why.\n\n### 1. Backbone: EVA-02 Large @ 448px\n\nEVA-02 is a ViT-based model (patch size 14, 448×448 input → 32×32 = 1024 patch tokens) pretrained with masked image modelling on ImageNet-21k then fine-tuned on ImageNet-1k. The key property for re-ID is the resolution: at 448px with 14px patches, each token covers a 14×14px region of the image. Jaguar spot patterns have detail at roughly that scale, so the model's attention can resolve individual spots rather than blurred regions.\n\nConvNeXt at 224px with its effective receptive fields was already working well, but EVA-02 at 448px sees nearly 4× more pixel area per decision. That's the main reason for the +15% LB jump.\n\nTwo seeds (42 and 123) are trained separately. They converge to slightly different regions of the embedding space — enough to reduce ensemble variance when averaged later.\n\n### 2. Embedding head: GeM + BN + ArcFace\n\n**GeM pooling** (Generalised Mean Pooling, learnable p≈3): takes the 1024 patch tokens from the ViT, reshapes them into a 32×32 spatial grid, and pools with a learnable power mean. When p=1 it's average pooling; as p increases it increasingly emphasises the highest-activating (most discriminative) spatial locations. For re-ID this matters — you want the spot pattern regions to dominate the embedding, not the background.\n\n**Batch normalisation** on the embedding before ArcFace stabilises training and makes the embedding norm uniform across samples, which is important for cosine similarity to be meaningful.\n\n**ArcFace** (s=30, m=0.5): the angular margin loss. For each batch, it adds a penalty of 0.5 radians (≈29°) to the angle between each embedding and its correct class centre. This forces the model to place embeddings closer to their class centre than any other centre — directly optimising the geometry used at inference.\n\n### 3. Training: EMA + gradient checkpointing\n\n**EMA** (exponential moving average, decay=0.999): maintains a shadow copy of the model weights that is a running average over the last ~1000 gradient steps. The shadow model is used at inference. This smooths out late-training oscillations and consistently produces better retrieval performance than the final raw checkpoint.\n\n**Gradient checkpointing**: at 448px with a 300M parameter ViT and batch size 4, we're near the P100's VRAM limit. Gradient checkpointing recomputes activations during the backward pass rather than storing them, trading ~30% compute overhead for ~40% VRAM savings. Without it, training would require batch size 1-2 which hurts convergence.\n\n**20 epochs matters**: v07a at 10 epochs scored 0.849. At 20 epochs it scored 0.918. The +0.069 gap is because EVA-02's large capacity takes time to converge on a 1,895-image dataset — it's learning slowly but surely.\n\n### 4. Pseudo-labeling\n\nAfter training all three models, we use them to label the 371 test images. Only images where all three models predicted the same identity with confidence > 0.90 were accepted as pseudo-labels. This 3-way agreement filter is conservative (~100-150 samples) but high precision — we're adding mostly correct labels.\n\nThe fine-tuning step (5 epochs at lr=1e-5 = 10× smaller than the original lr) is deliberately gentle. The goal is to reinforce known patterns with the new test-domain samples, not to overfit to potentially noisy pseudo-labels.\n\n### 5. Ensemble + post-processing\n\n**3-way average**: simple arithmetic mean of the L2-normalised embeddings from EVA-A, EVA-B, and DINOv2 (all 1024-dim). Then re-normalise. This is the right fusion strategy for same-dimension embeddings: it's equivalent to taking the centroid in the embedding space. We tried concatenation (EVA 1024d + ConvNeXt 1536d) earlier and it regressed results — concatenation puts architecturally different features in the same vector without any calibration, the distance metric becomes meaningless.\n\n**Query Expansion** (top-3 DBA): for each image, replace its embedding with the average of itself and its 3 nearest neighbours in the gallery. This uses the assumption that neighbours in embedding space are likely the same identity — a reasonable assumption for well-trained ArcFace models.\n\n**K-reciprocal reranking** (k1=20, λ=0.3): the Zhong et al. 2017 reranking algorithm. The idea: if image A is in image B's top-k neighbourhood AND B is in A's top-k neighbourhood, they are \"k-reciprocal nearest neighbours\" — very likely the same identity. It computes a Jaccard similarity over these reciprocal sets and blends it with the original cosine similarity (λ controls the blend). This consistently boosted scores by 0.01-0.02 throughout the competition.\n\n---\n\n## What didn't work\n\n- **Heavy augmentation**: over-regularises at 1,895 images. Jaguars are already diverse enough.\n- **LightGBM on pairwise features**: ArcFace geometry is already optimal for cosine; adding a nonlinear layer on top hurts.\n- **Cross-architecture concat**: mixing 1024-dim and 1536-dim embeddings. The Euclidean distances become dominated by the larger-norm component.\n- **Round-2 pseudo-labeling**: stacking PL rounds. Second round adds noise the first round already filtered.\n- ** Convnext underperformed compared to ViTs\n---\n\n## Project\n\nAll training and inference code is on GitHub: https://github.com/Smooth-Cactus0/jaguar-re-identification. The repo includes all notebook versions (v01–v07g), EDA notebooks, reusable `src/` modules, and benchmark results.\n\nMost of the notebooks in this project were built with the help of Claude Sonnet 4.6\n\n— Alexy",
    "3421400": "Great writeup! Thanks for sharing your journey, @alexycactus"
  },
  "source": "meta"
}