{
  "id": 682238,
  "title": "[Round 1: 2nd Place / Round 2: 1st Place] Solution - Jaguar Re-ID",
  "url": "/competitions/jaguar-re-id/writeups/round-1-2nd-place-round-2-1st-place-solution",
  "author_name": "",
  "post_date": "2026-03-18T03:21:34.063Z",
  "votes": 6,
  "comment_count": 1,
  "views": 0,
  "content": "<h2>Summary</h2>\n<p>I fine-tuned DINOv3 ViT-7B as the backbone using QLoRA, and trained discriminative embeddings with ArcFace Loss for individual re-identification. At inference time, I applied 5-fold ensemble + TTA + k-Reciprocal Re-ranking. The same model and pipeline were used for both Round 1 and Round 2 without any modifications.</p>\n<h2>Key Insight: DINOv3 Scaling Law</h2>\n<p>The most important finding in this competition was that <strong>DINOv3 is highly effective for individual re-identification tasks</strong>. Through self-supervised learning, DINOv3 acquires powerful visual representations that capture local features (e.g., patterns, spot arrangements), making it an excellent fit for fine-grained recognition tasks such as jaguar re-identification.\nFurthermore, <strong>larger DINOv3 models consistently yielded higher scores</strong>. Performance improved as I scaled from Large → Huge, and the best score was achieved with ViT-7B/16.\nHowever, ViT-7B/16 has approximately 7 billion parameters, making <strong>full fine-tuning infeasible due to GPU memory constraints</strong>. To overcome this, I adopted QLoRA — drastically reducing memory usage via 4-bit quantization while efficiently fine-tuning through LoRA adapters — which allowed me to leverage the benefits of this massive model.</p>\n<h2>Model Architecture</h2>\n<ul>\n<li><strong>Backbone</strong>: <code>facebook/dinov3-vit7b16-pretrain-lvd1689m</code> (DINOv3 ViT-7B)</li>\n<li><strong>Embedding Head</strong>: CLS token → Linear(hidden_size, 1024) → BatchNorm1d → L2 Normalize</li>\n<li><strong>Classification Head</strong>: ArcFace (s=30.0, m=0.5, num_classes=31)\nBy leveraging DINOv3's powerful visual representations and training with ArcFace Loss, the model learned a highly discriminative embedding space for distinguishing between individuals.</li>\n</ul>\n<h2>QLoRA Fine-tuning</h2>\n<p>As mentioned above, full fine-tuning of ViT-7B/16's 7 billion parameters was infeasible due to GPU memory constraints, so I adopted QLoRA.</p>\n<ul>\n<li><strong>Quantization</strong>: 4-bit NF4 + Double Quantization (BitsAndBytes)</li>\n<li><strong>LoRA</strong>: Applied to q_proj, k_proj, v_proj (r=16, alpha=32, dropout=0.1)</li>\n<li>All backbone parameters were frozen; only LoRA adapters, Embedding Head, and ArcFace Head were trained\nThis drastically reduced the number of trainable parameters while maintaining high performance.</li>\n</ul>\n<h2>Training Details</h2>\n<table>\n<thead>\n<tr>\n<th>Parameter</th>\n<th>Value</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Image Size</td>\n<td>512×512</td>\n</tr>\n<tr>\n<td>Batch Size</td>\n<td>8</td>\n</tr>\n<tr>\n<td>Optimizer</td>\n<td>AdamW (lr=1e-4, weight_decay=1e-4)</td>\n</tr>\n<tr>\n<td>Scheduler</td>\n<td>CosineAnnealingLR (eta_min=1e-6)</td>\n</tr>\n<tr>\n<td>Epochs</td>\n<td>20</td>\n</tr>\n<tr>\n<td>CV</td>\n<td>5-fold StratifiedKFold</td>\n</tr>\n<tr>\n<td>Precision</td>\n<td>Mixed Precision (FP16)</td>\n</tr>\n</tbody>\n</table>\n<h3>Data Augmentation</h3>\n<ul>\n<li>HorizontalFlip (p=0.5)</li>\n<li>Affine (translate=10%, scale=0.85~1.15, rotate=±15°, p=0.5)</li>\n<li>ColorJitter (brightness=0.2, contrast=0.2, saturation=0.2, hue=0.1, p=0.5)</li>\n<li>GaussNoise (std=0.02~0.1, p=0.3)</li>\n</ul>\n<h3>Balanced Sampling</h3>\n<p>Due to class imbalance among the 31 individuals, I implemented a custom sampler that uniformly samples from multiple classes per batch (4 samples per class per batch). For ArcFace training, it is important that each batch contains multiple samples of the same individual, and this strategy proved effective.</p>\n<h2>Inference</h2>\n<h3>5-fold Ensemble</h3>\n<p>Embeddings were extracted from each fold's model, averaged, and then L2-normalized.</p>\n<h3>Test Time Augmentation (TTA)</h3>\n<ul>\n<li>Averaged embeddings from Original + HorizontalFlip (2 augmentations)</li>\n</ul>\n<h3>k-Reciprocal Re-ranking</h3>\n<p>For the final similarity scores, I applied <a href=\"https://arxiv.org/abs/1701.08398\" target=\"_blank\">k-Reciprocal Re-ranking</a> (k1=15, k2=4, λ=0.3). By combining Jaccard distance from k-reciprocal nearest neighbors with the initial cosine distance-based ranking, the method produces more robust similarity scores.\nFinal score = 0.3 × Jaccard distance + 0.7 × Cosine distance</p>\n<h2>What Worked</h2>\n<ol>\n<li><strong>DINOv3</strong>: Self-supervised visual representations were highly effective for individual re-identification. Larger models consistently scored higher (Large → Huge → 7B), with clear scaling benefits</li>\n<li><strong>QLoRA</strong>: ViT-7B/16 was too large for full fine-tuning, but 4-bit quantization + LoRA made it trainable on a single GPU — this was the key enabler</li>\n<li><strong>ArcFace Loss</strong>: Explicitly enforces inter-class margins, making it a better fit for Re-ID tasks than standard classification losses</li>\n<li><strong>Balanced Sampling</strong>: Addressed class imbalance and ensured ArcFace could compare multiple individuals within each batch</li>\n<li><strong>k-Reciprocal Re-ranking</strong>: Significant score improvement from post-processing alone; a well-established technique in Re-ID</li>\n</ol>\n<h2>What Didn't Work / Tried</h2>\n<ul>\n<li>EMA (Exponential Moving Average): Incompatible with QLoRA (deepcopy of quantized models is not possible)</li>\n<li>Larger image size (1024): Abandoned due to memory constraints</li>\n</ul>",
  "messages": [
    {
      "id": "3422930",
      "postDate": "03/18/2026 03:20:55",
      "content": "<h2>Summary</h2>\n<p>I fine-tuned DINOv3 ViT-7B as the backbone using QLoRA, and trained discriminative embeddings with ArcFace Loss for individual re-identification. At inference time, I applied 5-fold ensemble + TTA + k-Reciprocal Re-ranking. The same model and pipeline were used for both Round 1 and Round 2 without any modifications.</p>\n<h2>Key Insight: DINOv3 Scaling Law</h2>\n<p>The most important finding in this competition was that <strong>DINOv3 is highly effective for individual re-identification tasks</strong>. Through self-supervised learning, DINOv3 acquires powerful visual representations that capture local features (e.g., patterns, spot arrangements), making it an excellent fit for fine-grained recognition tasks such as jaguar re-identification.\nFurthermore, <strong>larger DINOv3 models consistently yielded higher scores</strong>. Performance improved as I scaled from Large → Huge, and the best score was achieved with ViT-7B/16.\nHowever, ViT-7B/16 has approximately 7 billion parameters, making <strong>full fine-tuning infeasible due to GPU memory constraints</strong>. To overcome this, I adopted QLoRA — drastically reducing memory usage via 4-bit quantization while efficiently fine-tuning through LoRA adapters — which allowed me to leverage the benefits of this massive model.</p>\n<h2>Model Architecture</h2>\n<ul>\n<li><strong>Backbone</strong>: <code>facebook/dinov3-vit7b16-pretrain-lvd1689m</code> (DINOv3 ViT-7B)</li>\n<li><strong>Embedding Head</strong>: CLS token → Linear(hidden_size, 1024) → BatchNorm1d → L2 Normalize</li>\n<li><strong>Classification Head</strong>: ArcFace (s=30.0, m=0.5, num_classes=31)\nBy leveraging DINOv3's powerful visual representations and training with ArcFace Loss, the model learned a highly discriminative embedding space for distinguishing between individuals.</li>\n</ul>\n<h2>QLoRA Fine-tuning</h2>\n<p>As mentioned above, full fine-tuning of ViT-7B/16's 7 billion parameters was infeasible due to GPU memory constraints, so I adopted QLoRA.</p>\n<ul>\n<li><strong>Quantization</strong>: 4-bit NF4 + Double Quantization (BitsAndBytes)</li>\n<li><strong>LoRA</strong>: Applied to q_proj, k_proj, v_proj (r=16, alpha=32, dropout=0.1)</li>\n<li>All backbone parameters were frozen; only LoRA adapters, Embedding Head, and ArcFace Head were trained\nThis drastically reduced the number of trainable parameters while maintaining high performance.</li>\n</ul>\n<h2>Training Details</h2>\n<table>\n<thead>\n<tr>\n<th>Parameter</th>\n<th>Value</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Image Size</td>\n<td>512×512</td>\n</tr>\n<tr>\n<td>Batch Size</td>\n<td>8</td>\n</tr>\n<tr>\n<td>Optimizer</td>\n<td>AdamW (lr=1e-4, weight_decay=1e-4)</td>\n</tr>\n<tr>\n<td>Scheduler</td>\n<td>CosineAnnealingLR (eta_min=1e-6)</td>\n</tr>\n<tr>\n<td>Epochs</td>\n<td>20</td>\n</tr>\n<tr>\n<td>CV</td>\n<td>5-fold StratifiedKFold</td>\n</tr>\n<tr>\n<td>Precision</td>\n<td>Mixed Precision (FP16)</td>\n</tr>\n</tbody>\n</table>\n<h3>Data Augmentation</h3>\n<ul>\n<li>HorizontalFlip (p=0.5)</li>\n<li>Affine (translate=10%, scale=0.85~1.15, rotate=±15°, p=0.5)</li>\n<li>ColorJitter (brightness=0.2, contrast=0.2, saturation=0.2, hue=0.1, p=0.5)</li>\n<li>GaussNoise (std=0.02~0.1, p=0.3)</li>\n</ul>\n<h3>Balanced Sampling</h3>\n<p>Due to class imbalance among the 31 individuals, I implemented a custom sampler that uniformly samples from multiple classes per batch (4 samples per class per batch). For ArcFace training, it is important that each batch contains multiple samples of the same individual, and this strategy proved effective.</p>\n<h2>Inference</h2>\n<h3>5-fold Ensemble</h3>\n<p>Embeddings were extracted from each fold's model, averaged, and then L2-normalized.</p>\n<h3>Test Time Augmentation (TTA)</h3>\n<ul>\n<li>Averaged embeddings from Original + HorizontalFlip (2 augmentations)</li>\n</ul>\n<h3>k-Reciprocal Re-ranking</h3>\n<p>For the final similarity scores, I applied <a href=\"https://arxiv.org/abs/1701.08398\" target=\"_blank\">k-Reciprocal Re-ranking</a> (k1=15, k2=4, λ=0.3). By combining Jaccard distance from k-reciprocal nearest neighbors with the initial cosine distance-based ranking, the method produces more robust similarity scores.\nFinal score = 0.3 × Jaccard distance + 0.7 × Cosine distance</p>\n<h2>What Worked</h2>\n<ol>\n<li><strong>DINOv3</strong>: Self-supervised visual representations were highly effective for individual re-identification. Larger models consistently scored higher (Large → Huge → 7B), with clear scaling benefits</li>\n<li><strong>QLoRA</strong>: ViT-7B/16 was too large for full fine-tuning, but 4-bit quantization + LoRA made it trainable on a single GPU — this was the key enabler</li>\n<li><strong>ArcFace Loss</strong>: Explicitly enforces inter-class margins, making it a better fit for Re-ID tasks than standard classification losses</li>\n<li><strong>Balanced Sampling</strong>: Addressed class imbalance and ensured ArcFace could compare multiple individuals within each batch</li>\n<li><strong>k-Reciprocal Re-ranking</strong>: Significant score improvement from post-processing alone; a well-established technique in Re-ID</li>\n</ol>\n<h2>What Didn't Work / Tried</h2>\n<ul>\n<li>EMA (Exponential Moving Average): Incompatible with QLoRA (deepcopy of quantized models is not possible)</li>\n<li>Larger image size (1024): Abandoned due to memory constraints</li>\n</ul>",
      "rawMarkdown": "## Summary\n\nI fine-tuned DINOv3 ViT-7B as the backbone using QLoRA, and trained discriminative embeddings with ArcFace Loss for individual re-identification. At inference time, I applied 5-fold ensemble + TTA + k-Reciprocal Re-ranking. The same model and pipeline were used for both Round 1 and Round 2 without any modifications.\n\n## Key Insight: DINOv3 Scaling Law\n\nThe most important finding in this competition was that **DINOv3 is highly effective for individual re-identification tasks**. Through self-supervised learning, DINOv3 acquires powerful visual representations that capture local features (e.g., patterns, spot arrangements), making it an excellent fit for fine-grained recognition tasks such as jaguar re-identification.\n\nFurthermore, **larger DINOv3 models consistently yielded higher scores**. Performance improved as I scaled from Large → Huge, and the best score was achieved with ViT-7B/16.\n\nHowever, ViT-7B/16 has approximately 7 billion parameters, making **full fine-tuning infeasible due to GPU memory constraints**. To overcome this, I adopted QLoRA — drastically reducing memory usage via 4-bit quantization while efficiently fine-tuning through LoRA adapters — which allowed me to leverage the benefits of this massive model.\n\n## Model Architecture\n\n- **Backbone**: `facebook/dinov3-vit7b16-pretrain-lvd1689m` (DINOv3 ViT-7B)\n- **Embedding Head**: CLS token → Linear(hidden_size, 1024) → BatchNorm1d → L2 Normalize\n- **Classification Head**: ArcFace (s=30.0, m=0.5, num_classes=31)\n\nBy leveraging DINOv3's powerful visual representations and training with ArcFace Loss, the model learned a highly discriminative embedding space for distinguishing between individuals.\n\n## QLoRA Fine-tuning\n\nAs mentioned above, full fine-tuning of ViT-7B/16's 7 billion parameters was infeasible due to GPU memory constraints, so I adopted QLoRA.\n\n- **Quantization**: 4-bit NF4 + Double Quantization (BitsAndBytes)\n- **LoRA**: Applied to q_proj, k_proj, v_proj (r=16, alpha=32, dropout=0.1)\n- All backbone parameters were frozen; only LoRA adapters, Embedding Head, and ArcFace Head were trained\n\nThis drastically reduced the number of trainable parameters while maintaining high performance.\n\n## Training Details\n\n| Parameter | Value |\n|---|---|\n| Image Size | 512×512 |\n| Batch Size | 8 |\n| Optimizer | AdamW (lr=1e-4, weight_decay=1e-4) |\n| Scheduler | CosineAnnealingLR (eta_min=1e-6) |\n| Epochs | 20 |\n| CV | 5-fold StratifiedKFold |\n| Precision | Mixed Precision (FP16) |\n\n### Data Augmentation\n\n- HorizontalFlip (p=0.5)\n- Affine (translate=10%, scale=0.85~1.15, rotate=±15°, p=0.5)\n- ColorJitter (brightness=0.2, contrast=0.2, saturation=0.2, hue=0.1, p=0.5)\n- GaussNoise (std=0.02~0.1, p=0.3)\n\n### Balanced Sampling\n\nDue to class imbalance among the 31 individuals, I implemented a custom sampler that uniformly samples from multiple classes per batch (4 samples per class per batch). For ArcFace training, it is important that each batch contains multiple samples of the same individual, and this strategy proved effective.\n\n## Inference\n\n### 5-fold Ensemble\n\nEmbeddings were extracted from each fold's model, averaged, and then L2-normalized.\n\n### Test Time Augmentation (TTA)\n\n- Averaged embeddings from Original + HorizontalFlip (2 augmentations)\n\n### k-Reciprocal Re-ranking\n\nFor the final similarity scores, I applied [k-Reciprocal Re-ranking](https://arxiv.org/abs/1701.08398) (k1=15, k2=4, λ=0.3). By combining Jaccard distance from k-reciprocal nearest neighbors with the initial cosine distance-based ranking, the method produces more robust similarity scores.\n\nFinal score = 0.3 × Jaccard distance + 0.7 × Cosine distance\n\n## What Worked\n\n1. **DINOv3**: Self-supervised visual representations were highly effective for individual re-identification. Larger models consistently scored higher (Large → Huge → 7B), with clear scaling benefits\n2. **QLoRA**: ViT-7B/16 was too large for full fine-tuning, but 4-bit quantization + LoRA made it trainable on a single GPU — this was the key enabler\n3. **ArcFace Loss**: Explicitly enforces inter-class margins, making it a better fit for Re-ID tasks than standard classification losses\n4. **Balanced Sampling**: Addressed class imbalance and ensured ArcFace could compare multiple individuals within each batch\n5. **k-Reciprocal Re-ranking**: Significant score improvement from post-processing alone; a well-established technique in Re-ID\n\n## What Didn't Work / Tried\n\n- EMA (Exponential Moving Average): Incompatible with QLoRA (deepcopy of quantized models is not possible)\n- Larger image size (1024): Abandoned due to memory constraints",
      "votes": null
    },
    {
      "id": "3423246",
      "postDate": "03/18/2026 10:40:22",
      "content": "<p>Very nice writeup! Thank you</p>",
      "rawMarkdown": "Very nice writeup! Thank you",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3423246,
      "author_name": "andandand",
      "author_url": "",
      "post_date": "03/18/2026 10:40:22",
      "content": "<p>Very nice writeup! Thank you</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3422930": "## Summary\n\nI fine-tuned DINOv3 ViT-7B as the backbone using QLoRA, and trained discriminative embeddings with ArcFace Loss for individual re-identification. At inference time, I applied 5-fold ensemble + TTA + k-Reciprocal Re-ranking. The same model and pipeline were used for both Round 1 and Round 2 without any modifications.\n\n## Key Insight: DINOv3 Scaling Law\n\nThe most important finding in this competition was that **DINOv3 is highly effective for individual re-identification tasks**. Through self-supervised learning, DINOv3 acquires powerful visual representations that capture local features (e.g., patterns, spot arrangements), making it an excellent fit for fine-grained recognition tasks such as jaguar re-identification.\n\nFurthermore, **larger DINOv3 models consistently yielded higher scores**. Performance improved as I scaled from Large → Huge, and the best score was achieved with ViT-7B/16.\n\nHowever, ViT-7B/16 has approximately 7 billion parameters, making **full fine-tuning infeasible due to GPU memory constraints**. To overcome this, I adopted QLoRA — drastically reducing memory usage via 4-bit quantization while efficiently fine-tuning through LoRA adapters — which allowed me to leverage the benefits of this massive model.\n\n## Model Architecture\n\n- **Backbone**: `facebook/dinov3-vit7b16-pretrain-lvd1689m` (DINOv3 ViT-7B)\n- **Embedding Head**: CLS token → Linear(hidden_size, 1024) → BatchNorm1d → L2 Normalize\n- **Classification Head**: ArcFace (s=30.0, m=0.5, num_classes=31)\n\nBy leveraging DINOv3's powerful visual representations and training with ArcFace Loss, the model learned a highly discriminative embedding space for distinguishing between individuals.\n\n## QLoRA Fine-tuning\n\nAs mentioned above, full fine-tuning of ViT-7B/16's 7 billion parameters was infeasible due to GPU memory constraints, so I adopted QLoRA.\n\n- **Quantization**: 4-bit NF4 + Double Quantization (BitsAndBytes)\n- **LoRA**: Applied to q_proj, k_proj, v_proj (r=16, alpha=32, dropout=0.1)\n- All backbone parameters were frozen; only LoRA adapters, Embedding Head, and ArcFace Head were trained\n\nThis drastically reduced the number of trainable parameters while maintaining high performance.\n\n## Training Details\n\n| Parameter | Value |\n|---|---|\n| Image Size | 512×512 |\n| Batch Size | 8 |\n| Optimizer | AdamW (lr=1e-4, weight_decay=1e-4) |\n| Scheduler | CosineAnnealingLR (eta_min=1e-6) |\n| Epochs | 20 |\n| CV | 5-fold StratifiedKFold |\n| Precision | Mixed Precision (FP16) |\n\n### Data Augmentation\n\n- HorizontalFlip (p=0.5)\n- Affine (translate=10%, scale=0.85~1.15, rotate=±15°, p=0.5)\n- ColorJitter (brightness=0.2, contrast=0.2, saturation=0.2, hue=0.1, p=0.5)\n- GaussNoise (std=0.02~0.1, p=0.3)\n\n### Balanced Sampling\n\nDue to class imbalance among the 31 individuals, I implemented a custom sampler that uniformly samples from multiple classes per batch (4 samples per class per batch). For ArcFace training, it is important that each batch contains multiple samples of the same individual, and this strategy proved effective.\n\n## Inference\n\n### 5-fold Ensemble\n\nEmbeddings were extracted from each fold's model, averaged, and then L2-normalized.\n\n### Test Time Augmentation (TTA)\n\n- Averaged embeddings from Original + HorizontalFlip (2 augmentations)\n\n### k-Reciprocal Re-ranking\n\nFor the final similarity scores, I applied [k-Reciprocal Re-ranking](https://arxiv.org/abs/1701.08398) (k1=15, k2=4, λ=0.3). By combining Jaccard distance from k-reciprocal nearest neighbors with the initial cosine distance-based ranking, the method produces more robust similarity scores.\n\nFinal score = 0.3 × Jaccard distance + 0.7 × Cosine distance\n\n## What Worked\n\n1. **DINOv3**: Self-supervised visual representations were highly effective for individual re-identification. Larger models consistently scored higher (Large → Huge → 7B), with clear scaling benefits\n2. **QLoRA**: ViT-7B/16 was too large for full fine-tuning, but 4-bit quantization + LoRA made it trainable on a single GPU — this was the key enabler\n3. **ArcFace Loss**: Explicitly enforces inter-class margins, making it a better fit for Re-ID tasks than standard classification losses\n4. **Balanced Sampling**: Addressed class imbalance and ensured ArcFace could compare multiple individuals within each batch\n5. **k-Reciprocal Re-ranking**: Significant score improvement from post-processing alone; a well-established technique in Re-ID\n\n## What Didn't Work / Tried\n\n- EMA (Exponential Moving Average): Incompatible with QLoRA (deepcopy of quantized models is not possible)\n- Larger image size (1024): Abandoned due to memory constraints",
    "3423246": "Very nice writeup! Thank you"
  },
  "source": "meta"
}