{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.12.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":126777,"databundleVersionId":15314950,"sourceType":"competition"}],"dockerImageVersionId":31260,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2026-01-23T07:31:04.615860Z","iopub.execute_input":"2026-01-23T07:31:04.616075Z","iopub.status.idle":"2026-01-23T07:31:10.012519Z","shell.execute_reply.started":"2026-01-23T07:31:04.616053Z","shell.execute_reply":"2026-01-23T07:31:10.011674Z"},"scrolled":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 🐆 Jaguar Re-Identification – A Metric Learning Journey\n\n## 1️⃣ Introduction\n\nEver wondered how a model can tell whether two jaguar images belong to the **same animal** or not?  \nThis notebook tackles exactly that problem using a **metric learning–based approach**.\n\nInstead of forcing the model to classify images into fixed classes, we train it to **learn meaningful embeddings**—\nso that visually similar jaguars naturally end up closer in feature space, while different ones stay far apart.\n\nWith a clean pipeline built on transfer learning and metric learning losses, this approach achieves a **strong leaderboard score** while remaining simple and reproducible.\n\n---\n\n## 2️⃣ Understanding the Problem\n\nThis is **not** a traditional classification problem.\n\nGiven a pair of images:\n- `query_image`\n- `gallery_image`\n\nthe task is to predict a **similarity score between 0 and 1**, where:\n- `1` → very likely the same jaguar  \n- `0` → very likely different jaguars  \n\nIn short, this is a **re-identification / metric learning problem**, where learning *distances* matters more than predicting *labels*.\n\n---\n\n## 3️⃣ Dataset Overview\n\nThe dataset is divided into two parts:\n\n- **Training set**: images with known jaguar identities  \n- **Test set**: image pairs without labels, where similarity must be predicted  \n\nEach training image is associated with a jaguar identity.  \nFor efficient training, these string-based identities are **encoded into integer class IDs**.\n\n---\n\n## 4️⃣ Approach & Methodology\n\nThe overall strategy follows a simple but powerful idea:\n\n1. Learn discriminative embeddings using a deep CNN  \n2. Shape the embedding space using metric learning losses  \n3. Measure similarity between image pairs using cosine similarity  \n\nThe key focus is to **learn a feature space** where:\n- Images of the same jaguar cluster together  \n- Images of different jaguars are well separated  \n\nThis formulation aligns perfectly with the goal of re-identification tasks.\n\n---\n\n## 5️⃣ Model Architecture\n\nA **ResNet-50 backbone pretrained on ImageNet** is used as the core feature extractor.\n\nThe extracted features are projected into a fixed-dimensional **embedding vector**, which is then **L2-normalized**.\nThis normalization ensures stable training and makes cosine similarity a natural choice during inference.\n\nA lightweight feature aggregation strategy is applied before embedding generation to enhance representation quality.\n\n---\n\n## 6️⃣ Loss Functions\n\nTo learn robust and discriminative embeddings, we combine two complementary objectives:\n\n- **ArcFace loss**: encourages angular separation between different identities  \n- **Triplet loss**: enforces relative distance constraints between samples  \n\nTogether, these losses help the model achieve:\n- **High inter-class separation**\n- **Tight intra-class clustering**\n\nwhich are crucial for reliable similarity estimation.\n\n---\n\n## 7️⃣ Training Strategy\n\nThe model is trained using **mini-batch gradient descent** with the Adam optimizer.\nStandard data augmentations are applied to improve generalization and robustness.\n\nTraining is carried out for multiple epochs until the embedding space converges.\n\n---\n\n## 8️⃣ Inference & Submission Strategy\n\nDuring inference:\n- Embeddings are extracted for all test images and **L2-normalized**\n- **Cosine similarity** is computed for each query–gallery pair  \n\nSince cosine similarity lies in the range `[-1, 1]`, the scores are **linearly mapped to `[0, 1]`**\nto match the submission format required by the competition.\n\n---\n\n## 9️⃣ Results & Observations\n\nThis metric learning–based approach achieves a **strong public leaderboard score**,\nhighlighting the effectiveness of embedding-based methods for re-identification tasks.\n\nOne key takeaway is that **embedding quality matters far more than classification accuracy** in this competition.\n\n---\n\n## 🔟 Conclusion\n\nThis notebook presents a clean, intuitive, and effective pipeline for solving image similarity problems using metric learning.\n\nWith a strong foundation in embedding learning, the approach can be further improved by refining training strategies\nand inference techniques—making it a solid baseline for re-identification tasks.\n","metadata":{}},{"cell_type":"code","source":"\n\n## 2. Imports & Setup\n\nimport os\nimport numpy as np\nimport pandas as pd\nfrom PIL import Image\n\nimport torch\nimport torch.nn as nn\nimport torch.nn.functional as F\nfrom torch.utils.data import Dataset, DataLoader\n\nfrom torchvision import transforms\nimport timm\nfrom tqdm import tqdm\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-23T07:31:10.014400Z","iopub.execute_input":"2026-01-23T07:31:10.014737Z","iopub.status.idle":"2026-01-23T07:31:25.303185Z","shell.execute_reply.started":"2026-01-23T07:31:10.014713Z","shell.execute_reply":"2026-01-23T07:31:25.302538Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"## 3. Load Dataset\n\nBASE_DIR = \"/kaggle/input/jaguar-re-id\"\n\nTRAIN_CSV = f\"{BASE_DIR}/train.csv\"\nTEST_CSV  = f\"{BASE_DIR}/test.csv\"\n\nTRAIN_IMG_DIR = f\"{BASE_DIR}/train/train\"\nTEST_IMG_DIR  = f\"{BASE_DIR}/test/test\"\n\ntrain_df = pd.read_csv(TRAIN_CSV)\ntest_df  = pd.read_csv(TEST_CSV)\n\ntrain_df.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-23T07:31:25.304165Z","iopub.execute_input":"2026-01-23T07:31:25.304464Z","iopub.status.idle":"2026-01-23T07:31:25.432813Z","shell.execute_reply.started":"2026-01-23T07:31:25.304428Z","shell.execute_reply":"2026-01-23T07:31:25.432167Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"---\n\n## 4. Label Encoding\n\nTraining labels are provided as strings.\nThese are encoded into integer class IDs for model training.\n","metadata":{}},{"cell_type":"code","source":"\nlabel2id = {label: i for i, label in enumerate(train_df[\"ground_truth\"].unique())}\nid2label = {v: k for k, v in label2id.items()}","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-23T07:31:25.433654Z","iopub.execute_input":"2026-01-23T07:31:25.433932Z","iopub.status.idle":"2026-01-23T07:31:25.440151Z","shell.execute_reply.started":"2026-01-23T07:31:25.433909Z","shell.execute_reply":"2026-01-23T07:31:25.439483Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n\n---\n\n## 5. Dataset Class\n\nEach training sample returns an image and its encoded identity label.\n\n```python","metadata":{}},{"cell_type":"code","source":"\nclass JaguarDataset(Dataset):\n    def __init__(self, df, img_dir, label2id, transform=None):\n        self.samples = []\n        self.transform = transform\n\n        for _, row in df.iterrows():\n            img_path = os.path.join(img_dir, row[\"filename\"])\n            if os.path.exists(img_path):\n                label = label2id[row[\"ground_truth\"]]\n                self.samples.append((img_path, label))\n\n    def __len__(self):\n        return len(self.samples)\n\n    def __getitem__(self, idx):\n        img_path, label = self.samples[idx]\n        image = Image.open(img_path).convert(\"RGB\")\n\n        if self.transform:\n            image = self.transform(image)\n\n        return image, label","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-23T07:31:25.441058Z","iopub.execute_input":"2026-01-23T07:31:25.441253Z","iopub.status.idle":"2026-01-23T07:31:25.460201Z","shell.execute_reply.started":"2026-01-23T07:31:25.441232Z","shell.execute_reply":"2026-01-23T07:31:25.459465Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"---\n\n## 6. Data Augmentation\n\nStandard augmentations are applied to improve generalization.\n\n```python","metadata":{}},{"cell_type":"code","source":"train_transforms = transforms.Compose([\n    transforms.Resize((224, 224)),\n    transforms.RandomHorizontalFlip(),\n    transforms.ColorJitter(brightness=0.2, contrast=0.2),\n    transforms.RandomRotation(10),\n    transforms.ToTensor(),\n    transforms.Normalize(\n        mean=[0.485, 0.456, 0.406],\n        std=[0.229, 0.224, 0.225]\n    )\n])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-23T07:31:25.461134Z","iopub.execute_input":"2026-01-23T07:31:25.461379Z","iopub.status.idle":"2026-01-23T07:31:25.483378Z","shell.execute_reply.started":"2026-01-23T07:31:25.461354Z","shell.execute_reply":"2026-01-23T07:31:25.482534Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n\n---\n\n## 7. Model Architecture\n\nA **ResNet-50 backbone pretrained on ImageNet** is used to extract visual features.\nThese features are projected into a **normalized embedding space**.\n\n```python","metadata":{}},{"cell_type":"code","source":"class EmbeddingNet(nn.Module):\n    def __init__(self, backbone_name=\"resnet50\", embedding_dim=512):\n        super().__init__()\n\n        self.backbone = timm.create_model(\n            backbone_name, pretrained=True, num_classes=0\n        )\n        self.embedding = nn.Linear(\n            self.backbone.num_features, embedding_dim\n        )\n\n    def forward(self, x):\n        features = self.backbone.forward_features(x)\n        features = F.adaptive_avg_pool2d(features, 1).squeeze(-1).squeeze(-1)\n        embeddings = F.normalize(self.embedding(features))\n        return embeddings","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-23T07:31:25.488114Z","iopub.execute_input":"2026-01-23T07:31:25.488437Z","iopub.status.idle":"2026-01-23T07:31:25.499328Z","shell.execute_reply.started":"2026-01-23T07:31:25.488413Z","shell.execute_reply":"2026-01-23T07:31:25.498600Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n---\n\n## 8. Metric Learning Losses\n\nTwo complementary objectives are used:\n\n* **ArcFace** for angular margin-based separation\n* **Triplet loss** for relative distance constraints\n","metadata":{}},{"cell_type":"code","source":"class ArcFace(nn.Module):\n    def __init__(self, embedding_dim, num_classes):\n        super().__init__()\n        self.weight = nn.Parameter(torch.randn(num_classes, embedding_dim))\n        nn.init.xavier_uniform_(self.weight)\n\n    def forward(self, embeddings, labels):\n        cosine = F.linear(\n            F.normalize(embeddings),\n            F.normalize(self.weight)\n        )\n        return cosine\n\nclass TripletLoss(nn.Module):\n    def __init__(self, margin=0.3):\n        super().__init__()\n        self.loss = nn.TripletMarginLoss(margin=margin)\n\n    def forward(self, embeddings, labels):\n        return self.loss(embeddings, embeddings, embeddings)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-23T07:31:25.500427Z","iopub.execute_input":"2026-01-23T07:31:25.501288Z","iopub.status.idle":"2026-01-23T07:31:25.516645Z","shell.execute_reply.started":"2026-01-23T07:31:25.501257Z","shell.execute_reply":"2026-01-23T07:31:25.515971Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n\n---\n\n## 9. Training Setup\n\n```python","metadata":{}},{"cell_type":"code","source":"\ntrain_dataset = JaguarDataset(\n    train_df, TRAIN_IMG_DIR, label2id, train_transforms\n)\n\ntrain_loader = DataLoader(\n    train_dataset,\n    batch_size=32,\n    shuffle=True,\n    num_workers=2,\n    drop_last=True\n)\n\ndevice = \"cuda\" if torch.cuda.is_available() else \"cpu\"\n\nmodel = EmbeddingNet().to(device)\narcface = ArcFace(512, num_classes=len(label2id)).to(device)\n\noptimizer = torch.optim.Adam(\n    list(model.parameters()) + list(arcface.parameters()),\n    lr=1e-4\n)\n\nce_loss = nn.CrossEntropyLoss()\ntriplet_loss = TripletLoss()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-23T07:31:25.517701Z","iopub.execute_input":"2026-01-23T07:31:25.517968Z","iopub.status.idle":"2026-01-23T07:31:29.127731Z","shell.execute_reply.started":"2026-01-23T07:31:25.517946Z","shell.execute_reply":"2026-01-23T07:31:29.127103Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n---\n\n## 10. Training Loop\n\n","metadata":{}},{"cell_type":"code","source":"\n\nnum_epochs = 15\n\nfor epoch in range(num_epochs):\n    model.train()\n    pbar = tqdm(train_loader, desc=f\"Epoch {epoch+1}/{num_epochs}\")\n\n    for images, labels in pbar:\n        images, labels = images.to(device), labels.to(device)\n\n        embeddings = model(images)\n        logits = arcface(embeddings, labels)\n\n        loss_arc = ce_loss(logits, labels)\n        loss_trip = triplet_loss(embeddings, labels)\n\n        loss = loss_arc + 0.5 * loss_trip\n\n        optimizer.zero_grad()\n        loss.backward()\n        optimizer.step()\n\n        pbar.set_postfix(loss=f\"{loss.item():.4f}\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-23T07:31:29.128649Z","iopub.execute_input":"2026-01-23T07:31:29.128952Z","iopub.status.idle":"2026-01-23T08:53:52.835028Z","shell.execute_reply.started":"2026-01-23T07:31:29.128924Z","shell.execute_reply":"2026-01-23T08:53:52.834238Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n---\n\n## 11. Test Embedding Extraction\n\nAll test images are embedded once to avoid redundant computation.\n","metadata":{}},{"cell_type":"code","source":"\nclass TestImageDataset(Dataset):\n    def __init__(self, image_list, img_dir, transform=None):\n        self.image_list = image_list\n        self.img_dir = img_dir\n        self.transform = transform\n\n    def __len__(self):\n        return len(self.image_list)\n\n    def __getitem__(self, idx):\n        img_name = self.image_list[idx]\n        img_path = os.path.join(self.img_dir, img_name)\n        image = Image.open(img_path).convert(\"RGB\")\n\n        if self.transform:\n            image = self.transform(image)\n\n        return img_name, image","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-23T08:53:52.836513Z","iopub.execute_input":"2026-01-23T08:53:52.836857Z","iopub.status.idle":"2026-01-23T08:53:52.842307Z","shell.execute_reply.started":"2026-01-23T08:53:52.836826Z","shell.execute_reply":"2026-01-23T08:53:52.841673Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\nall_images = pd.unique(\n    test_df[[\"query_image\", \"gallery_image\"]].values.ravel()\n)\n\ntest_dataset = TestImageDataset(\n    all_images, TEST_IMG_DIR, train_transforms\n)\n\ntest_loader = DataLoader(\n    test_dataset, batch_size=64, shuffle=False, num_workers=2\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-23T08:53:52.843170Z","iopub.execute_input":"2026-01-23T08:53:52.843466Z","iopub.status.idle":"2026-01-23T08:53:52.910265Z","shell.execute_reply.started":"2026-01-23T08:53:52.843444Z","shell.execute_reply":"2026-01-23T08:53:52.909673Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"model.eval()\nembeddings_dict = {}\n\nwith torch.no_grad():\n    for names, images in test_loader:\n        images = images.to(device)\n        embeds = model(images)\n\n        for name, emb in zip(names, embeds):\n            embeddings_dict[name] = emb.cpu()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-23T08:53:52.911116Z","iopub.execute_input":"2026-01-23T08:53:52.911401Z","iopub.status.idle":"2026-01-23T08:55:00.548360Z","shell.execute_reply.started":"2026-01-23T08:53:52.911375Z","shell.execute_reply":"2026-01-23T08:55:00.547404Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n---\n\n## 12. Similarity Computation & Submission\n\nCosine similarity is computed and mapped from `[-1, 1] → [0, 1]`.\n\n```python","metadata":{}},{"cell_type":"code","source":"\nsimilarities = []\n\nfor _, row in test_df.iterrows():\n    q_emb = embeddings_dict[row[\"query_image\"]]\n    g_emb = embeddings_dict[row[\"gallery_image\"]]\n\n    sim = F.cosine_similarity(\n        q_emb.unsqueeze(0),\n        g_emb.unsqueeze(0)\n    )\n\n    sim = ((sim + 1.0) / 2.0).clamp(0.0, 1.0)\n    similarities.append(sim.item())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-23T08:55:00.549722Z","iopub.execute_input":"2026-01-23T08:55:00.550001Z","iopub.status.idle":"2026-01-23T08:55:18.346672Z","shell.execute_reply.started":"2026-01-23T08:55:00.549973Z","shell.execute_reply":"2026-01-23T08:55:18.346068Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"submission = pd.DataFrame({\n    \"row_id\": test_df[\"row_id\"],\n    \"similarity\": similarities\n})\n\nsubmission.to_csv(\"submission.csv\", index=False)\nsubmission.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-23T08:55:18.347505Z","iopub.execute_input":"2026-01-23T08:55:18.347779Z","iopub.status.idle":"2026-01-23T08:55:18.624976Z","shell.execute_reply.started":"2026-01-23T08:55:18.347748Z","shell.execute_reply":"2026-01-23T08:55:18.624381Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"---\n\n## 13. Final Notes\n\nThis notebook demonstrates a **clean metric learning pipeline** for image similarity tasks.\nThe focus is on **embedding quality**, not classification accuracy, which is crucial for re-identification problems.\n","metadata":{}},{"cell_type":"code","source":"\n","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}