{"metadata":{"kernelspec":{"name":"python3","display_name":"Python 3","language":"python"},"language_info":{"name":"python","version":"3.11.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceId":99249,"databundleVersionId":12808057,"sourceType":"competition"},{"sourceId":11425807,"sourceType":"datasetVersion","datasetId":7154629},{"sourceId":12476940,"sourceType":"datasetVersion","datasetId":7867242},{"sourceId":12479224,"sourceType":"datasetVersion","datasetId":7873381}],"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# 🐾 Cup-ybara Demo Submission\nThis notebook provides a fully self-contained baseline submission for the Cup Ybara challenge.\nIt defines the required `BaseModel` interface, implements a dummy `CustomModel`,\nand outputs a properly formatted `submission.csv` file.","metadata":{},"attachments":{}},{"cell_type":"code","source":"# Requirements\n\n!pip install megadetector==5.0.28\n!pip install speciesnet==5.0.0\n!pip install opencv-python\n!pip install protobuf==3.20.*\n\n#!pip install --no-index --find-links=/kaggle/input/cupybara-ariel-malowany-dependencies/packages megadetector --no-deps\n#!pip install --no-index --find-links=/kaggle/input/cupybara-ariel-malowany-dependencies/packages speciesnet --no-deps\n#!pip install --no-index --find-links=/kaggle/input/cupybara-ariel-malowany-dependencies/packages opencv-python --no-deps\n#!pip install --no-index --find-links=/kaggle/input/cupybara-ariel-malowany-dependencies/packages protobuf --no-deps\n#!pip install --no-index --find-links=/kaggle/input/cupybara-ariel-malowany-dependencies/packages humanfriendly --no-deps","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Import packages\n\nimport os\nimport sys\nimport ast\nimport kagglehub\nimport random\nimport subprocess\nfrom datetime import datetime\nimport json\nimport joblib\nimport shutil\nimport pandas as pd\nimport numpy as np\nimport tqdm\nfrom typing import List\nimport cv2\nfrom PIL import Image\nfrom sklearn.metrics import f1_score\nimport matplotlib.pyplot as plt\nimport matplotlib.patches as patches\nfrom megadetector.detection import run_detector\nfrom sklearn.metrics.pairwise import cosine_similarity\nimport torch\nfrom speciesnet import SpeciesNet","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-15T19:06:57.093303Z","iopub.execute_input":"2025-07-15T19:06:57.093618Z","execution_failed":"2025-07-15T19:07:12.145Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"LABELS = [\"'armadillo'\", \"'bird'\", \"'capybara'\", \"'cow'\", \"'dusky_legged_guan'\",\n    \"'gray_brocket'\", \"'hare'\", \"'human'\", \"'insect'\", \"'margay'\", \"'no_animal'\",\n    \"'skunk'\", \"'unknown_animal'\", \"'wild_boar'\"]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-15T18:11:08.455556Z","iopub.execute_input":"2025-07-15T18:11:08.456316Z","iopub.status.idle":"2025-07-15T18:11:08.459885Z","shell.execute_reply.started":"2025-07-15T18:11:08.456291Z","shell.execute_reply":"2025-07-15T18:11:08.459285Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"!ls /kaggle/input/cupybara/dataset/","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-15T18:11:09.974413Z","iopub.execute_input":"2025-07-15T18:11:09.975108Z","iopub.status.idle":"2025-07-15T18:11:10.126305Z","shell.execute_reply.started":"2025-07-15T18:11:09.975083Z","shell.execute_reply":"2025-07-15T18:11:10.125578Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🏋️ About Train Mode\n\nThe boolean variable TRAIN is used to determine if the portion of the trainng code should be run or not. It will always be set to `False` when running the submission manually. Ensure that if a lengthy part of your notebook is used for training, that is toggled away by this variable","metadata":{}},{"cell_type":"code","source":"DATASET_ROOT = \"/kaggle/input/cupybara/dataset/dataset/\"\nTRAIN_DIR = os.path.join(DATASET_ROOT,\"train\")  # Directory containing train `.mp4` files\nEVAL_DIR =  os.path.join(DATASET_ROOT,\"test\")  # Directory containing evaluation `.mp4` files\nTRAIN = False","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-15T18:11:11.489851Z","iopub.execute_input":"2025-07-15T18:11:11.490140Z","iopub.status.idle":"2025-07-15T18:11:11.495062Z","shell.execute_reply.started":"2025-07-15T18:11:11.490116Z","shell.execute_reply":"2025-07-15T18:11:11.494146Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"!rm -rf model_weights/ model_weights.zip submission.csv","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-15T18:11:13.490094Z","iopub.execute_input":"2025-07-15T18:11:13.490406Z","iopub.status.idle":"2025-07-15T18:11:13.633758Z","shell.execute_reply.started":"2025-07-15T18:11:13.490376Z","shell.execute_reply":"2025-07-15T18:11:13.632959Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🤖 Uploading Trained Model as Public Dataset\n\n1. Manually download the generated `model_weights.zip` file from the `Output` section on the side panel.\n2. Upload the `model_weights.zip` file as a new **PUBLIC** dataset using the `Input` sectino on the side panel (Upload -> New Dataset).\n3. You can refer to the uploaded dataset by its name (all lowercase and replacing spaces with \"-\"). In our example: `Native Fauna Dummy Model` becomes `native-fauna-dummy-model`\n\n> Ensure that the dataset is **Public** and the license is Attributtion 4.0 International (CC BY 4.0)\n\n> If you had already created a dataset, you can upload a new version by clicking `Open in New Tab`, and then checking for updates. ","metadata":{}},{"cell_type":"code","source":"!ls /kaggle/input/native-fauna-dummy-model","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-15T18:11:16.599787Z","iopub.execute_input":"2025-07-15T18:11:16.600060Z","iopub.status.idle":"2025-07-15T18:11:16.757971Z","shell.execute_reply.started":"2025-07-15T18:11:16.600036Z","shell.execute_reply":"2025-07-15T18:11:16.757185Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🧱 BaseModel Interface\nThis class defines the interface expected by the competition evaluators.\nYou must subclass `BaseModel` and implement `_load_model()` and `_predict()`.","metadata":{}},{"cell_type":"code","source":"# Helper functions\n\ndef save_image(frame, file_name, append = None, save_dir = '/kaggle/working/extracted_images'):\n  base_dir = os.path.join(save_dir, str(file_name))\n  os.makedirs(base_dir, exist_ok = True)\n  if append is not None:\n    file_name = f\"{file_name}_{append}\"\n  save_path = os.path.join(base_dir, f'{file_name}.jpg')\n  cv2.imwrite(save_path, frame)\n\ndef save_json(json_file, file_name, save_dir = '/kaggle/working/extracted_images'):\n  save_dir = os.path.join(save_dir, str(file_name))\n  os.makedirs(save_dir, exist_ok = True)\n  file_dir = os.path.join(save_dir, 'yolo_metadata.json')\n  with open(file_dir, 'w') as f:\n    yolo_metadata = json.dumps(json_file)\n    f.write(yolo_metadata)\n\ndef open_video(file_path):\n  cap = cv2.VideoCapture(file_path)\n  if not cap.isOpened():\n    print(\"Error: Could not open video.\")\n  else:\n    frame_count = int(cap.get(cv2.CAP_PROP_FRAME_COUNT))\n  return cap, frame_count\n\ndef extract_frame(cap_obj, frame_number=1, save_img=False, save_path=None):\n    cap_obj.set(cv2.CAP_PROP_POS_FRAMES, frame_number)\n    ret, frame = cap_obj.read()\n    \n    if not ret:\n        print(f\"No fue posible extraer el frame {frame_number}\")\n        return None\n\n    if save_img and save_path is not None:\n        save_image(frame, save_path)\n\n    return frame","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-15T18:11:18.491179Z","iopub.execute_input":"2025-07-15T18:11:18.491497Z","iopub.status.idle":"2025-07-15T18:11:18.500062Z","shell.execute_reply.started":"2025-07-15T18:11:18.491469Z","shell.execute_reply":"2025-07-15T18:11:18.499405Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"class BaseModel:\n    def __init__(self):\n        self._load_model()\n\n    def _load_model(self) -> None:\n        raise NotImplementedError(\"You must implement `_load_model`.\")\n\n    def _predict(self, video_path: str) -> str:\n        raise NotImplementedError(\"You must implement `_predict`.\")\n\n    def predict(self, video_path: str) -> str:\n        return self._predict(video_path) # Changed to return strings instead of indexes\n\n    def generate_submission(self, eval_dir: str) -> None:\n\n        output_path: str = \"submission.csv\"\n        filenames: List[str] = sorted(os.listdir(eval_dir))\n        submission = []\n        for filename in filenames:\n            if filename.endswith(\".mp4\"):\n                video_path = os.path.join(eval_dir, filename)\n                prediction = self.predict(video_path)\n                submission.append((filename.split(\".\")[0], f\"{prediction}\"))  # Label in single quotes\n        \n        submission_df = pd.DataFrame(submission, columns=[\"Filename\", \"Species\"])\n        submission_df.to_csv(output_path, index=False)\n        print(f\"✅ Submission file saved to {output_path}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-15T18:11:22.529396Z","iopub.execute_input":"2025-07-15T18:11:22.530054Z","iopub.status.idle":"2025-07-15T18:11:22.536020Z","shell.execute_reply.started":"2025-07-15T18:11:22.530031Z","shell.execute_reply":"2025-07-15T18:11:22.535279Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🧠 Megadetector - SpeciesNet Ensemble","metadata":{}},{"cell_type":"code","source":"class CustomModel(BaseModel):\n\n    def _load_model(self) -> None:\n        \"\"\"\n        Loads the model from a public dataset previously uploaded.\n        This guarantees that the model runs OFFLINE\n        \"\"\"\n        os.makedirs('/kaggle/working/model', exist_ok=True)\n        species_net_files = os.listdir('/kaggle/input/cupybara-ariel-malowany-model')\n        for file in species_net_files:\n          shutil.copyfile(f\"\"\"/kaggle/input/cupybara-ariel-malowany-model/{file}\"\"\", f\"\"\"/kaggle/working/model/{file}\"\"\")\n        self.custom_megadetector_model = run_detector.load_detector('/kaggle/working/model/md_v5a.0.0.pt')\n        self.species_net = SpeciesNet('/kaggle/working/model')\n        with open('/kaggle/working/model/species_dict.json', 'r') as species_dict:\n          self.species_dict = json.load(species_dict)\n\n    def _detect_image_objects(self, frame, threshold=0.2, model = None):\n          if model is None:\n            model = self.custom_megadetector_model\n            result = model.generate_detections_one_image(frame)\n            detections = result.get('detections', [])\n            filtered = {str(i): {\"category\": d[\"category\"], \"confidence\": d[\"conf\"], \"bbox\": d[\"bbox\"]}\n                              for i, d in enumerate(detections) if d[\"conf\"] > threshold}\n            return filtered\n\n    def _retrieve_prediction(self, img, frame, obj):\n        try:\n            predictions_dict = self.species_net.classify(\n                filepaths=[f'/kaggle/working/cropped_images/{img}/{img}_{frame}_{obj}.jpg'],\n                country = ['ARG', 'BRA', 'PRY', 'URY']\n            )\n    \n            preds = predictions_dict[\"predictions\"][0]\n            if \"classifications\" not in preds:\n                return \"unknown\", 0.0\n    \n            scores = preds[\"classifications\"][\"scores\"]\n            classes = preds[\"classifications\"][\"classes\"]\n    \n            return [(x, y) for x, y in zip(classes, scores)]\n    \n        except Exception as e:\n            print(f\"Prediction failed for {img}_{frame}_{obj}: {e}\")\n            return \"error\", 0.0\n\n    def _crop_and_save_image(self, array, detection_metadata, file_name, frame_number,\n                            return_dict=False, save_dir='/kaggle/working/cropped_images', full_image = False):\n        height, width = array.shape[:2]\n        for idx_str, obj in detection_metadata.items():\n            bbx = obj[\"bbox\"]\n            x1 = int(bbx[0] * width)\n            y1 = int(bbx[1] * height)\n            x2 = int((bbx[0] + bbx[2]) * width)\n            y2 = int((bbx[1] + bbx[3]) * height)\n    \n            crop_img = array[y1:y2, x1:x2]\n            append = f\"{frame_number}_{idx_str}\"\n            if full_image:\n              to_save = array\n            else:\n              to_save = crop_img\n            save_image(to_save, file_name, append=append, save_dir=save_dir)\n    \n            predictions = self._retrieve_prediction(file_name, str(frame_number), idx_str)\n            if return_dict:\n                obj[\"pred_class\"] = predictions\n    \n        if return_dict:\n            return detection_metadata\n\n    def _detect_and_predict_image(self, file_path, open_video_data = None, steps=10, find_n_frames = 3, threshold=0.5, not_indexes = None, save_dir_path='/kaggle/working/cropped_images'):\n        if open_video_data is not None:\n          cap_obj, frame_count = open_video_data\n        else: \n          cap_obj, frame_count = open_video(file_path)\n        if not_indexes is None:\n            indexes = np.linspace(0, frame_count - 1, steps, dtype=int)\n            iterated_frames = indexes.tolist()\n        else:\n            frame_seq = list(range(frame_count - 1))\n            frame_seq = list(set(frame_seq) - set(not_indexes))\n            frame_idx = list(np.linspace(0, len(frame_seq) - 1, steps, dtype=int))\n            indexes = [frame_seq[i] for i in frame_idx]\n            iterated_frames = not_indexes + indexes\n        image_yolo_metadata = {}\n        video_predictions = {}\n        file_name = os.path.basename(file_path)\n        save_file_name = file_name.replace('.mp4', '')\n    \n        frames_with_objects = 0\n        i = 0\n        category = []\n        while i < len(indexes) and frames_with_objects < find_n_frames:\n            frame_num = indexes[i]\n            frame = extract_frame(cap_obj, frame_num)\n            if frame is None:\n                continue\n            \n            detection_metadata = self._detect_image_objects(frame, threshold = threshold)\n            if detection_metadata:\n                frames_with_objects += 1\n                # Update detection_metadata with predictions\n                detection_metadata = self._crop_and_save_image(\n                    frame, detection_metadata, save_file_name, frame_num, return_dict=True, save_dir = save_dir_path\n                )\n                image_yolo_metadata[str(frame_num)] = detection_metadata\n                for obj in list(detection_metadata.keys()):\n                    obj_metadata = detection_metadata.get(obj)\n                    video_pred = obj_metadata[\"pred_class\"]\n                    category.append(int(obj_metadata[\"category\"]))\n                    found_classes = video_predictions.keys()\n                    for c, s in video_pred:\n                      if c not in found_classes:\n                          video_predictions[c] = s\n                      max_score = video_predictions[c]\n                      if s > max_score:\n                        video_predictions[c] = s\n            i +=1\n        if frames_with_objects > 0:\n          image_yolo_metadata[\"category\"] = np.mean(category)\n        image_yolo_metadata[\"frames_with_objects\"] = frames_with_objects\n        image_yolo_metadata[\"iterated_frames\"] = iterated_frames\n        image_yolo_metadata[\"video_predictions\"] = dict(sorted(video_predictions.items(), key=lambda item: item[1], reverse = True))\n        save_json(image_yolo_metadata, save_file_name, save_dir_path)\n    \n        return image_yolo_metadata\n\n    def _species_net_to_cupybara(self, yolo_metadata, species_dict = None):\n        if species_dict is None:\n          species_dict = self.species_dict\n        yolo_dict = yolo_metadata\n        yolo_keys = yolo_dict.keys()\n        video_predictions = yolo_dict[\"video_predictions\"]\n        species_list = species_dict.keys()\n        mapped_predicted_species = {}\n        for pred_class, score in video_predictions.items():\n            for species in species_list:\n                if species in pred_class:\n                    label = species_dict[species]\n                    if label not in mapped_predicted_species.keys() and score >= 0.05:\n                        mapped_predicted_species[label] = score\n                    if score >= 0.05 and mapped_predicted_species[label] < score:\n                        mapped_predicted_species[label] = score\n        return mapped_predicted_species\n\n    def _final_predict(self, yolo_metadata, speciesnet_preds):\n        frames_with_objects = yolo_metadata.get('frames_with_objects', 0)\n        md_category = yolo_metadata.get('category')\n        prediction = 'no_animal'\n    \n        if frames_with_objects == 0:\n            return prediction \n    \n        predicted_classes = list(speciesnet_preds.keys())\n        unk_score = speciesnet_preds.get('unknown_animal', 0)\n        no_unk_preds = [cls for cls in predicted_classes if cls != 'unknown_animal']\n    \n        if not no_unk_preds:\n            if md_category == 1:\n                return 'unknown_animal'\n            elif md_category == 2:\n                return 'human'\n            return prediction  \n    \n        top_no_unk_class = no_unk_preds[0]\n        top_no_unk_score = speciesnet_preds.get(top_no_unk_class, 0)\n    \n        all_no_unk_scores = [speciesnet_preds[cls] for cls in no_unk_preds]\n        confounded_threshold = np.mean(all_no_unk_scores) + 2.5 * np.std(all_no_unk_scores)\n    \n        bird_like = {'dusky_legged_guan', 'bird'}\n        confounded_birds = bird_like | {'squirrel'}\n        birds_in_preds = [cls for cls in no_unk_preds if cls in bird_like]\n        confounded_in_preds = [cls for cls in no_unk_preds if cls in confounded_birds]\n    \n        if len(predicted_classes) == 1:\n            prediction = predicted_classes[0]\n        elif top_no_unk_score > unk_score and top_no_unk_score >= 0.25:\n            prediction = top_no_unk_class\n        elif len(top_no_unk_class) == 1 and top_no_unk_score >= 0.10:\n            prediction = top_no_unk_class\n        elif set(no_unk_preds) == bird_like:\n            prediction = 'dusky_legged_guan'\n        elif set(no_unk_preds) == confounded_birds:\n            prediction = 'bird'\n        elif unk_score > top_no_unk_score:\n            if top_no_unk_score > 0.50 or top_no_unk_score > confounded_threshold:\n                prediction = top_no_unk_class\n            else:\n                prediction = 'unknown_animal'\n        else:\n            prediction = top_no_unk_class \n    \n        if prediction in {'domestic_cat', 'fox', 'squirrel'}:\n            prediction = 'unknown_animal'\n    \n        return prediction\n\n    def _predict(self, video_path: str) -> int:\n        \"\"\"\n        Uses the loaded model to generate a prediction.\n        \"\"\"\n        yolo_metadata = self._detect_and_predict_image(video_path, open_video_data = None, steps = 20, find_n_frames = 5, threshold = 0.5)\n        speciesnet_preds = self._species_net_to_cupybara(yolo_metadata)\n        prediction = self._final_predict(yolo_metadata, speciesnet_preds)\n        return prediction","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-15T18:12:53.626449Z","iopub.execute_input":"2025-07-15T18:12:53.626742Z","iopub.status.idle":"2025-07-15T18:12:54.205205Z","shell.execute_reply.started":"2025-07-15T18:12:53.626722Z","shell.execute_reply":"2025-07-15T18:12:54.204598Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🚀 Run Inference and Generate Submission File","metadata":{}},{"cell_type":"code","source":"model = CustomModel()\nmodel.generate_submission(eval_dir=EVAL_DIR)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🎯 Scoring the submission\n\nAt this point, we are ready to calculate the submission's score. Because the `test` dataset is split between the public and private leaderboard, the score below might not match exactly the one shown in the public leaderboard.\n\nTo generate your final submission, click `Submit` on the side panel and ensure that your internet connection is off.","metadata":{}},{"cell_type":"code","source":"df_test = pd.read_csv(os.path.join(DATASET_ROOT, \"test.csv\")).sort_values(\"Filename\")\ndf_submission = pd.read_csv(\"submission.csv\").sort_values(\"Filename\")","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"f1_score(df_test[\"Species\"].values, df_submission[\"Species\"].values, average=\"weighted\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-07-15T14:02:05.432752Z","iopub.execute_input":"2025-07-15T14:02:05.433456Z","iopub.status.idle":"2025-07-15T14:02:05.451020Z","shell.execute_reply.started":"2025-07-15T14:02:05.433433Z","shell.execute_reply":"2025-07-15T14:02:05.450464Z"}},"outputs":[],"execution_count":null}]}