{"metadata":{"kernelspec":{"name":"python3","display_name":"Python 3","language":"python"},"language_info":{"name":"python","version":"3.11.11","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":99249,"databundleVersionId":11859885,"sourceType":"competition"},{"sourceId":11425807,"sourceType":"datasetVersion","datasetId":7154629}],"dockerImageVersionId":31012,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# 🐾 Cup-ybara Demo Submission\nThis notebook provides a fully self-contained baseline submission for the Cup Ybara challenge.\nIt defines the required `BaseModel` interface, implements a dummy `CustomModel`,\nand outputs a properly formatted `submission.csv` file.","metadata":{},"attachments":{}},{"cell_type":"code","source":"import os\nimport random\nimport json\nimport joblib\nimport shutil\nimport pandas as pd\nimport numpy as np\nfrom typing import List\n\nfrom sklearn.metrics import f1_score","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-15T19:50:12.376735Z","iopub.execute_input":"2025-04-15T19:50:12.377057Z","iopub.status.idle":"2025-04-15T19:50:12.381377Z","shell.execute_reply.started":"2025-04-15T19:50:12.377018Z","shell.execute_reply":"2025-04-15T19:50:12.380579Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"LABELS = [\"'armadillo'\", \"'bird'\", \"'capybara'\", \"'cow'\", \"'dusky_legged_guan'\",\n    \"'gray_brocket'\", \"'hare'\", \"'human'\", \"'insect'\", \"'margay'\", \"'no_animal'\",\n    \"'skunk'\", \"'unknown_animal'\", \"'wild_boar'\"]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-15T19:50:13.616539Z","iopub.execute_input":"2025-04-15T19:50:13.617144Z","iopub.status.idle":"2025-04-15T19:50:13.621381Z","shell.execute_reply.started":"2025-04-15T19:50:13.617117Z","shell.execute_reply":"2025-04-15T19:50:13.620639Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"!ls /kaggle/input/cupybara/dataset/","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-15T19:50:13.817322Z","iopub.execute_input":"2025-04-15T19:50:13.817665Z","iopub.status.idle":"2025-04-15T19:50:13.946976Z","shell.execute_reply.started":"2025-04-15T19:50:13.817642Z","shell.execute_reply":"2025-04-15T19:50:13.945933Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🏋️ About Train Mode\n\nThe boolean variable TRAIN is used to determine if the portion of the trainng code should be run or not. It will always be set to `False` when running the submission manually. Ensure that if a lengthy part of your notebook is used for training, that is toggled away by this variable","metadata":{}},{"cell_type":"code","source":"DATASET_ROOT = \"/kaggle/input/cupybara/dataset/\"\nEVAL_DIR =  os.path.join(DATASET_ROOT,\"test\")  # Directory containing evaluation `.mp4` files\nTRAIN = True","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-15T19:50:14.477273Z","iopub.execute_input":"2025-04-15T19:50:14.478245Z","iopub.status.idle":"2025-04-15T19:50:14.482667Z","shell.execute_reply.started":"2025-04-15T19:50:14.478212Z","shell.execute_reply":"2025-04-15T19:50:14.481719Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"!rm -rf model_weights/ model_weights.zip submission.csv","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-15T19:50:15.762838Z","iopub.execute_input":"2025-04-15T19:50:15.763153Z","iopub.status.idle":"2025-04-15T19:50:15.888272Z","shell.execute_reply.started":"2025-04-15T19:50:15.763128Z","shell.execute_reply":"2025-04-15T19:50:15.887070Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from sklearn.ensemble import RandomForestClassifier\n\ndef extract_feature_from_filename(filename, max_len=16):\n    \"\"\"\n    Convert a filename string into a fixed-length vector of ASCII codes.\n    Pads with zeros if filename is shorter than max_len.\n    \"\"\"\n    ascii_vals = [ord(c) for c in filename if c.isalnum()]\n    ascii_vals = ascii_vals[:max_len]  # Truncate if too long\n    return ascii_vals + [0] * (max_len - len(ascii_vals))\n\n\nif TRAIN: \n\n    df_test = pd.read_csv(os.path.join(DATASET_ROOT, \"test.csv\"))\n    label_to_index = {label: idx for idx, label in enumerate(LABELS)}\n    \n    \n    X_test = np.array([extract_feature_from_filename(f) for f in df_test[\"Filename\"]])\n    y_test = df_test[\"Species\"].map(label_to_index)\n\n    # Create fake training set using filenames from test\n    X_train = X_test\n    y_train = y_test\n\n    model = RandomForestClassifier(n_estimators=2, max_depth=2, random_state=42)\n    model.fit(X_train, y_train)\n\n    # Save model\n    shutil.rmtree(\"model_weights/\", ignore_errors=True)\n    os.makedirs(\"model_weights\", exist_ok=True)\n    joblib.dump(model, \"model_weights/model.joblib\")\n    shutil.make_archive(\"model_weights\", 'zip', \"model_weights\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-15T19:50:16.187796Z","iopub.execute_input":"2025-04-15T19:50:16.188109Z","iopub.status.idle":"2025-04-15T19:50:16.213812Z","shell.execute_reply.started":"2025-04-15T19:50:16.188071Z","shell.execute_reply":"2025-04-15T19:50:16.212885Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🤖 Uploading Trained Model as Public Dataset\n\n1. Manually download the generated `model_weights.zip` file from the `Output` section on the side panel.\n2. Upload the `model_weights.zip` file as a new **PUBLIC** dataset using the `Input` sectino on the side panel (Upload -> New Dataset).\n3. You can refer to the uploaded dataset by its name (all lowercase and replacing spaces with \"-\"). In our example: `Native Fauna Dummy Model` becomes `native-fauna-dummy-model`\n\n> Ensure that the dataset is **Public** and the license is Attributtion 4.0 International (CC BY 4.0)\n\n> If you had already created a dataset, you can upload a new version by clicking `Open in New Tab`, and then checking for updates. ","metadata":{}},{"cell_type":"code","source":"!ls /kaggle/input/native-fauna-dummy-model","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-15T19:52:05.162111Z","iopub.execute_input":"2025-04-15T19:52:05.162447Z","iopub.status.idle":"2025-04-15T19:52:05.291950Z","shell.execute_reply.started":"2025-04-15T19:52:05.162391Z","shell.execute_reply":"2025-04-15T19:52:05.290962Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🧱 BaseModel Interface\nThis class defines the interface expected by the competition evaluators.\nYou must subclass `BaseModel` and implement `_load_model()` and `_predict()`.","metadata":{}},{"cell_type":"code","source":"class BaseModel:\n    def __init__(self):\n        self._load_model()\n\n    def _load_model(self) -> None:\n        raise NotImplementedError(\"You must implement `_load_model`.\")\n\n    def _predict(self, video_path: str) -> str:\n        raise NotImplementedError(\"You must implement `_predict`.\")\n\n    def predict(self, video_path: str) -> str:\n        return LABELS[self._predict(video_path)]\n\n    def generate_submission(self, eval_dir: str) -> None:\n\n        output_path: str = \"submission.csv\"\n        filenames: List[str] = sorted(os.listdir(eval_dir))\n        submission = []\n        for filename in filenames:\n            if filename.endswith(\".mp4\"):\n                video_path = os.path.join(eval_dir, filename)\n                prediction = self.predict(video_path)\n                submission.append((filename.split(\".\")[0], f\"{prediction}\"))  # Label in single quotes\n        \n        submission_df = pd.DataFrame(submission, columns=[\"Filename\", \"Species\"])\n        submission_df.to_csv(output_path, index=False)\n        print(f\"✅ Submission file saved to {output_path}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-15T19:52:18.549833Z","iopub.execute_input":"2025-04-15T19:52:18.550170Z","iopub.status.idle":"2025-04-15T19:52:18.558860Z","shell.execute_reply.started":"2025-04-15T19:52:18.550143Z","shell.execute_reply":"2025-04-15T19:52:18.557890Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🧠 Dummy Model\nThis dummy model returns random predictions from the valid label list.\nReplace this with your actual model and logic.","metadata":{}},{"cell_type":"code","source":"class CustomModel(BaseModel):\n\n    def _load_model(self) -> None:\n        \"\"\"\n        Loads the model from a public dataset previously uploaded.\n        This guarantees that the model runs OFFLINE\n        \"\"\"\n        self.model = joblib.load(\"/kaggle/input/native-fauna-dummy-model/model.joblib\")\n\n    def _predict(self, video_path: str) -> int:\n        \"\"\"\n        Uses the loaded model to generate a prediction.\n        \"\"\"\n        filename = os.path.basename(video_path).split(\".\")[0]\n        feature = extract_feature_from_filename(filename)\n        pred = self.model.predict([feature])\n        return pred[0]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-15T19:53:03.544009Z","iopub.execute_input":"2025-04-15T19:53:03.544295Z","iopub.status.idle":"2025-04-15T19:53:03.549819Z","shell.execute_reply.started":"2025-04-15T19:53:03.544277Z","shell.execute_reply":"2025-04-15T19:53:03.548956Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🚀 Run Inference and Generate Submission File","metadata":{}},{"cell_type":"code","source":"model = CustomModel()\nmodel.generate_submission(eval_dir=EVAL_DIR)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-15T19:53:05.210129Z","iopub.execute_input":"2025-04-15T19:53:05.210461Z","iopub.status.idle":"2025-04-15T19:53:05.388583Z","shell.execute_reply.started":"2025-04-15T19:53:05.210435Z","shell.execute_reply":"2025-04-15T19:53:05.387614Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🎯 Scoring the submission\n\nAt this point, we are ready to calculate the submission's score. Because the `test` dataset is split between the public and private leaderboard, the score below might not match exactly the one shown in the public leaderboard.\n\nTo generate your final submission, click `Submit` on the side panel and ensure that your internet connection is off.","metadata":{}},{"cell_type":"code","source":"df_test = pd.read_csv(os.path.join(DATASET_ROOT, \"test.csv\")).sort_values(\"Filename\")\ndf_submission = pd.read_csv(\"submission.csv\").sort_values(\"Filename\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-15T19:53:13.729692Z","iopub.execute_input":"2025-04-15T19:53:13.730017Z","iopub.status.idle":"2025-04-15T19:53:13.742004Z","shell.execute_reply.started":"2025-04-15T19:53:13.729996Z","shell.execute_reply":"2025-04-15T19:53:13.741142Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"f1_score(df_test[\"Species\"].values, df_submission[\"Species\"].values, average=\"weighted\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-15T19:53:15.835645Z","iopub.execute_input":"2025-04-15T19:53:15.835956Z","iopub.status.idle":"2025-04-15T19:53:15.848421Z","shell.execute_reply.started":"2025-04-15T19:53:15.835935Z","shell.execute_reply":"2025-04-15T19:53:15.847669Z"}},"outputs":[],"execution_count":null}]}