{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceId":107469,"databundleVersionId":13024000,"sourceType":"competition"},{"sourceId":470538,"sourceType":"modelInstanceVersion","modelInstanceId":379586,"modelId":399493}],"dockerImageVersionId":31090,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"**# Sheyda_Asadi**\n\n# ##Step 1: Library and Model Setup\n*First, we install the necessary library, ultralytics, and import all required packages. Then, we define the paths to the pre-trained YOLO models and load them. I've renamed the model variables to make them more descriptive.*****","metadata":{}},{"cell_type":"code","source":"# Install the ultralytics library for YOLO models\n!pip install ultralytics > /dev/null\n\nimport os\nimport cv2\nimport csv\nimport random\nimport matplotlib.pyplot as plt\nfrom pathlib import Path\nfrom ultralytics import YOLO\n\n# Define the file paths for the two object detection models\nmodel_path_a = '/kaggle/input/2-top-models/pytorch/default/1/habijabii.pt'\nmodel_path_b = '/kaggle/input/2-top-models/pytorch/default/1/nadiatriki.pt'\n\n# Load the models using the YOLO class\ndetector_a = YOLO(model_path_a, verbose=False)\ndetector_b = YOLO(model_path_b, verbose=False)\n\n# Set the directory for test images\ntest_image_folder = '/kaggle/input/multi-class-object-detection-challenge/testImages/images'\ntest_image_files = [f for f in os.listdir(test_image_folder) if f.endswith(('.jpg', '.png'))]","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Step 2: Bounding Box Formatting Function\nThis helper function processes the raw output from the YOLO models. It extracts the class ID, confidence score, and bounding box coordinates, then formats them into a single string. This specific format is required for the competition's submission file. I've renamed the function and its internal variables for clarity.","metadata":{}},{"cell_type":"code","source":"def serialize_predictions(yolo_results, class_id_offset=0):\n    \"\"\"\n    Converts YOLO detection results into a formatted string for submission.\n\n    Args:\n        yolo_results: The prediction results from a single image.\n        class_id_offset: An integer to add to the class ID, used to handle different\n                         class mappings between models.\n\n    Returns:\n        A string of space-separated predictions, or an empty string if no boxes are found.\n    \"\"\"\n    detected_boxes = yolo_results.boxes\n    img_width, img_height = yolo_results.orig_shape[1], yolo_results.orig_shape[0]\n\n    if detected_boxes is None or len(detected_boxes) == 0:\n        return \"\"\n\n    prediction_strings = []\n    for box_info in detected_boxes:\n        class_id = int(box_info.cls.cpu().numpy()) + class_id_offset\n        confidence = float(box_info.conf.cpu().numpy())\n        x_center, y_center, box_width, box_height = box_info.xywh[0].cpu().numpy()\n\n        # Normalize coordinates and dimensions to be between 0 and 1\n        x_norm = x_center / img_width\n        y_norm = y_center / img_height\n        w_norm = box_width / img_width\n        h_norm = box_height / img_height\n\n        prediction_strings.append(f\"{class_id} {confidence:.6f} {x_norm:.6f} {y_norm:.6f} {w_norm:.6f} {h_norm:.6f}\")\n\n    return \" \".join(prediction_strings)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Step 3: Inference and Ensembling\nThis is the core of the script. We iterate through each test image, run both models, and then combine their predictions. The class_id_offset is crucial here to ensure the class IDs are consistent across both models for the final submission.","metadata":{}},{"cell_type":"code","source":"submission_data = []\n\nfor image_filename in test_image_files:\n    image_full_path = os.path.join(test_image_folder, image_filename)\n\n    # Run inference on both models for the same image\n    results_a = detector_a.predict(image_full_path, conf=1e-6, device=0, verbose=False)[0]\n    results_b = detector_b.predict(image_full_path, conf=1e-6, device=0, verbose=False)[0]\n\n    # Format predictions from each model\n    predictions_model_a = serialize_predictions(results_a, class_id_offset=1)\n    predictions_model_b = serialize_predictions(results_b, class_id_offset=0)\n\n    # Combine the prediction strings from both models (ensembling)\n    final_prediction_string = (predictions_model_a + \" \" + predictions_model_b).strip()\n\n    # Handle the case where no detections were made\n    if final_prediction_string == \"\":\n        final_prediction_string = \"no boxes\"\n\n    image_identifier = os.path.splitext(image_filename)[0]\n\n    # Append the combined prediction to our list\n    submission_data.append({\n        \"image_id\": image_identifier,\n        \"prediction_string\": final_prediction_string\n    })","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Step 4: Save Submission File\nAfter processing all images, we write the collected data to a submission.csv file, which is the required output format for the competition.","metadata":{}},{"cell_type":"code","source":"output_csv_path = \"submission.csv\"\nwith open(output_csv_path, 'w', newline='') as f:\n    writer = csv.DictWriter(f, fieldnames=[\"image_id\", \"prediction_string\"])\n    writer.writeheader()\n    writer.writerows(submission_data)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Step 5: Visualization\nThis section is for visual inspection and debugging. It randomly selects 10 images and displays the bounding boxes from both models, each with a different color and label prefix. This helps us see how the two models' predictions differ or complement each other.","metadata":{}}]}