{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":117876,"databundleVersionId":14198377,"sourceType":"competition"}],"dockerImageVersionId":31234,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Robust Cotton Weed Detection with YOLOv8\n\n### Data Quality, Confidence Thresholding & Inference Lessons  \n**The 3LC Cotton Weed Detection Challenge**\n","metadata":{}},{"cell_type":"markdown","source":"## Why This Notebook\n\nThis notebook summarizes my complete workflow and lessons learned from\n**The 3LC Cotton Weed Detection Challenge**.\n\nInstead of focusing on larger models or heavy ensembling, this project\ngradually revealed a different bottleneck:\n\n> **Most performance issues were caused by data quality and inference choices,\n> not by model capacity.**\n\nThe goal of this notebook is to document:\n- what was tried,\n- what failed (and why),\n- and what finally worked in a stable and reproducible way.\n","metadata":{}},{"cell_type":"markdown","source":"## Competition Constraints & Compliance\n\nThis work strictly follows all competition constraints:\n\n- Only YOLOv8n was used (input size fixed to 640).\n- No larger models, model ensembling, or stacking were applied.\n- All improvements focus on data quality, confidence calibration,\n  and inference robustness.\n\nThe approach aligns with the competition’s goal of simulating\nreal-world, resource-constrained production environments.\n","metadata":{}},{"cell_type":"markdown","source":"## Competition Overview\n\nThe task is a standard object detection problem:\n\n- Input: a single RGB image of a cotton field\n- Output: bounding boxes and class labels for weeds\n- Evaluation metric: F1 score on the hidden test set\n\nThe dataset was obtained via:\n\n`kaggle competitions download -c the-3lc-cotton-weed-detection-challenge`\n","metadata":{}},{"cell_type":"markdown","source":"## A Critical Early Mistake: Test Set Corruption\n\nDuring early local experiments, the test set was accidentally overwritten.\n\nThis led to several confusing symptoms:\n- the number of test images silently changed (170 → 286),\n- image filenames no longer matched the official solution,\n- submissions contained empty rows or NaN values,\n- Kaggle reported image_id mismatch errors.\n\nAt first, this looked like a submission-format issue.\nThe real cause, however, was **dataset integrity loss**.\n\n**Lesson learned:**  \nAlways verify dataset integrity before trusting any metric or submission.\n","metadata":{}},{"cell_type":"markdown","source":"## Why Data Quality Became Central\n\nInitial training instability was not caused by model architecture,\nbut by subtle label issues:\n\n- missing bounding boxes,\n- incorrect box placement,\n- statistically unlikely class configurations.\n\nThis shifted the project focus from “better models” to\n**better understanding the data**.\n","metadata":{}},{"cell_type":"markdown","source":"## Automatic Label Fixing: Why It Failed\n\nSeveral heuristic scripts were tested to automatically fix labels\nat scale (hundreds of images).\n\nAlthough well-intentioned, these approaches:\n- modified many correct labels,\n- introduced additional noise,\n- and degraded validation stability.\n\nTraining became less reliable instead of better.\n\n**Conclusion:**  \nLarge-scale automatic label fixing caused more harm than benefit.\n","metadata":{}},{"cell_type":"markdown","source":"## Targeted Manual Label Fixing\n\nInstead of fixing many labels, I switched to a **risk-based inspection strategy**.\n\nOnly a very small number of samples (around 5–10 images) were manually corrected.\nThese samples were selected based on:\n- invalid or extreme bounding box geometry,\n- statistically unlikely class distributions,\n- disagreement between strong model predictions and existing labels.\n\nDespite the small number, this change had a clear positive impact\non training stability and validation performance.\n\n**Key insight:**  \nFixing a few critical samples is far more effective than fixing many blindly.\n","metadata":{}},{"cell_type":"markdown","source":"## Model Choice\n\nThe main model used throughout the project was **YOLOv8n**, trained for 80 epochs.\n\nThis lightweight architecture proved sufficient once:\n- data quality was controlled,\n- inference parameters were properly tuned.\n\nIncreasing model size did not yield gains comparable to\nimprovements from cleaner data and better inference design.\n","metadata":{}},{"cell_type":"markdown","source":"## Confidence Threshold Selection\n\nOne of the most impactful findings was the importance of\nthe confidence threshold during inference.\n\nSmall changes in the threshold caused large variations in F1 score.\nAfter systematic experiments, the threshold was fixed to:\n\n`conf = 0.25`\n\nThis value provided a good balance between recall and precision\nand significantly reduced unstable predictions.\n\n**Lesson:**  \nConfidence tuning mattered more than changing the model architecture.\n","metadata":{}},{"cell_type":"markdown","source":"## Test-Time Augmentation (TTA)\n\nTo further improve robustness, simple test-time augmentation was applied:\n\n- original image\n- horizontal flip\n- (vertical flip in some experiments)\n\nPredictions from augmented images were merged to reduce missed detections\nand smooth prediction variance.\n\nThis lightweight TTA strategy consistently outperformed\nmore complex ensemble approaches.\n","metadata":{}},{"cell_type":"markdown","source":"## Ensemble Attempts: Why Simpler Was Better\n\nSeveral ensemble strategies were tested, including:\n- combining predictions from different confidence thresholds,\n- merging multiple submissions.\n\nAlthough some local improvements were observed, these methods:\n- increased result variance,\n- reduced interpretability,\n- did not provide consistent gains on the leaderboard.\n\n**Conclusion:**  \nComplex ensembles increased instability without reliable benefits.\n","metadata":{}},{"cell_type":"markdown","source":"## Final Submission Strategy\n\nThe final submission pipeline combined:\n\n- a stable YOLOv8n model (80 epochs),\n- a fixed confidence threshold (0.25),\n- lightweight test-time augmentation,\n- careful handling of empty predictions and image_id alignment.\n\nSpecial care was taken to avoid:\n- empty submission rows,\n- mismatched image identifiers,\n- silent dataset inconsistencies.\n","metadata":{}},{"cell_type":"markdown","source":"## Key Takeaways\n\nThis project reinforced several important principles:\n\n- Data issues can dominate model performance.\n- Small, targeted fixes outperform large automatic changes.\n- Inference design is as important as training.\n- Robustness matters more than chasing marginal gains.\n","metadata":{}},{"cell_type":"markdown","source":"## Conclusion\n\nThis notebook documents a real-world object detection project\nwhere progress came not from larger models,\nbut from **better control of data and inference decisions**.\n\nThe lessons learned here are broadly applicable to applied\ncomputer vision workflows involving imperfect data and\nreal-world constraints.\n","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}}]}