{
  "id": 564579,
  "title": "Previous Videos Based Competitions' Solutions and Takeaways",
  "url": "/competitions/nexar-collision-prediction/discussion/564579",
  "author_name": "",
  "post_date": "2025-02-23T17:58:24.782762600Z",
  "votes": 15,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Below are the publicly available links to previous video‐based competitions with a brief summary of key techniques used in each top solution. I hope we can go through them to get an idea or two on how to approach this competition. Happy Kaggling!</p>\n<hr>\n<p><strong>Competition - 1: 1st and Future – Player Contact Detection</strong>  <br>\n<a href=\"https://www.kaggle.com/competitions/nfl-player-contact-detection/overview\" target=\"_blank\">Overview</a>  </p>\n<ul>\n<li><p><strong>1st Place Solution:</strong> <a href=\"https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/391635\" target=\"_blank\">Discussion Post</a>  </p>\n<ul>\n<li>3D CNN on multi-view inputs (endzone &amp; sideline) with 18-frame sampling and head masking.  </li>\n<li>Simulated tracking data overlaid on images using cv2.circle.  </li>\n<li>Extensive image augmentations and cosine LR scheduling.  </li>\n<li>Postprocessing using XGBoost on temporal neighboring predictions.</li></ul></li>\n<li><p><strong>2nd Place Solution:</strong> <a href=\"https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/391740\" target=\"_blank\">Discussion Post</a>  </p>\n<ul>\n<li>Two-stage pipeline: YOLOv5 for helmet detection with optical flow-based tracking.  </li>\n<li>Ensemble of 3D CNNs (efficientnets, transformer decoders) on cropped helmet regions.  </li>\n<li>GroupKFold validation and temporal filtering (NMS, topK filtering).  </li>\n<li>Blending stage outputs via simple averaging for robust final predictions.</li></ul></li>\n<li><p><strong>3rd Place Solution:</strong> <a href=\"https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/392182\" target=\"_blank\">Discussion Post</a>  </p>\n<ul>\n<li>Single-stage, end-to-end model per player and step using video encoder + transformer decoder.  </li>\n<li>Predicts contact for each player across overlapping time windows.  </li>\n<li>Efficient integration of temporal and tracking features with model ensembling.</li></ul></li>\n</ul>\n<hr>\n<p><strong>Competition - 2: NFL 1st and Future – Impact Detection</strong>  <br>\n<a href=\"https://www.kaggle.com/competitions/nfl-impact-detection/overview\" target=\"_blank\">Overview</a>  </p>\n<ul>\n<li><p><strong>1st Place Solution:</strong> <a href=\"https://www.kaggle.com/competitions/nfl-impact-detection/discussion/209403\" target=\"_blank\">Discussion Post</a>  </p>\n<ul>\n<li>Helmet detection using YOLOv5 (or EfficientDet) with full-resolution training.  </li>\n<li>Helmet tracking with optical flow (OpenCV/RAFT) to estimate movement.  </li>\n<li>2.5D classification using cropped helmet regions and Temporal Shift Modules (TSM).  </li>\n<li>Postprocessing to suppress duplicate detections via temporal NMS and thresholding.</li></ul></li>\n<li><p><strong>2nd Place Solution:</strong> <a href=\"https://www.kaggle.com/competitions/nfl-impact-detection/discussion/208979\" target=\"_blank\">Discussion Post</a>  </p>\n<ul>\n<li>Two-stage pipeline: YOLOv5-based helmet detection followed by 3D CNN impact classification.  </li>\n<li>GroupKFold validation and diverse augmentations (affine, flips, dropout).  </li>\n<li>Extensive postprocessing (filtering low-confidence, video NMS, topK filtering).  </li>\n<li>Ensemble blending of multiple 3D CNN models for final prediction.</li></ul></li>\n<li><p><strong>3rd Place Solution:</strong> <a href=\"https://www.kaggle.com/competitions/nfl-impact-detection/discussion/208787\" target=\"_blank\">Discussion Post</a>  </p>\n<ul>\n<li>Generate candidate impact boxes using EfficientDet.  </li>\n<li>Binary image classification on helmet crops over 9-frame sequences.  </li>\n<li>Multi-view postprocessing: adjust thresholds across views and drop similar boxes.  </li>\n<li>Ensemble of 7 EfficientDet and 18 binary classifiers to boost score.</li></ul></li>\n</ul>\n<hr>\n<p><strong>Competition - 3: The 3rd YouTube-8M Video Understanding Challenge</strong>  <br>\n<a href=\"https://www.kaggle.com/competitions/youtube8m-2019\" target=\"_blank\">Overview</a>  </p>\n<ul>\n<li><p><strong>1st Place Solution:</strong> <a href=\"https://www.kaggle.com/competitions/youtube8m-2019/discussion/112869\" target=\"_blank\">Discussion Post</a>  </p>\n<ul>\n<li>Split modeling into video-level candidate generation (using NeXtVLAD/ResNet-like models) and segment-level re-ranking.  </li>\n<li>High-recall candidate generation followed by temporal filtering (3-frame kernel).  </li>\n<li>Focus on maximizing mAP with joint training on video and segment levels.</li></ul></li>\n<li><p><strong>2nd Place Solution:</strong> <a href=\"https://www.kaggle.com/competitions/youtube8m-2019/discussion/113663\" target=\"_blank\">Discussion Post</a>  </p>\n<ul>\n<li>Pre-train base models on 2018 data; fine-tune on 2019 segments with joint CTC-Attention decoding.  </li>\n<li>Refinement inference strategy combining video-level predictions with segment-level re-ranking.  </li>\n<li>Ensemble of mixture models (NeXtVLAD, GatedDBOF, ResNet-like) with candidate filtering.</li></ul></li>\n<li><p><strong>3rd Place Solution:</strong> <a href=\"https://www.kaggle.com/competitions/youtube8m-2019/discussion/112929\" target=\"_blank\">Discussion Post</a>  </p>\n<ul>\n<li>Deep mixture model with online distillation using multiple NeXtVLAD submodels.  </li>\n<li>Candidate generation focused on top 20 topics covering 97% of positives.  </li>\n<li>Two-layer mixture network to prevent overfitting and effective ensemble averaging.</li></ul></li>\n</ul>\n<hr>\n<p><strong>Competition - 4: The 2nd YouTube-8M Video Understanding Challenge</strong>  <br>\n<a href=\"https://www.kaggle.com/competitions/youtube8m-2018\" target=\"_blank\">Overview</a>  </p>\n<ul>\n<li><p><strong>1st Place Solution Summary:</strong> <a href=\"https://www.kaggle.com/competitions/youtube8m-2018/discussion/62781\" target=\"_blank\">Discussion Post</a>  </p>\n<ul>\n<li>Ensemble of 9 submodels from 4 families: NetVLAD, DBoF, FV, and RNNs.  </li>\n<li>Multilayer distillation and 8-bit quantization for a compact (&lt;1GB) model.  </li>\n<li>Inference-time sampling and exponential moving average for model blending.</li></ul></li>\n<li><p><strong>3rd Place Solution Sharing (NeXtVLAD):</strong> <a href=\"https://www.kaggle.com/competitions/youtube8m-2018/discussion/63223\" target=\"_blank\">Discussion Post</a>  </p>\n<ul>\n<li>Introduced NeXtVLAD to decompose high-dimensional features into low-dimensional vectors with attention.  </li>\n<li>Efficient NetVLAD aggregation over time with a lightweight parameter count.  </li>\n<li>Ensemble of 3 NeXtVLAD models yielding strong GAP performance.</li></ul></li>\n</ul>\n<hr>\n<p><strong>Competition - 5: Google – American Sign Language Fingerspelling Recognition</strong>  <br>\n<a href=\"https://www.kaggle.com/competitions/asl-fingerspelling\" target=\"_blank\">Overview</a>  </p>\n<ul>\n<li><p><strong>1st Place Solution:</strong> <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/434485\" target=\"_blank\">Discussion Post</a>  </p>\n<ul>\n<li>Encoder-decoder architecture with an improved Squeezeformer encoder and 2-layer Transformer decoder.  </li>\n<li>Custom augmentations like CutMix, FingerDropout, and TimeStretch tailored for landmark data.  </li>\n<li>Cross-validation split by signer and multiple seed ensemble.  </li>\n<li>Converted from PyTorch to TensorFlow Lite for on-device inference.</li></ul></li>\n<li><p><strong>2nd Place Solution:</strong> <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/434588\" target=\"_blank\">Discussion Post</a>  </p>\n<ul>\n<li>Adopted ASR-inspired methods using joint CTC-Attention decoding on landmark sequences.  </li>\n<li>Extensive exploration of CTC, Attention, and Transducer strategies for robust decoding.  </li>\n<li>Emphasis on efficient training and inference compatible with TensorFlow Lite.</li></ul></li>\n<li><p><strong>3rd Place Solution:</strong> <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/434393\" target=\"_blank\">Discussion Post</a>  </p>\n<ul>\n<li>Ensemble of six Conv1D models with additional transformer-based models for landmark sequences.  </li>\n<li>Robust preprocessing with heavy spatial/temporal augmentations and normalization using reference points.  </li>\n<li>Lightweight Keras architectures optimized for fast TFLite inference.</li></ul></li>\n</ul>\n<hr>\n<p><strong>Competition - 6: Google – Isolated Sign Language Recognition</strong>  <br>\n<a href=\"https://www.kaggle.com/competitions/asl-signs\" target=\"_blank\">Overview</a>  </p>\n<ul>\n<li><p><strong>1st Place Solution:</strong> <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/406684\" target=\"_blank\">Discussion Post</a>  </p>\n<ul>\n<li>Hybrid 1D CNN + Transformer model to capture sequential dynamics of landmarks.  </li>\n<li>Normalize landmarks using a central reference (e.g. nose) and include motion features.  </li>\n<li>Heavy dropout and stochastic depth for regularization; converted to TensorFlow Lite.</li></ul></li>\n<li><p><strong>2nd Place Solution:</strong> <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/406306\" target=\"_blank\">Discussion Post</a>  </p>\n<ul>\n<li>EfficientNet-B0 based model with extensive augmentations and helper transformer models (BERT/DeBERTa).  </li>\n<li>Extraction of key landmark subsets and fixed-size sequence interpolation.  </li>\n<li>Ensemble without softmax for improved accuracy while meeting TFLite constraints.</li></ul></li>\n<li><p><strong>3rd Place Solution:</strong> <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/406568\" target=\"_blank\">Discussion Post</a>  </p>\n<ul>\n<li>Ensemble of six Conv1D and two transformer models focused on strong data preprocessing and hard augmentations.  </li>\n<li>Custom TTA strategies (time-based padding, frame dropping) and signer-based cross-validation.  </li>\n<li>Lightweight architectures optimized for on-device inference.</li></ul></li>\n</ul>",
  "messages": [
    {
      "id": "3132096",
      "postDate": "02/23/2025 17:58:24",
      "content": "<p>Below are the publicly available links to previous video‐based competitions with a brief summary of key techniques used in each top solution. I hope we can go through them to get an idea or two on how to approach this competition. Happy Kaggling!</p>\n<hr>\n<p><strong>Competition - 1: 1st and Future – Player Contact Detection</strong>  <br>\n<a href=\"https://www.kaggle.com/competitions/nfl-player-contact-detection/overview\" target=\"_blank\">Overview</a>  </p>\n<ul>\n<li><p><strong>1st Place Solution:</strong> <a href=\"https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/391635\" target=\"_blank\">Discussion Post</a>  </p>\n<ul>\n<li>3D CNN on multi-view inputs (endzone &amp; sideline) with 18-frame sampling and head masking.  </li>\n<li>Simulated tracking data overlaid on images using cv2.circle.  </li>\n<li>Extensive image augmentations and cosine LR scheduling.  </li>\n<li>Postprocessing using XGBoost on temporal neighboring predictions.</li></ul></li>\n<li><p><strong>2nd Place Solution:</strong> <a href=\"https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/391740\" target=\"_blank\">Discussion Post</a>  </p>\n<ul>\n<li>Two-stage pipeline: YOLOv5 for helmet detection with optical flow-based tracking.  </li>\n<li>Ensemble of 3D CNNs (efficientnets, transformer decoders) on cropped helmet regions.  </li>\n<li>GroupKFold validation and temporal filtering (NMS, topK filtering).  </li>\n<li>Blending stage outputs via simple averaging for robust final predictions.</li></ul></li>\n<li><p><strong>3rd Place Solution:</strong> <a href=\"https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/392182\" target=\"_blank\">Discussion Post</a>  </p>\n<ul>\n<li>Single-stage, end-to-end model per player and step using video encoder + transformer decoder.  </li>\n<li>Predicts contact for each player across overlapping time windows.  </li>\n<li>Efficient integration of temporal and tracking features with model ensembling.</li></ul></li>\n</ul>\n<hr>\n<p><strong>Competition - 2: NFL 1st and Future – Impact Detection</strong>  <br>\n<a href=\"https://www.kaggle.com/competitions/nfl-impact-detection/overview\" target=\"_blank\">Overview</a>  </p>\n<ul>\n<li><p><strong>1st Place Solution:</strong> <a href=\"https://www.kaggle.com/competitions/nfl-impact-detection/discussion/209403\" target=\"_blank\">Discussion Post</a>  </p>\n<ul>\n<li>Helmet detection using YOLOv5 (or EfficientDet) with full-resolution training.  </li>\n<li>Helmet tracking with optical flow (OpenCV/RAFT) to estimate movement.  </li>\n<li>2.5D classification using cropped helmet regions and Temporal Shift Modules (TSM).  </li>\n<li>Postprocessing to suppress duplicate detections via temporal NMS and thresholding.</li></ul></li>\n<li><p><strong>2nd Place Solution:</strong> <a href=\"https://www.kaggle.com/competitions/nfl-impact-detection/discussion/208979\" target=\"_blank\">Discussion Post</a>  </p>\n<ul>\n<li>Two-stage pipeline: YOLOv5-based helmet detection followed by 3D CNN impact classification.  </li>\n<li>GroupKFold validation and diverse augmentations (affine, flips, dropout).  </li>\n<li>Extensive postprocessing (filtering low-confidence, video NMS, topK filtering).  </li>\n<li>Ensemble blending of multiple 3D CNN models for final prediction.</li></ul></li>\n<li><p><strong>3rd Place Solution:</strong> <a href=\"https://www.kaggle.com/competitions/nfl-impact-detection/discussion/208787\" target=\"_blank\">Discussion Post</a>  </p>\n<ul>\n<li>Generate candidate impact boxes using EfficientDet.  </li>\n<li>Binary image classification on helmet crops over 9-frame sequences.  </li>\n<li>Multi-view postprocessing: adjust thresholds across views and drop similar boxes.  </li>\n<li>Ensemble of 7 EfficientDet and 18 binary classifiers to boost score.</li></ul></li>\n</ul>\n<hr>\n<p><strong>Competition - 3: The 3rd YouTube-8M Video Understanding Challenge</strong>  <br>\n<a href=\"https://www.kaggle.com/competitions/youtube8m-2019\" target=\"_blank\">Overview</a>  </p>\n<ul>\n<li><p><strong>1st Place Solution:</strong> <a href=\"https://www.kaggle.com/competitions/youtube8m-2019/discussion/112869\" target=\"_blank\">Discussion Post</a>  </p>\n<ul>\n<li>Split modeling into video-level candidate generation (using NeXtVLAD/ResNet-like models) and segment-level re-ranking.  </li>\n<li>High-recall candidate generation followed by temporal filtering (3-frame kernel).  </li>\n<li>Focus on maximizing mAP with joint training on video and segment levels.</li></ul></li>\n<li><p><strong>2nd Place Solution:</strong> <a href=\"https://www.kaggle.com/competitions/youtube8m-2019/discussion/113663\" target=\"_blank\">Discussion Post</a>  </p>\n<ul>\n<li>Pre-train base models on 2018 data; fine-tune on 2019 segments with joint CTC-Attention decoding.  </li>\n<li>Refinement inference strategy combining video-level predictions with segment-level re-ranking.  </li>\n<li>Ensemble of mixture models (NeXtVLAD, GatedDBOF, ResNet-like) with candidate filtering.</li></ul></li>\n<li><p><strong>3rd Place Solution:</strong> <a href=\"https://www.kaggle.com/competitions/youtube8m-2019/discussion/112929\" target=\"_blank\">Discussion Post</a>  </p>\n<ul>\n<li>Deep mixture model with online distillation using multiple NeXtVLAD submodels.  </li>\n<li>Candidate generation focused on top 20 topics covering 97% of positives.  </li>\n<li>Two-layer mixture network to prevent overfitting and effective ensemble averaging.</li></ul></li>\n</ul>\n<hr>\n<p><strong>Competition - 4: The 2nd YouTube-8M Video Understanding Challenge</strong>  <br>\n<a href=\"https://www.kaggle.com/competitions/youtube8m-2018\" target=\"_blank\">Overview</a>  </p>\n<ul>\n<li><p><strong>1st Place Solution Summary:</strong> <a href=\"https://www.kaggle.com/competitions/youtube8m-2018/discussion/62781\" target=\"_blank\">Discussion Post</a>  </p>\n<ul>\n<li>Ensemble of 9 submodels from 4 families: NetVLAD, DBoF, FV, and RNNs.  </li>\n<li>Multilayer distillation and 8-bit quantization for a compact (&lt;1GB) model.  </li>\n<li>Inference-time sampling and exponential moving average for model blending.</li></ul></li>\n<li><p><strong>3rd Place Solution Sharing (NeXtVLAD):</strong> <a href=\"https://www.kaggle.com/competitions/youtube8m-2018/discussion/63223\" target=\"_blank\">Discussion Post</a>  </p>\n<ul>\n<li>Introduced NeXtVLAD to decompose high-dimensional features into low-dimensional vectors with attention.  </li>\n<li>Efficient NetVLAD aggregation over time with a lightweight parameter count.  </li>\n<li>Ensemble of 3 NeXtVLAD models yielding strong GAP performance.</li></ul></li>\n</ul>\n<hr>\n<p><strong>Competition - 5: Google – American Sign Language Fingerspelling Recognition</strong>  <br>\n<a href=\"https://www.kaggle.com/competitions/asl-fingerspelling\" target=\"_blank\">Overview</a>  </p>\n<ul>\n<li><p><strong>1st Place Solution:</strong> <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/434485\" target=\"_blank\">Discussion Post</a>  </p>\n<ul>\n<li>Encoder-decoder architecture with an improved Squeezeformer encoder and 2-layer Transformer decoder.  </li>\n<li>Custom augmentations like CutMix, FingerDropout, and TimeStretch tailored for landmark data.  </li>\n<li>Cross-validation split by signer and multiple seed ensemble.  </li>\n<li>Converted from PyTorch to TensorFlow Lite for on-device inference.</li></ul></li>\n<li><p><strong>2nd Place Solution:</strong> <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/434588\" target=\"_blank\">Discussion Post</a>  </p>\n<ul>\n<li>Adopted ASR-inspired methods using joint CTC-Attention decoding on landmark sequences.  </li>\n<li>Extensive exploration of CTC, Attention, and Transducer strategies for robust decoding.  </li>\n<li>Emphasis on efficient training and inference compatible with TensorFlow Lite.</li></ul></li>\n<li><p><strong>3rd Place Solution:</strong> <a href=\"https://www.kaggle.com/competitions/asl-fingerspelling/discussion/434393\" target=\"_blank\">Discussion Post</a>  </p>\n<ul>\n<li>Ensemble of six Conv1D models with additional transformer-based models for landmark sequences.  </li>\n<li>Robust preprocessing with heavy spatial/temporal augmentations and normalization using reference points.  </li>\n<li>Lightweight Keras architectures optimized for fast TFLite inference.</li></ul></li>\n</ul>\n<hr>\n<p><strong>Competition - 6: Google – Isolated Sign Language Recognition</strong>  <br>\n<a href=\"https://www.kaggle.com/competitions/asl-signs\" target=\"_blank\">Overview</a>  </p>\n<ul>\n<li><p><strong>1st Place Solution:</strong> <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/406684\" target=\"_blank\">Discussion Post</a>  </p>\n<ul>\n<li>Hybrid 1D CNN + Transformer model to capture sequential dynamics of landmarks.  </li>\n<li>Normalize landmarks using a central reference (e.g. nose) and include motion features.  </li>\n<li>Heavy dropout and stochastic depth for regularization; converted to TensorFlow Lite.</li></ul></li>\n<li><p><strong>2nd Place Solution:</strong> <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/406306\" target=\"_blank\">Discussion Post</a>  </p>\n<ul>\n<li>EfficientNet-B0 based model with extensive augmentations and helper transformer models (BERT/DeBERTa).  </li>\n<li>Extraction of key landmark subsets and fixed-size sequence interpolation.  </li>\n<li>Ensemble without softmax for improved accuracy while meeting TFLite constraints.</li></ul></li>\n<li><p><strong>3rd Place Solution:</strong> <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/406568\" target=\"_blank\">Discussion Post</a>  </p>\n<ul>\n<li>Ensemble of six Conv1D and two transformer models focused on strong data preprocessing and hard augmentations.  </li>\n<li>Custom TTA strategies (time-based padding, frame dropping) and signer-based cross-validation.  </li>\n<li>Lightweight architectures optimized for on-device inference.</li></ul></li>\n</ul>",
      "rawMarkdown": "Below are the publicly available links to previous video‐based competitions with a brief summary of key techniques used in each top solution. I hope we can go through them to get an idea or two on how to approach this competition. Happy Kaggling!\n\n---\n\n**Competition - 1: 1st and Future – Player Contact Detection**  \n[Overview](https://www.kaggle.com/competitions/nfl-player-contact-detection/overview)  \n\n- **1st Place Solution:** [Discussion Post](https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/391635)  \n  - 3D CNN on multi-view inputs (endzone & sideline) with 18-frame sampling and head masking.  \n  - Simulated tracking data overlaid on images using cv2.circle.  \n  - Extensive image augmentations and cosine LR scheduling.  \n  - Postprocessing using XGBoost on temporal neighboring predictions.\n\n- **2nd Place Solution:** [Discussion Post](https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/391740)  \n  - Two-stage pipeline: YOLOv5 for helmet detection with optical flow-based tracking.  \n  - Ensemble of 3D CNNs (efficientnets, transformer decoders) on cropped helmet regions.  \n  - GroupKFold validation and temporal filtering (NMS, topK filtering).  \n  - Blending stage outputs via simple averaging for robust final predictions.\n\n- **3rd Place Solution:** [Discussion Post](https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/392182)  \n  - Single-stage, end-to-end model per player and step using video encoder + transformer decoder.  \n  - Predicts contact for each player across overlapping time windows.  \n  - Efficient integration of temporal and tracking features with model ensembling.\n\n---\n\n**Competition - 2: NFL 1st and Future – Impact Detection**  \n[Overview](https://www.kaggle.com/competitions/nfl-impact-detection/overview)  \n\n- **1st Place Solution:** [Discussion Post](https://www.kaggle.com/competitions/nfl-impact-detection/discussion/209403)  \n  - Helmet detection using YOLOv5 (or EfficientDet) with full-resolution training.  \n  - Helmet tracking with optical flow (OpenCV/RAFT) to estimate movement.  \n  - 2.5D classification using cropped helmet regions and Temporal Shift Modules (TSM).  \n  - Postprocessing to suppress duplicate detections via temporal NMS and thresholding.\n\n- **2nd Place Solution:** [Discussion Post](https://www.kaggle.com/competitions/nfl-impact-detection/discussion/208979)  \n  - Two-stage pipeline: YOLOv5-based helmet detection followed by 3D CNN impact classification.  \n  - GroupKFold validation and diverse augmentations (affine, flips, dropout).  \n  - Extensive postprocessing (filtering low-confidence, video NMS, topK filtering).  \n  - Ensemble blending of multiple 3D CNN models for final prediction.\n\n- **3rd Place Solution:** [Discussion Post](https://www.kaggle.com/competitions/nfl-impact-detection/discussion/208787)  \n  - Generate candidate impact boxes using EfficientDet.  \n  - Binary image classification on helmet crops over 9-frame sequences.  \n  - Multi-view postprocessing: adjust thresholds across views and drop similar boxes.  \n  - Ensemble of 7 EfficientDet and 18 binary classifiers to boost score.\n\n---\n\n**Competition - 3: The 3rd YouTube-8M Video Understanding Challenge**  \n[Overview](https://www.kaggle.com/competitions/youtube8m-2019)  \n\n- **1st Place Solution:** [Discussion Post](https://www.kaggle.com/competitions/youtube8m-2019/discussion/112869)  \n  - Split modeling into video-level candidate generation (using NeXtVLAD/ResNet-like models) and segment-level re-ranking.  \n  - High-recall candidate generation followed by temporal filtering (3-frame kernel).  \n  - Focus on maximizing mAP with joint training on video and segment levels.\n\n- **2nd Place Solution:** [Discussion Post](https://www.kaggle.com/competitions/youtube8m-2019/discussion/113663)  \n  - Pre-train base models on 2018 data; fine-tune on 2019 segments with joint CTC-Attention decoding.  \n  - Refinement inference strategy combining video-level predictions with segment-level re-ranking.  \n  - Ensemble of mixture models (NeXtVLAD, GatedDBOF, ResNet-like) with candidate filtering.\n\n- **3rd Place Solution:** [Discussion Post](https://www.kaggle.com/competitions/youtube8m-2019/discussion/112929)  \n  - Deep mixture model with online distillation using multiple NeXtVLAD submodels.  \n  - Candidate generation focused on top 20 topics covering 97% of positives.  \n  - Two-layer mixture network to prevent overfitting and effective ensemble averaging.\n\n---\n\n**Competition - 4: The 2nd YouTube-8M Video Understanding Challenge**  \n[Overview](https://www.kaggle.com/competitions/youtube8m-2018)  \n\n- **1st Place Solution Summary:** [Discussion Post](https://www.kaggle.com/competitions/youtube8m-2018/discussion/62781)  \n  - Ensemble of 9 submodels from 4 families: NetVLAD, DBoF, FV, and RNNs.  \n  - Multilayer distillation and 8-bit quantization for a compact (<1GB) model.  \n  - Inference-time sampling and exponential moving average for model blending.\n\n- **3rd Place Solution Sharing (NeXtVLAD):** [Discussion Post](https://www.kaggle.com/competitions/youtube8m-2018/discussion/63223)  \n  - Introduced NeXtVLAD to decompose high-dimensional features into low-dimensional vectors with attention.  \n  - Efficient NetVLAD aggregation over time with a lightweight parameter count.  \n  - Ensemble of 3 NeXtVLAD models yielding strong GAP performance.\n\n---\n\n**Competition - 5: Google – American Sign Language Fingerspelling Recognition**  \n[Overview](https://www.kaggle.com/competitions/asl-fingerspelling)  \n\n- **1st Place Solution:** [Discussion Post](https://www.kaggle.com/competitions/asl-fingerspelling/discussion/434485)  \n  - Encoder-decoder architecture with an improved Squeezeformer encoder and 2-layer Transformer decoder.  \n  - Custom augmentations like CutMix, FingerDropout, and TimeStretch tailored for landmark data.  \n  - Cross-validation split by signer and multiple seed ensemble.  \n  - Converted from PyTorch to TensorFlow Lite for on-device inference.\n\n- **2nd Place Solution:** [Discussion Post](https://www.kaggle.com/competitions/asl-fingerspelling/discussion/434588)  \n  - Adopted ASR-inspired methods using joint CTC-Attention decoding on landmark sequences.  \n  - Extensive exploration of CTC, Attention, and Transducer strategies for robust decoding.  \n  - Emphasis on efficient training and inference compatible with TensorFlow Lite.\n\n- **3rd Place Solution:** [Discussion Post](https://www.kaggle.com/competitions/asl-fingerspelling/discussion/434393)  \n  - Ensemble of six Conv1D models with additional transformer-based models for landmark sequences.  \n  - Robust preprocessing with heavy spatial/temporal augmentations and normalization using reference points.  \n  - Lightweight Keras architectures optimized for fast TFLite inference.\n\n---\n\n**Competition - 6: Google – Isolated Sign Language Recognition**  \n[Overview](https://www.kaggle.com/competitions/asl-signs)  \n\n- **1st Place Solution:** [Discussion Post](https://www.kaggle.com/competitions/asl-signs/discussion/406684)  \n  - Hybrid 1D CNN + Transformer model to capture sequential dynamics of landmarks.  \n  - Normalize landmarks using a central reference (e.g. nose) and include motion features.  \n  - Heavy dropout and stochastic depth for regularization; converted to TensorFlow Lite.\n\n- **2nd Place Solution:** [Discussion Post](https://www.kaggle.com/competitions/asl-signs/discussion/406306)  \n  - EfficientNet-B0 based model with extensive augmentations and helper transformer models (BERT/DeBERTa).  \n  - Extraction of key landmark subsets and fixed-size sequence interpolation.  \n  - Ensemble without softmax for improved accuracy while meeting TFLite constraints.\n\n- **3rd Place Solution:** [Discussion Post](https://www.kaggle.com/competitions/asl-signs/discussion/406568)  \n  - Ensemble of six Conv1D and two transformer models focused on strong data preprocessing and hard augmentations.  \n  - Custom TTA strategies (time-based padding, frame dropping) and signer-based cross-validation.  \n  - Lightweight architectures optimized for on-device inference.",
      "votes": null
    },
    {
      "id": "3189399",
      "postDate": "04/29/2025 06:20:26",
      "content": "<p>kudos for making a thread for this</p>",
      "rawMarkdown": "kudos for making a thread for this",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3189399,
      "author_name": "abhyudaya456",
      "author_url": "",
      "post_date": "04/29/2025 06:20:26",
      "content": "<p>kudos for making a thread for this</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3132096": "Below are the publicly available links to previous video‐based competitions with a brief summary of key techniques used in each top solution. I hope we can go through them to get an idea or two on how to approach this competition. Happy Kaggling!\n\n---\n\n**Competition - 1: 1st and Future – Player Contact Detection**  \n[Overview](https://www.kaggle.com/competitions/nfl-player-contact-detection/overview)  \n\n- **1st Place Solution:** [Discussion Post](https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/391635)  \n  - 3D CNN on multi-view inputs (endzone & sideline) with 18-frame sampling and head masking.  \n  - Simulated tracking data overlaid on images using cv2.circle.  \n  - Extensive image augmentations and cosine LR scheduling.  \n  - Postprocessing using XGBoost on temporal neighboring predictions.\n\n- **2nd Place Solution:** [Discussion Post](https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/391740)  \n  - Two-stage pipeline: YOLOv5 for helmet detection with optical flow-based tracking.  \n  - Ensemble of 3D CNNs (efficientnets, transformer decoders) on cropped helmet regions.  \n  - GroupKFold validation and temporal filtering (NMS, topK filtering).  \n  - Blending stage outputs via simple averaging for robust final predictions.\n\n- **3rd Place Solution:** [Discussion Post](https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/392182)  \n  - Single-stage, end-to-end model per player and step using video encoder + transformer decoder.  \n  - Predicts contact for each player across overlapping time windows.  \n  - Efficient integration of temporal and tracking features with model ensembling.\n\n---\n\n**Competition - 2: NFL 1st and Future – Impact Detection**  \n[Overview](https://www.kaggle.com/competitions/nfl-impact-detection/overview)  \n\n- **1st Place Solution:** [Discussion Post](https://www.kaggle.com/competitions/nfl-impact-detection/discussion/209403)  \n  - Helmet detection using YOLOv5 (or EfficientDet) with full-resolution training.  \n  - Helmet tracking with optical flow (OpenCV/RAFT) to estimate movement.  \n  - 2.5D classification using cropped helmet regions and Temporal Shift Modules (TSM).  \n  - Postprocessing to suppress duplicate detections via temporal NMS and thresholding.\n\n- **2nd Place Solution:** [Discussion Post](https://www.kaggle.com/competitions/nfl-impact-detection/discussion/208979)  \n  - Two-stage pipeline: YOLOv5-based helmet detection followed by 3D CNN impact classification.  \n  - GroupKFold validation and diverse augmentations (affine, flips, dropout).  \n  - Extensive postprocessing (filtering low-confidence, video NMS, topK filtering).  \n  - Ensemble blending of multiple 3D CNN models for final prediction.\n\n- **3rd Place Solution:** [Discussion Post](https://www.kaggle.com/competitions/nfl-impact-detection/discussion/208787)  \n  - Generate candidate impact boxes using EfficientDet.  \n  - Binary image classification on helmet crops over 9-frame sequences.  \n  - Multi-view postprocessing: adjust thresholds across views and drop similar boxes.  \n  - Ensemble of 7 EfficientDet and 18 binary classifiers to boost score.\n\n---\n\n**Competition - 3: The 3rd YouTube-8M Video Understanding Challenge**  \n[Overview](https://www.kaggle.com/competitions/youtube8m-2019)  \n\n- **1st Place Solution:** [Discussion Post](https://www.kaggle.com/competitions/youtube8m-2019/discussion/112869)  \n  - Split modeling into video-level candidate generation (using NeXtVLAD/ResNet-like models) and segment-level re-ranking.  \n  - High-recall candidate generation followed by temporal filtering (3-frame kernel).  \n  - Focus on maximizing mAP with joint training on video and segment levels.\n\n- **2nd Place Solution:** [Discussion Post](https://www.kaggle.com/competitions/youtube8m-2019/discussion/113663)  \n  - Pre-train base models on 2018 data; fine-tune on 2019 segments with joint CTC-Attention decoding.  \n  - Refinement inference strategy combining video-level predictions with segment-level re-ranking.  \n  - Ensemble of mixture models (NeXtVLAD, GatedDBOF, ResNet-like) with candidate filtering.\n\n- **3rd Place Solution:** [Discussion Post](https://www.kaggle.com/competitions/youtube8m-2019/discussion/112929)  \n  - Deep mixture model with online distillation using multiple NeXtVLAD submodels.  \n  - Candidate generation focused on top 20 topics covering 97% of positives.  \n  - Two-layer mixture network to prevent overfitting and effective ensemble averaging.\n\n---\n\n**Competition - 4: The 2nd YouTube-8M Video Understanding Challenge**  \n[Overview](https://www.kaggle.com/competitions/youtube8m-2018)  \n\n- **1st Place Solution Summary:** [Discussion Post](https://www.kaggle.com/competitions/youtube8m-2018/discussion/62781)  \n  - Ensemble of 9 submodels from 4 families: NetVLAD, DBoF, FV, and RNNs.  \n  - Multilayer distillation and 8-bit quantization for a compact (<1GB) model.  \n  - Inference-time sampling and exponential moving average for model blending.\n\n- **3rd Place Solution Sharing (NeXtVLAD):** [Discussion Post](https://www.kaggle.com/competitions/youtube8m-2018/discussion/63223)  \n  - Introduced NeXtVLAD to decompose high-dimensional features into low-dimensional vectors with attention.  \n  - Efficient NetVLAD aggregation over time with a lightweight parameter count.  \n  - Ensemble of 3 NeXtVLAD models yielding strong GAP performance.\n\n---\n\n**Competition - 5: Google – American Sign Language Fingerspelling Recognition**  \n[Overview](https://www.kaggle.com/competitions/asl-fingerspelling)  \n\n- **1st Place Solution:** [Discussion Post](https://www.kaggle.com/competitions/asl-fingerspelling/discussion/434485)  \n  - Encoder-decoder architecture with an improved Squeezeformer encoder and 2-layer Transformer decoder.  \n  - Custom augmentations like CutMix, FingerDropout, and TimeStretch tailored for landmark data.  \n  - Cross-validation split by signer and multiple seed ensemble.  \n  - Converted from PyTorch to TensorFlow Lite for on-device inference.\n\n- **2nd Place Solution:** [Discussion Post](https://www.kaggle.com/competitions/asl-fingerspelling/discussion/434588)  \n  - Adopted ASR-inspired methods using joint CTC-Attention decoding on landmark sequences.  \n  - Extensive exploration of CTC, Attention, and Transducer strategies for robust decoding.  \n  - Emphasis on efficient training and inference compatible with TensorFlow Lite.\n\n- **3rd Place Solution:** [Discussion Post](https://www.kaggle.com/competitions/asl-fingerspelling/discussion/434393)  \n  - Ensemble of six Conv1D models with additional transformer-based models for landmark sequences.  \n  - Robust preprocessing with heavy spatial/temporal augmentations and normalization using reference points.  \n  - Lightweight Keras architectures optimized for fast TFLite inference.\n\n---\n\n**Competition - 6: Google – Isolated Sign Language Recognition**  \n[Overview](https://www.kaggle.com/competitions/asl-signs)  \n\n- **1st Place Solution:** [Discussion Post](https://www.kaggle.com/competitions/asl-signs/discussion/406684)  \n  - Hybrid 1D CNN + Transformer model to capture sequential dynamics of landmarks.  \n  - Normalize landmarks using a central reference (e.g. nose) and include motion features.  \n  - Heavy dropout and stochastic depth for regularization; converted to TensorFlow Lite.\n\n- **2nd Place Solution:** [Discussion Post](https://www.kaggle.com/competitions/asl-signs/discussion/406306)  \n  - EfficientNet-B0 based model with extensive augmentations and helper transformer models (BERT/DeBERTa).  \n  - Extraction of key landmark subsets and fixed-size sequence interpolation.  \n  - Ensemble without softmax for improved accuracy while meeting TFLite constraints.\n\n- **3rd Place Solution:** [Discussion Post](https://www.kaggle.com/competitions/asl-signs/discussion/406568)  \n  - Ensemble of six Conv1D and two transformer models focused on strong data preprocessing and hard augmentations.  \n  - Custom TTA strategies (time-based padding, frame dropping) and signer-based cross-validation.  \n  - Lightweight architectures optimized for on-device inference.",
    "3189399": "kudos for making a thread for this"
  },
  "source": "meta"
}