{
  "id": 576289,
  "title": "2nd Place solution(public&private lb):VideoMAEv2-based Vehicle Collision Prediction",
  "url": "/competitions/nexar-collision-prediction/discussion/576289",
  "author_name": "Wang Zhiyao (王致尧)",
  "post_date": "2025-05-04T06:26:49.694000",
  "votes": 13,
  "comment_count": 9,
  "views": 0,
  "content": "<h1>VideoMAE-based Vehicle Collision Prediction Solution: 2nd Place on Kaggle Public Leaderboard</h1>\n<h2>Solution Overview</h2>\n<p>Hey everyone,</p>\n<p>I'd like to share my approach for the Nexar Safe Driving Video Analysis competition. This solution uses a deep learning method based on the <strong>VideoMAEv2-giant</strong> pre-trained model to predict collision and near-miss risks in driving videos.</p>\n<p>The process was relatively straightforward, and it achieved <strong>2nd place on the public leaderboard</strong> </p>\n<h2>Why VideoMAEv2?</h2>\n<p>We framed the accident prediction task as a video classification problem: classifying videos into \"imminent collision/near-miss\" or \"normal driving\". After experimenting with several approaches, VideoMAEv2 delivered the best performance.</p>\n<ul>\n<li><p><strong>Initial Attempts</strong>: We first tried more traditional architectures like CNN+LSTM and CNN+Attention, using the same data processing pipeline. However, these models generally scored between 0.7 and 0.8, falling short of VideoMAEv2.</p></li>\n<li><p><strong>VideoMAEv2 Advantages</strong>: As a state-of-the-art pre-trained model for video understanding, VideoMAEv2 offers several key benefits:</p>\n<ol>\n<li><strong>Strong Spatio-temporal Feature Extraction</strong>: It excels at capturing the dynamic information crucial for collision prediction.</li>\n<li><strong>Large-Scale Pre-training</strong>: Pre-training on massive video datasets provides excellent generalization capabilities.</li>\n<li><strong>Action Recognition Performance</strong>: Its proven success in action recognition tasks suggests suitability for understanding dynamic events like near-misses and collisions.</li></ol></li>\n</ul>\n<p>Its strong performance on various video benchmarks further motivated its selection for this task.</p>\n<h2>Technical Implementation</h2>\n<ul>\n<li><p><strong>Label Assignment</strong></p>\n<ol>\n<li>For positive videos (collisions/near-misses), only frames preceding the event time were processed.</li></ol>\n<pre><code>  last_time &gt;= event_time:\n       \n</code></pre>\n<ol>\n<li><strong>Threshold Selection &amp; Handling of <code>alert_time</code></strong></li></ol>\n<ul>\n<li>Initially, we did not utilize the provided human-labeled <code>alert_time</code> (earliest prediction time). Instead, we labeled a window as positive if its last frame fell within <strong>1.5 seconds</strong> before the event; earlier frames were labeled negative by default.</li>\n<li>The <strong>1.5-second</strong> threshold was chosen because the evaluation metric specifically tests predictions at 0.5 s, 1.0 s, and 1.5 s TTAs. Among these, the 1.5-second condition generally covers most evaluation scenarios, allowing us to directly align our labeling strategy with the main evaluation criterion.</li>\n<li>Later, we conducted a \"cleaner\" experiment where we strictly adhered to the official labeling standard (marking windows that ended before the human-labeled <code>alert_time</code> as negative). Interestingly, this approach slightly <strong>decreased validation accuracy and leaderboard score</strong>, suggesting that our original, looser labeling method might have inadvertently captured useful early signals beneficial for model generalization.</li>\n<li>Due to limited computing resources and inefficient preprocessing pipelines, we couldn't perform extensive ablation studies to investigate this issue further.</li></ul>\n<ol>\n<li>For negative videos (normal driving), all windows were labeled negative.</li></ol></li>\n</ul>\n<p><strong>Important Considerations</strong>:</p>\n<ul>\n<li><strong>Ignoring <code>alert_time</code></strong>: The competition provided <code>alert_time</code> (human-annotated earliest prediction time). We didn't directly use it, hypothesizing that a deep learning model might capture longer-term dependencies better than human annotation. Instead, we opted for the fixed 1.5-second threshold before the event. Frankly, this lacks empirical rigor, as we didn't run experiments comparing different thresholds. This is a key area for future improvement.</li>\n<li><strong>Data Balancing Details</strong>: This process yielded 4,258 positive samples, and we matched this count with randomly selected negative samples. Our sliding window approach generated far more negative than positive windows initially. This aggressive undersampling, while simple, likely discarded many potentially informative or hard-negative examples, possibly hindering the model's ability to learn fine-grained distinctions. Exploring techniques like Focal Loss could be beneficial.</li>\n</ul>\n<pre><code>\n is_positive_video:\n     last_time &gt;= event_time:\n          \n    label =   (event_time -  &lt;= last_time &lt; event_time)  \n:\n    label = \n</code></pre>\n<h3>Model Architecture</h3>\n<p>The model architecture is straightforward:</p>\n<ul>\n<li><p><strong>Backbone</strong>: VideoMAEv2-giant pre-trained model.<br>\nIt’s already <em>giant</em>—in name and in GPU appetite—so we figured it was doing more than enough of the heavy lifting.</p></li>\n<li><p><strong>Classification Head</strong>: A simple linear layer followed by Dropout (p=0.1).<br>\nWe considered using a transformer-based or attention pooling classification head to further model temporal relationships between tokens. However, initial experiments did not show significant performance gains. Given the trade-off between computational cost and model simplicity—and the fact that the backbone was already doing the equivalent of a full-time job—we ultimately chose a simple linear head as the classifier.</p></li>\n<li><p><strong>Training Technique</strong>: Temperature scaling (T=2.0) applied to logits during training for smoother probability distributions.</p></li>\n</ul>\n<h3>Inference Method</h3>\n<p>For the test set, we used a simple yet effective inference strategy:</p>\n<ul>\n<li><strong>Extract only the last 2 seconds</strong> of each test video.</li>\n<li>Feed this 2-second clip into the trained model for prediction.</li>\n</ul>\n<p>This approach relies on the intuition that the most critical information for predicting an imminent event is likely concentrated towards the end of the video clip (especially since test videos were trimmed). This simplification performed well, likely because:</p>\n<ol>\n<li>Positive samples (trimmed near events) have strong signals in the final moments.</li>\n<li>Negative samples (normal driving) shouldn't have strong signals regardless of the segment.</li>\n</ol>\n<p>While testing various window lengths would be ideal, this focused approach reduced computational load and aligned with the intuition that later predictions tend to be more accurate for imminent events. It yielded good results despite its simplicity.</p>\n<h3>Training Environment and Strategy</h3>\n<ul>\n<li><strong>Hardware</strong>: Trained on an H20 server (96GB VRAM). Training is also feasible on an A100 (40GB). Smaller GPUs might require using smaller VideoMAE variants.</li>\n<li><strong>Reproducibility</strong>: Fixed random seed (<code>seed=42</code>) for Python, NumPy, and PyTorch (including CUDA) ensures deterministic runs.</li>\n<li><strong>Key Hyperparameters</strong>:<ul>\n<li>Optimizer: AdamW</li>\n<li>Learning Rate: 1e-5 (with Cosine Annealing scheduler)</li>\n<li>Batch Size: 3 (effective batch size 24 due to gradient accumulation over 8 steps)</li>\n<li>Epochs: 10 (saving the best model based on validation accuracy)</li></ul></li>\n</ul>\n<h2>Limitations and Room for Improvement</h2>\n<p>Given academic and resource constraints, this solution has several limitations and areas for future work:</p>\n<ol>\n<li><strong>Fixed Sampling Strategy</strong>: The 2-second window (16 frames) was chosen somewhat arbitrarily. Systematic experiments comparing different window lengths and frame rates are needed.</li>\n<li><strong>Validation Mismatch</strong>: Our local validation setup didn't perfectly mirror the competition's evaluation metric (Mean Average Precision over time-to-accident thresholds). This might mean the submitted model wasn't the true optimal one based on the leaderboard metric.</li>\n<li><strong>Single Model Submission</strong>: Although the prediction script includes ensemble capabilities, the final submission relied on a single model, missing potential gains from ensembling.</li>\n<li><strong>Limited Data Augmentation</strong>: More sophisticated video-specific augmentations (e.g., temporal jittering, different cropping strategies) could be explored.</li>\n<li><strong>Crude Class Imbalance Handling</strong>: As noted, simple undersampling might not be optimal. Exploring methods like Focal Loss or more intelligent negative mining could improve performance.</li>\n</ol>\n<h2>Code and Model Weights</h2>\n<p>All code has been uploaded to the GitHub repository:<br>\n<a href=\"https://github.com/wzyfromhust/nexer-solution\" target=\"_blank\">https://github.com/wzyfromhust/nexer-solution</a></p>\n<p><strong>Note</strong>: The final model weights are quite large and will be uploaded later when feasible. However, the code includes the fixed random seed (<code>seed=42</code>), so reproducing the results by following the data processing and training scripts should be relatively straightforward.</p>\n<hr>\n<p>Achieving a decent rank with what feels like a somewhat unrefined and insufficiently experimented approach suggests either a good dose of luck, or perhaps, significant untapped potential in the VideoMAEv2 model for this specific task.</p>",
  "messages": [
    {
      "id": 3193266,
      "postDate": "2025-05-04T06:26:49.693Z",
      "content": "<h1>VideoMAE-based Vehicle Collision Prediction Solution: 2nd Place on Kaggle Public Leaderboard</h1>\n<h2>Solution Overview</h2>\n<p>Hey everyone,</p>\n<p>I'd like to share my approach for the Nexar Safe Driving Video Analysis competition. This solution uses a deep learning method based on the <strong>VideoMAEv2-giant</strong> pre-trained model to predict collision and near-miss risks in driving videos.</p>\n<p>The process was relatively straightforward, and it achieved <strong>2nd place on the public leaderboard</strong> </p>\n<h2>Why VideoMAEv2?</h2>\n<p>We framed the accident prediction task as a video classification problem: classifying videos into \"imminent collision/near-miss\" or \"normal driving\". After experimenting with several approaches, VideoMAEv2 delivered the best performance.</p>\n<ul>\n<li><p><strong>Initial Attempts</strong>: We first tried more traditional architectures like CNN+LSTM and CNN+Attention, using the same data processing pipeline. However, these models generally scored between 0.7 and 0.8, falling short of VideoMAEv2.</p></li>\n<li><p><strong>VideoMAEv2 Advantages</strong>: As a state-of-the-art pre-trained model for video understanding, VideoMAEv2 offers several key benefits:</p>\n<ol>\n<li><strong>Strong Spatio-temporal Feature Extraction</strong>: It excels at capturing the dynamic information crucial for collision prediction.</li>\n<li><strong>Large-Scale Pre-training</strong>: Pre-training on massive video datasets provides excellent generalization capabilities.</li>\n<li><strong>Action Recognition Performance</strong>: Its proven success in action recognition tasks suggests suitability for understanding dynamic events like near-misses and collisions.</li></ol></li>\n</ul>\n<p>Its strong performance on various video benchmarks further motivated its selection for this task.</p>\n<h2>Technical Implementation</h2>\n<ul>\n<li><p><strong>Label Assignment</strong></p>\n<ol>\n<li>For positive videos (collisions/near-misses), only frames preceding the event time were processed.</li></ol>\n<pre><code>  last_time &gt;= event_time:\n       \n</code></pre>\n<ol>\n<li><strong>Threshold Selection &amp; Handling of <code>alert_time</code></strong></li></ol>\n<ul>\n<li>Initially, we did not utilize the provided human-labeled <code>alert_time</code> (earliest prediction time). Instead, we labeled a window as positive if its last frame fell within <strong>1.5 seconds</strong> before the event; earlier frames were labeled negative by default.</li>\n<li>The <strong>1.5-second</strong> threshold was chosen because the evaluation metric specifically tests predictions at 0.5 s, 1.0 s, and 1.5 s TTAs. Among these, the 1.5-second condition generally covers most evaluation scenarios, allowing us to directly align our labeling strategy with the main evaluation criterion.</li>\n<li>Later, we conducted a \"cleaner\" experiment where we strictly adhered to the official labeling standard (marking windows that ended before the human-labeled <code>alert_time</code> as negative). Interestingly, this approach slightly <strong>decreased validation accuracy and leaderboard score</strong>, suggesting that our original, looser labeling method might have inadvertently captured useful early signals beneficial for model generalization.</li>\n<li>Due to limited computing resources and inefficient preprocessing pipelines, we couldn't perform extensive ablation studies to investigate this issue further.</li></ul>\n<ol>\n<li>For negative videos (normal driving), all windows were labeled negative.</li></ol></li>\n</ul>\n<p><strong>Important Considerations</strong>:</p>\n<ul>\n<li><strong>Ignoring <code>alert_time</code></strong>: The competition provided <code>alert_time</code> (human-annotated earliest prediction time). We didn't directly use it, hypothesizing that a deep learning model might capture longer-term dependencies better than human annotation. Instead, we opted for the fixed 1.5-second threshold before the event. Frankly, this lacks empirical rigor, as we didn't run experiments comparing different thresholds. This is a key area for future improvement.</li>\n<li><strong>Data Balancing Details</strong>: This process yielded 4,258 positive samples, and we matched this count with randomly selected negative samples. Our sliding window approach generated far more negative than positive windows initially. This aggressive undersampling, while simple, likely discarded many potentially informative or hard-negative examples, possibly hindering the model's ability to learn fine-grained distinctions. Exploring techniques like Focal Loss could be beneficial.</li>\n</ul>\n<pre><code>\n is_positive_video:\n     last_time &gt;= event_time:\n          \n    label =   (event_time -  &lt;= last_time &lt; event_time)  \n:\n    label = \n</code></pre>\n<h3>Model Architecture</h3>\n<p>The model architecture is straightforward:</p>\n<ul>\n<li><p><strong>Backbone</strong>: VideoMAEv2-giant pre-trained model.<br>\nIt’s already <em>giant</em>—in name and in GPU appetite—so we figured it was doing more than enough of the heavy lifting.</p></li>\n<li><p><strong>Classification Head</strong>: A simple linear layer followed by Dropout (p=0.1).<br>\nWe considered using a transformer-based or attention pooling classification head to further model temporal relationships between tokens. However, initial experiments did not show significant performance gains. Given the trade-off between computational cost and model simplicity—and the fact that the backbone was already doing the equivalent of a full-time job—we ultimately chose a simple linear head as the classifier.</p></li>\n<li><p><strong>Training Technique</strong>: Temperature scaling (T=2.0) applied to logits during training for smoother probability distributions.</p></li>\n</ul>\n<h3>Inference Method</h3>\n<p>For the test set, we used a simple yet effective inference strategy:</p>\n<ul>\n<li><strong>Extract only the last 2 seconds</strong> of each test video.</li>\n<li>Feed this 2-second clip into the trained model for prediction.</li>\n</ul>\n<p>This approach relies on the intuition that the most critical information for predicting an imminent event is likely concentrated towards the end of the video clip (especially since test videos were trimmed). This simplification performed well, likely because:</p>\n<ol>\n<li>Positive samples (trimmed near events) have strong signals in the final moments.</li>\n<li>Negative samples (normal driving) shouldn't have strong signals regardless of the segment.</li>\n</ol>\n<p>While testing various window lengths would be ideal, this focused approach reduced computational load and aligned with the intuition that later predictions tend to be more accurate for imminent events. It yielded good results despite its simplicity.</p>\n<h3>Training Environment and Strategy</h3>\n<ul>\n<li><strong>Hardware</strong>: Trained on an H20 server (96GB VRAM). Training is also feasible on an A100 (40GB). Smaller GPUs might require using smaller VideoMAE variants.</li>\n<li><strong>Reproducibility</strong>: Fixed random seed (<code>seed=42</code>) for Python, NumPy, and PyTorch (including CUDA) ensures deterministic runs.</li>\n<li><strong>Key Hyperparameters</strong>:<ul>\n<li>Optimizer: AdamW</li>\n<li>Learning Rate: 1e-5 (with Cosine Annealing scheduler)</li>\n<li>Batch Size: 3 (effective batch size 24 due to gradient accumulation over 8 steps)</li>\n<li>Epochs: 10 (saving the best model based on validation accuracy)</li></ul></li>\n</ul>\n<h2>Limitations and Room for Improvement</h2>\n<p>Given academic and resource constraints, this solution has several limitations and areas for future work:</p>\n<ol>\n<li><strong>Fixed Sampling Strategy</strong>: The 2-second window (16 frames) was chosen somewhat arbitrarily. Systematic experiments comparing different window lengths and frame rates are needed.</li>\n<li><strong>Validation Mismatch</strong>: Our local validation setup didn't perfectly mirror the competition's evaluation metric (Mean Average Precision over time-to-accident thresholds). This might mean the submitted model wasn't the true optimal one based on the leaderboard metric.</li>\n<li><strong>Single Model Submission</strong>: Although the prediction script includes ensemble capabilities, the final submission relied on a single model, missing potential gains from ensembling.</li>\n<li><strong>Limited Data Augmentation</strong>: More sophisticated video-specific augmentations (e.g., temporal jittering, different cropping strategies) could be explored.</li>\n<li><strong>Crude Class Imbalance Handling</strong>: As noted, simple undersampling might not be optimal. Exploring methods like Focal Loss or more intelligent negative mining could improve performance.</li>\n</ol>\n<h2>Code and Model Weights</h2>\n<p>All code has been uploaded to the GitHub repository:<br>\n<a href=\"https://github.com/wzyfromhust/nexer-solution\" target=\"_blank\">https://github.com/wzyfromhust/nexer-solution</a></p>\n<p><strong>Note</strong>: The final model weights are quite large and will be uploaded later when feasible. However, the code includes the fixed random seed (<code>seed=42</code>), so reproducing the results by following the data processing and training scripts should be relatively straightforward.</p>\n<hr>\n<p>Achieving a decent rank with what feels like a somewhat unrefined and insufficiently experimented approach suggests either a good dose of luck, or perhaps, significant untapped potential in the VideoMAEv2 model for this specific task.</p>",
      "rawMarkdown": "# VideoMAE-based Vehicle Collision Prediction Solution: 2nd Place on Kaggle Public Leaderboard\n\n## Solution Overview\n\nHey everyone,\n\nI'd like to share my approach for the Nexar Safe Driving Video Analysis competition. This solution uses a deep learning method based on the **VideoMAEv2-giant** pre-trained model to predict collision and near-miss risks in driving videos.\n\nThe process was relatively straightforward, and it achieved **2nd place on the public leaderboard** \n\n## Why VideoMAEv2?\n\nWe framed the accident prediction task as a video classification problem: classifying videos into \"imminent collision/near-miss\" or \"normal driving\". After experimenting with several approaches, VideoMAEv2 delivered the best performance.\n\n*   **Initial Attempts**: We first tried more traditional architectures like CNN+LSTM and CNN+Attention, using the same data processing pipeline. However, these models generally scored between 0.7 and 0.8, falling short of VideoMAEv2.\n\n*   **VideoMAEv2 Advantages**: As a state-of-the-art pre-trained model for video understanding, VideoMAEv2 offers several key benefits:\n    1.  **Strong Spatio-temporal Feature Extraction**: It excels at capturing the dynamic information crucial for collision prediction.\n    2.  **Large-Scale Pre-training**: Pre-training on massive video datasets provides excellent generalization capabilities.\n    3.  **Action Recognition Performance**: Its proven success in action recognition tasks suggests suitability for understanding dynamic events like near-misses and collisions.\n\nIts strong performance on various video benchmarks further motivated its selection for this task.\n\n## Technical Implementation\n\n* **Label Assignment**\n\n  1. For positive videos (collisions/near-misses), only frames preceding the event time were processed.\n\n     ```python\n     if last_time >= event_time:\n         break  # Stop processing if past the event time\n     ```\n\n  2. **Threshold Selection & Handling of `alert_time`**\n\n     * Initially, we did not utilize the provided human-labeled `alert_time` (earliest prediction time). Instead, we labeled a window as positive if its last frame fell within **1.5 seconds** before the event; earlier frames were labeled negative by default.\n     * The **1.5-second** threshold was chosen because the evaluation metric specifically tests predictions at 0.5 s, 1.0 s, and 1.5 s TTAs. Among these, the 1.5-second condition generally covers most evaluation scenarios, allowing us to directly align our labeling strategy with the main evaluation criterion.\n     * Later, we conducted a \"cleaner\" experiment where we strictly adhered to the official labeling standard (marking windows that ended before the human-labeled `alert_time` as negative). Interestingly, this approach slightly **decreased validation accuracy and leaderboard score**, suggesting that our original, looser labeling method might have inadvertently captured useful early signals beneficial for model generalization.\n     * Due to limited computing resources and inefficient preprocessing pipelines, we couldn't perform extensive ablation studies to investigate this issue further.\n\n  3. For negative videos (normal driving), all windows were labeled negative.\n\n\n**Important Considerations**:\n\n*   **Ignoring `alert_time`**: The competition provided `alert_time` (human-annotated earliest prediction time). We didn't directly use it, hypothesizing that a deep learning model might capture longer-term dependencies better than human annotation. Instead, we opted for the fixed 1.5-second threshold before the event. Frankly, this lacks empirical rigor, as we didn't run experiments comparing different thresholds. This is a key area for future improvement.\n*   **Data Balancing Details**: This process yielded 4,258 positive samples, and we matched this count with randomly selected negative samples. Our sliding window approach generated far more negative than positive windows initially. This aggressive undersampling, while simple, likely discarded many potentially informative or hard-negative examples, possibly hindering the model's ability to learn fine-grained distinctions. Exploring techniques like Focal Loss could be beneficial.\n\n```python\n# Core label assignment logic\nif is_positive_video:\n    if last_time >= event_time:\n        break  # Stop processing if past the event time\n    label = 1 if (event_time - 1.5 <= last_time < event_time) else 0\nelse:\n    label = 0\n```\n\n### Model Architecture\n\nThe model architecture is straightforward:\n\n* **Backbone**: VideoMAEv2-giant pre-trained model.\n  It’s already *giant*—in name and in GPU appetite—so we figured it was doing more than enough of the heavy lifting.\n\n* **Classification Head**: A simple linear layer followed by Dropout (p=0.1).\n  We considered using a transformer-based or attention pooling classification head to further model temporal relationships between tokens. However, initial experiments did not show significant performance gains. Given the trade-off between computational cost and model simplicity—and the fact that the backbone was already doing the equivalent of a full-time job—we ultimately chose a simple linear head as the classifier.\n\n* **Training Technique**: Temperature scaling (T=2.0) applied to logits during training for smoother probability distributions.\n\n### Inference Method\n\nFor the test set, we used a simple yet effective inference strategy:\n\n*   **Extract only the last 2 seconds** of each test video.\n*   Feed this 2-second clip into the trained model for prediction.\n\nThis approach relies on the intuition that the most critical information for predicting an imminent event is likely concentrated towards the end of the video clip (especially since test videos were trimmed). This simplification performed well, likely because:\n\n1.  Positive samples (trimmed near events) have strong signals in the final moments.\n2.  Negative samples (normal driving) shouldn't have strong signals regardless of the segment.\n\nWhile testing various window lengths would be ideal, this focused approach reduced computational load and aligned with the intuition that later predictions tend to be more accurate for imminent events. It yielded good results despite its simplicity.\n\n### Training Environment and Strategy\n\n*   **Hardware**: Trained on an H20 server (96GB VRAM). Training is also feasible on an A100 (40GB). Smaller GPUs might require using smaller VideoMAE variants.\n*   **Reproducibility**: Fixed random seed (`seed=42`) for Python, NumPy, and PyTorch (including CUDA) ensures deterministic runs.\n*   **Key Hyperparameters**:\n    *   Optimizer: AdamW\n    *   Learning Rate: 1e-5 (with Cosine Annealing scheduler)\n    *   Batch Size: 3 (effective batch size 24 due to gradient accumulation over 8 steps)\n    *   Epochs: 10 (saving the best model based on validation accuracy)\n\n## Limitations and Room for Improvement\n\nGiven academic and resource constraints, this solution has several limitations and areas for future work:\n\n1.  **Fixed Sampling Strategy**: The 2-second window (16 frames) was chosen somewhat arbitrarily. Systematic experiments comparing different window lengths and frame rates are needed.\n2.  **Validation Mismatch**: Our local validation setup didn't perfectly mirror the competition's evaluation metric (Mean Average Precision over time-to-accident thresholds). This might mean the submitted model wasn't the true optimal one based on the leaderboard metric.\n3.  **Single Model Submission**: Although the prediction script includes ensemble capabilities, the final submission relied on a single model, missing potential gains from ensembling.\n4.  **Limited Data Augmentation**: More sophisticated video-specific augmentations (e.g., temporal jittering, different cropping strategies) could be explored.\n5.  **Crude Class Imbalance Handling**: As noted, simple undersampling might not be optimal. Exploring methods like Focal Loss or more intelligent negative mining could improve performance.\n\n## Code and Model Weights\n\nAll code has been uploaded to the GitHub repository:\n[https://github.com/wzyfromhust/nexer-solution](https://github.com/wzyfromhust/nexer-solution)\n\n**Note**: The final model weights are quite large and will be uploaded later when feasible. However, the code includes the fixed random seed (`seed=42`), so reproducing the results by following the data processing and training scripts should be relatively straightforward.\n\n---\n\nAchieving a decent rank with what feels like a somewhat unrefined and insufficiently experimented approach suggests either a good dose of luck, or perhaps, significant untapped potential in the VideoMAEv2 model for this specific task.\n",
      "votes": 13
    },
    {
      "id": 3200026,
      "postDate": "2025-05-11T22:10:25.473Z",
      "content": "<p>Very much appreciate the details of your submission.    I like the thought process and the code details to learn from.</p>",
      "rawMarkdown": "Very much appreciate the details of your submission.    I like the thought process and the code details to learn from.",
      "votes": 1
    },
    {
      "id": 3194253,
      "postDate": "2025-05-05T15:20:15.020Z",
      "content": "<p>Great job! I also tried VideoMAE-Large on a 24GB RTX 4090, and the performance was very good. (Score 0.837)</p>",
      "rawMarkdown": "Great job! I also tried VideoMAE-Large on a 24GB RTX 4090, and the performance was very good. (Score 0.837)",
      "votes": 1
    },
    {
      "id": 3204197,
      "postDate": "2025-05-17T22:01:35.287Z",
      "content": "<p>This approach is quite interesting. I'm curious how you selected the 1.5-second threshold for positive samples?</p>",
      "rawMarkdown": "This approach is quite interesting. I'm curious how you selected the 1.5-second threshold for positive samples?",
      "replies": [
        {
          "id": 3204199,
          "postDate": "2025-05-17T22:15:59.713Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 3205042,
          "postDate": "2025-05-19T07:45:09.530Z",
          "content": "<h1>Hey, thanks for the question—happy to clarify!</h1>\n<h2>Why we ignored <code>alert_time</code></h2>\n<p>Initially, we hypothesized that a large pre-trained transformer model (VideoMAEv2-giant) might catch some subtle predictive cues even earlier than human annotators' marked \"alert_time.\" Therefore, instead of strictly using the provided human-labeled earliest prediction timestamp (<code>alert_time</code>), we simply labelled a window positive if its last frame fell within 1.5 s before the event. Anything earlier became negative by default. This simplified labeling allowed us to easily align training with the primary evaluation threshold used in testing.</p>\n<h2>Why exactly 1.5 s?</h2>\n<p>The evaluation metric explicitly tests at TTAs of 0.5 s, 1.0 s, and 1.5 s. Importantly, videos at each TTA are included in the evaluation only if their human-labeled <code>alert_time</code> is earlier than <code>(event_time - TTA)</code>. Among these thresholds, the 1.5 s condition generally covers most evaluation scenarios, thus we directly adopted 1.5 s as our positive labeling threshold for simplicity and better alignment with the main evaluation criterion.</p>\n<h2>What happened when we tried respecting <code>alert_time</code>?</h2>\n<p>Later, we realized our original labeling strategy didn't fully match the official definition—windows ending earlier than the human-labeled <code>alert_time</code> should actually be labeled as negative, according to the official guidelines. To test this, we conducted a \"cleaner\" experiment, carefully re-labeling all windows that ended before the official <code>alert_time</code> as negative.</p>\n<p>➜ <strong>Result:</strong> Surprisingly, our validation acc &amp; lb score slightly decreased.</p>\n<p>We expected performance to remain stable or perhaps slightly improve, since we were now strictly following the official labeling standard. But the validation drop indicates that our initial, looser labeling unintentionally provided some helpful information. We're honestly not sure why—perhaps those early frames contain subtle signals that benefit model generalization, or maybe it was just random variation. In any case, it felt a bit mysterious (\"deep-learning feng-shui,\" as we jokingly called it 😅).</p>\n<h2>Why didn't we explore further?</h2>\n<p>server resources were costly, and our preprocessing pipeline was slow (probably due to our suboptimal optimization skills, haha). Given these constraints, we couldn't perform extensive ablation studies or systematically investigate this labeling issue further. As mentioned in our write-up, the whole solution was rather intuition-driven and far from rigorously tested.</p>\n<p>There's definitely room left for more systematic exploration—but this was our honest state of affairs at submission time. Hope this clarifies things better!</p>",
          "rawMarkdown": "# Hey, thanks for the question—happy to clarify!\n\n## Why we ignored `alert_time`\n\nInitially, we hypothesized that a large pre-trained transformer model (VideoMAEv2-giant) might catch some subtle predictive cues even earlier than human annotators' marked \"alert\\_time.\" Therefore, instead of strictly using the provided human-labeled earliest prediction timestamp (`alert_time`), we simply labelled a window positive if its last frame fell within 1.5 s before the event. Anything earlier became negative by default. This simplified labeling allowed us to easily align training with the primary evaluation threshold used in testing.\n\n## Why exactly 1.5 s?\n\nThe evaluation metric explicitly tests at TTAs of 0.5 s, 1.0 s, and 1.5 s. Importantly, videos at each TTA are included in the evaluation only if their human-labeled `alert_time` is earlier than `(event_time - TTA)`. Among these thresholds, the 1.5 s condition generally covers most evaluation scenarios, thus we directly adopted 1.5 s as our positive labeling threshold for simplicity and better alignment with the main evaluation criterion.\n\n## What happened when we tried respecting `alert_time`?\n\nLater, we realized our original labeling strategy didn't fully match the official definition—windows ending earlier than the human-labeled `alert_time` should actually be labeled as negative, according to the official guidelines. To test this, we conducted a \"cleaner\" experiment, carefully re-labeling all windows that ended before the official `alert_time` as negative.\n\n➜ **Result:** Surprisingly, our validation acc & lb score slightly decreased.\n\nWe expected performance to remain stable or perhaps slightly improve, since we were now strictly following the official labeling standard. But the validation drop indicates that our initial, looser labeling unintentionally provided some helpful information. We're honestly not sure why—perhaps those early frames contain subtle signals that benefit model generalization, or maybe it was just random variation. In any case, it felt a bit mysterious (\"deep-learning feng-shui,\" as we jokingly called it 😅).\n\n## Why didn't we explore further?\n\nserver resources were costly, and our preprocessing pipeline was slow (probably due to our suboptimal optimization skills, haha). Given these constraints, we couldn't perform extensive ablation studies or systematically investigate this labeling issue further. As mentioned in our write-up, the whole solution was rather intuition-driven and far from rigorously tested.\n\nThere's definitely room left for more systematic exploration—but this was our honest state of affairs at submission time. Hope this clarifies things better!"
        }
      ]
    },
    {
      "id": 3193308,
      "postDate": "2025-05-04T07:36:59.390Z",
      "content": "<p>Due to GitHub's file size restrictions, the pretrained weights have been migrated to Hugging Face Hub:<br>\n<a href=\"https://huggingface.co/zhiyaowang/VideoMaev2-giant-nexar-solution\" target=\"_blank\">https://huggingface.co/zhiyaowang/VideoMaev2-giant-nexar-solution</a></p>",
      "rawMarkdown": "Due to GitHub's file size restrictions, the pretrained weights have been migrated to Hugging Face Hub:\nhttps://huggingface.co/zhiyaowang/VideoMaev2-giant-nexar-solution"
    },
    {
      "id": 3193282,
      "postDate": "2025-05-04T07:04:33.740Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 3193295,
          "postDate": "2025-05-04T07:23:37.447Z",
          "rawMarkdown": "",
          "isDeleted": true,
          "replies": [
            {
              "id": 3193325,
              "postDate": "2025-05-04T08:04:57.027Z",
              "rawMarkdown": "",
              "votes": 1,
              "isDeleted": true
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3200026,
      "author_name": "Ray Ashby",
      "author_url": "",
      "post_date": "2025-05-11T22:10:25.473000",
      "content": "<p>Very much appreciate the details of your submission.    I like the thought process and the code details to learn from.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3194253,
      "author_name": "Peace.LU",
      "author_url": "",
      "post_date": "2025-05-05T15:20:15.020000",
      "content": "<p>Great job! I also tried VideoMAE-Large on a 24GB RTX 4090, and the performance was very good. (Score 0.837)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3204197,
      "author_name": "sftwre",
      "author_url": "",
      "post_date": "2025-05-17T22:01:35.287000",
      "content": "<p>This approach is quite interesting. I'm curious how you selected the 1.5-second threshold for positive samples?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3204199,
          "author_name": "",
          "author_url": "",
          "post_date": "2025-05-17T22:15:59.713000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3205042,
          "author_name": "Wang Zhiyao (王致尧)",
          "author_url": "",
          "post_date": "2025-05-19T07:45:09.530000",
          "content": "<h1>Hey, thanks for the question—happy to clarify!</h1>\n<h2>Why we ignored <code>alert_time</code></h2>\n<p>Initially, we hypothesized that a large pre-trained transformer model (VideoMAEv2-giant) might catch some subtle predictive cues even earlier than human annotators' marked \"alert_time.\" Therefore, instead of strictly using the provided human-labeled earliest prediction timestamp (<code>alert_time</code>), we simply labelled a window positive if its last frame fell within 1.5 s before the event. Anything earlier became negative by default. This simplified labeling allowed us to easily align training with the primary evaluation threshold used in testing.</p>\n<h2>Why exactly 1.5 s?</h2>\n<p>The evaluation metric explicitly tests at TTAs of 0.5 s, 1.0 s, and 1.5 s. Importantly, videos at each TTA are included in the evaluation only if their human-labeled <code>alert_time</code> is earlier than <code>(event_time - TTA)</code>. Among these thresholds, the 1.5 s condition generally covers most evaluation scenarios, thus we directly adopted 1.5 s as our positive labeling threshold for simplicity and better alignment with the main evaluation criterion.</p>\n<h2>What happened when we tried respecting <code>alert_time</code>?</h2>\n<p>Later, we realized our original labeling strategy didn't fully match the official definition—windows ending earlier than the human-labeled <code>alert_time</code> should actually be labeled as negative, according to the official guidelines. To test this, we conducted a \"cleaner\" experiment, carefully re-labeling all windows that ended before the official <code>alert_time</code> as negative.</p>\n<p>➜ <strong>Result:</strong> Surprisingly, our validation acc &amp; lb score slightly decreased.</p>\n<p>We expected performance to remain stable or perhaps slightly improve, since we were now strictly following the official labeling standard. But the validation drop indicates that our initial, looser labeling unintentionally provided some helpful information. We're honestly not sure why—perhaps those early frames contain subtle signals that benefit model generalization, or maybe it was just random variation. In any case, it felt a bit mysterious (\"deep-learning feng-shui,\" as we jokingly called it 😅).</p>\n<h2>Why didn't we explore further?</h2>\n<p>server resources were costly, and our preprocessing pipeline was slow (probably due to our suboptimal optimization skills, haha). Given these constraints, we couldn't perform extensive ablation studies or systematically investigate this labeling issue further. As mentioned in our write-up, the whole solution was rather intuition-driven and far from rigorously tested.</p>\n<p>There's definitely room left for more systematic exploration—but this was our honest state of affairs at submission time. Hope this clarifies things better!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3193308,
      "author_name": "Wang Zhiyao (王致尧)",
      "author_url": "",
      "post_date": "2025-05-04T07:36:59.390000",
      "content": "<p>Due to GitHub's file size restrictions, the pretrained weights have been migrated to Hugging Face Hub:<br>\n<a href=\"https://huggingface.co/zhiyaowang/VideoMaev2-giant-nexar-solution\" target=\"_blank\">https://huggingface.co/zhiyaowang/VideoMaev2-giant-nexar-solution</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3193282,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-05-04T07:04:33.740000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 3193295,
          "author_name": "",
          "author_url": "",
          "post_date": "2025-05-04T07:23:37.447000",
          "content": "",
          "votes": 0,
          "replies": [
            {
              "id": 3193325,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-05-04T08:04:57.027000",
              "content": "",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3193266": "# VideoMAE-based Vehicle Collision Prediction Solution: 2nd Place on Kaggle Public Leaderboard\n\n## Solution Overview\n\nHey everyone,\n\nI'd like to share my approach for the Nexar Safe Driving Video Analysis competition. This solution uses a deep learning method based on the **VideoMAEv2-giant** pre-trained model to predict collision and near-miss risks in driving videos.\n\nThe process was relatively straightforward, and it achieved **2nd place on the public leaderboard** \n\n## Why VideoMAEv2?\n\nWe framed the accident prediction task as a video classification problem: classifying videos into \"imminent collision/near-miss\" or \"normal driving\". After experimenting with several approaches, VideoMAEv2 delivered the best performance.\n\n*   **Initial Attempts**: We first tried more traditional architectures like CNN+LSTM and CNN+Attention, using the same data processing pipeline. However, these models generally scored between 0.7 and 0.8, falling short of VideoMAEv2.\n\n*   **VideoMAEv2 Advantages**: As a state-of-the-art pre-trained model for video understanding, VideoMAEv2 offers several key benefits:\n    1.  **Strong Spatio-temporal Feature Extraction**: It excels at capturing the dynamic information crucial for collision prediction.\n    2.  **Large-Scale Pre-training**: Pre-training on massive video datasets provides excellent generalization capabilities.\n    3.  **Action Recognition Performance**: Its proven success in action recognition tasks suggests suitability for understanding dynamic events like near-misses and collisions.\n\nIts strong performance on various video benchmarks further motivated its selection for this task.\n\n## Technical Implementation\n\n* **Label Assignment**\n\n  1. For positive videos (collisions/near-misses), only frames preceding the event time were processed.\n\n     ```python\n     if last_time >= event_time:\n         break  # Stop processing if past the event time\n     ```\n\n  2. **Threshold Selection & Handling of `alert_time`**\n\n     * Initially, we did not utilize the provided human-labeled `alert_time` (earliest prediction time). Instead, we labeled a window as positive if its last frame fell within **1.5 seconds** before the event; earlier frames were labeled negative by default.\n     * The **1.5-second** threshold was chosen because the evaluation metric specifically tests predictions at 0.5 s, 1.0 s, and 1.5 s TTAs. Among these, the 1.5-second condition generally covers most evaluation scenarios, allowing us to directly align our labeling strategy with the main evaluation criterion.\n     * Later, we conducted a \"cleaner\" experiment where we strictly adhered to the official labeling standard (marking windows that ended before the human-labeled `alert_time` as negative). Interestingly, this approach slightly **decreased validation accuracy and leaderboard score**, suggesting that our original, looser labeling method might have inadvertently captured useful early signals beneficial for model generalization.\n     * Due to limited computing resources and inefficient preprocessing pipelines, we couldn't perform extensive ablation studies to investigate this issue further.\n\n  3. For negative videos (normal driving), all windows were labeled negative.\n\n\n**Important Considerations**:\n\n*   **Ignoring `alert_time`**: The competition provided `alert_time` (human-annotated earliest prediction time). We didn't directly use it, hypothesizing that a deep learning model might capture longer-term dependencies better than human annotation. Instead, we opted for the fixed 1.5-second threshold before the event. Frankly, this lacks empirical rigor, as we didn't run experiments comparing different thresholds. This is a key area for future improvement.\n*   **Data Balancing Details**: This process yielded 4,258 positive samples, and we matched this count with randomly selected negative samples. Our sliding window approach generated far more negative than positive windows initially. This aggressive undersampling, while simple, likely discarded many potentially informative or hard-negative examples, possibly hindering the model's ability to learn fine-grained distinctions. Exploring techniques like Focal Loss could be beneficial.\n\n```python\n# Core label assignment logic\nif is_positive_video:\n    if last_time >= event_time:\n        break  # Stop processing if past the event time\n    label = 1 if (event_time - 1.5 <= last_time < event_time) else 0\nelse:\n    label = 0\n```\n\n### Model Architecture\n\nThe model architecture is straightforward:\n\n* **Backbone**: VideoMAEv2-giant pre-trained model.\n  It’s already *giant*—in name and in GPU appetite—so we figured it was doing more than enough of the heavy lifting.\n\n* **Classification Head**: A simple linear layer followed by Dropout (p=0.1).\n  We considered using a transformer-based or attention pooling classification head to further model temporal relationships between tokens. However, initial experiments did not show significant performance gains. Given the trade-off between computational cost and model simplicity—and the fact that the backbone was already doing the equivalent of a full-time job—we ultimately chose a simple linear head as the classifier.\n\n* **Training Technique**: Temperature scaling (T=2.0) applied to logits during training for smoother probability distributions.\n\n### Inference Method\n\nFor the test set, we used a simple yet effective inference strategy:\n\n*   **Extract only the last 2 seconds** of each test video.\n*   Feed this 2-second clip into the trained model for prediction.\n\nThis approach relies on the intuition that the most critical information for predicting an imminent event is likely concentrated towards the end of the video clip (especially since test videos were trimmed). This simplification performed well, likely because:\n\n1.  Positive samples (trimmed near events) have strong signals in the final moments.\n2.  Negative samples (normal driving) shouldn't have strong signals regardless of the segment.\n\nWhile testing various window lengths would be ideal, this focused approach reduced computational load and aligned with the intuition that later predictions tend to be more accurate for imminent events. It yielded good results despite its simplicity.\n\n### Training Environment and Strategy\n\n*   **Hardware**: Trained on an H20 server (96GB VRAM). Training is also feasible on an A100 (40GB). Smaller GPUs might require using smaller VideoMAE variants.\n*   **Reproducibility**: Fixed random seed (`seed=42`) for Python, NumPy, and PyTorch (including CUDA) ensures deterministic runs.\n*   **Key Hyperparameters**:\n    *   Optimizer: AdamW\n    *   Learning Rate: 1e-5 (with Cosine Annealing scheduler)\n    *   Batch Size: 3 (effective batch size 24 due to gradient accumulation over 8 steps)\n    *   Epochs: 10 (saving the best model based on validation accuracy)\n\n## Limitations and Room for Improvement\n\nGiven academic and resource constraints, this solution has several limitations and areas for future work:\n\n1.  **Fixed Sampling Strategy**: The 2-second window (16 frames) was chosen somewhat arbitrarily. Systematic experiments comparing different window lengths and frame rates are needed.\n2.  **Validation Mismatch**: Our local validation setup didn't perfectly mirror the competition's evaluation metric (Mean Average Precision over time-to-accident thresholds). This might mean the submitted model wasn't the true optimal one based on the leaderboard metric.\n3.  **Single Model Submission**: Although the prediction script includes ensemble capabilities, the final submission relied on a single model, missing potential gains from ensembling.\n4.  **Limited Data Augmentation**: More sophisticated video-specific augmentations (e.g., temporal jittering, different cropping strategies) could be explored.\n5.  **Crude Class Imbalance Handling**: As noted, simple undersampling might not be optimal. Exploring methods like Focal Loss or more intelligent negative mining could improve performance.\n\n## Code and Model Weights\n\nAll code has been uploaded to the GitHub repository:\n[https://github.com/wzyfromhust/nexer-solution](https://github.com/wzyfromhust/nexer-solution)\n\n**Note**: The final model weights are quite large and will be uploaded later when feasible. However, the code includes the fixed random seed (`seed=42`), so reproducing the results by following the data processing and training scripts should be relatively straightforward.\n\n---\n\nAchieving a decent rank with what feels like a somewhat unrefined and insufficiently experimented approach suggests either a good dose of luck, or perhaps, significant untapped potential in the VideoMAEv2 model for this specific task.\n",
    "3200026": "Very much appreciate the details of your submission.    I like the thought process and the code details to learn from.",
    "3194253": "Great job! I also tried VideoMAE-Large on a 24GB RTX 4090, and the performance was very good. (Score 0.837)",
    "3204197": "This approach is quite interesting. I'm curious how you selected the 1.5-second threshold for positive samples?",
    "3193308": "Due to GitHub's file size restrictions, the pretrained weights have been migrated to Hugging Face Hub:\nhttps://huggingface.co/zhiyaowang/VideoMaev2-giant-nexar-solution",
    "3193282": ""
  }
}