{
  "id": 578712,
  "title": "3rd place solution: Two Frames Is All You Need",
  "url": "/competitions/nexar-collision-prediction/discussion/578712",
  "author_name": "",
  "post_date": "2025-05-12T18:59:23.572289600Z",
  "votes": 13,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hi everyone! This post contains my 3rd place solution to this Kaggle competition. Thanks a lot for the organizers for setting up this competition, and congrats to all the winners!</p>\n<p><strong>Summary</strong>:</p>\n<ul>\n<li>Transformed the problem to image classification.</li>\n<li>Extracted vehicle mask using YOLO-v8 and optical flow.</li>\n<li>Used two CNN (ResNet18 pre-trained on ImageNet) as backbones with single classificication head.<ul>\n<li>One CNN used the masked image, the other one used the mask + masked optical flow.</li></ul></li>\n<li>Averaged the predictions of an ensemble of 2 models.</li>\n</ul>\n<h3>Two Frames Is All You Need</h3>\n<p>To prevent overfitting on only 1500 video's with large video-based deep learning models, I took a very simple CNN-based approach that was inspired by the Dutch drivers' licence theoretical exam (all my fellow Dutchies should recognize this!). In this exam the student is presented with an image of a road scene where potentially an accident is about to happen. The student has three options to choose from: 'break', 'let go of the gas' or 'do nothing'. I'm convinced that if driving students are able to predict if an accident will happen from a single image, than a computer vision model should be too!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7949671%2Fee5a073dce99e9ae5c3c12008e542f18%2Fgevaarherkenning.jpg?generation=1747074061294039&amp;alt=media\" alt=\"\"></p>\n<p>We further simplify the classification problem by only presenting the model with the position, direction and velocity of any vehicles in the frame. This ensures that the model will not learn wrong associations (e.g. predicting that an accident will happen based on the color of a car).</p>\n<p>For example, for video 00325 at 00:18 we have the following frame:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7949671%2F7234c1cf9c9fbafbca54ea6ef37d3109%2Fframe.png?generation=1747074492112612&amp;alt=media\" alt=\"\"></p>\n<p>From which we can extract the following vehicle mask:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7949671%2Fa4445e01e1a4bfd5707023aec82c43c4%2Fplot_mask.png?generation=1747074511114670&amp;alt=media\" alt=\"\"></p>\n<p>And for every vehicle add the direction and velocity using Farneback Optical Flow, resulting in a 3-channel image where the channels represent the <strong>position</strong>, <strong>direction</strong> and <strong>velocity</strong> of each vehicle (if any) respectively:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7949671%2F83504009c309745f1ee31f02eb279d1b%2Foptical_flow.png?generation=1747074568521329&amp;alt=media\" alt=\"\"><br>\n<strong>Note</strong>: These green arrows are not literally in the image that is passed to the CNN. I plotted the direction and velocity of the optical flow as arrows to make it easier to understand.</p>\n<p>This last image is all we should need to predict if an accident is about to happen. We clearly see that there is a vehicle right in front of us rapidly moving in our direction. And we only needed <strong>two frames</strong> to get this information! </p>\n<h2>Data Processing</h2>\n<p>For extracting the vehicle mask I used the YOLO-v8 segmentation model by <a href=\"https://docs.ultralytics.com/tasks/segment/\" target=\"_blank\">Ultralytics</a>. For calculating the optical flow I used the <code>calcOpticalFlowFarneback</code> method from OpenCV, where the difference between each target frame <em>t</em> and frame <em>t-3</em> was used (delta 0.1 sec at 30FPS).</p>\n<p>Different data processing strategies were used depending on the dataset and label:</p>\n<ul>\n<li><strong>Negative training samples</strong>: 40 frames (every 1 second at 30FPS) were sampled and the frame (RGB), mask and flows were saved.</li>\n<li><strong>Positive training samples</strong>: For all frames between <code>time_of_alert</code> and <code>time_of_event</code> the frame (RGB), mask and flows were saved.</li>\n<li><strong>Test</strong>: For the last three frames the frame (RGB), mask and flows were saved.</li>\n</ul>\n<p>The data is then stored as pytorch tensors in the following folder structure:</p>\n<pre><code>/\n├── flows/\n│   ├── \n│   ├── \n│   └── \n├── frames/\n│   ├── \n│   ├── \n│   └── \n└── masks/\n    ├── \n    ├── \n    └── \n\n/\n...\n</code></pre>\n<h2>Model architecture</h2>\n<p>The model consists of two CNN backbones that are fed to a single classifier:</p>\n<ul>\n<li><strong>CNN for mask and flow</strong>: ResNet18 pre-trained on imagenet. Last fully connected layer is replaced by Identity operation.</li>\n<li><strong>CNN for frame</strong>: ResNet18 pre-trained on imagenet (CNN weights frozen). Last fully connected layer is replaced by a trainable linear layer of 512 to 512.</li>\n<li><strong>Classifier</strong>: Linear layer of (512 + 512) to 1.</li>\n</ul>\n<p>Note that the <strong>CNN for frame</strong> model is only fed the frame where <code>mask &gt; 0</code>:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7949671%2F3329ea5bc464f75b2bdef4878c0bb089%2Fmasked_RGB.png?generation=1747120748458166&amp;alt=media\" alt=\"\"></p>\n<h2>Training</h2>\n<p>The following hyperparameters were used:</p>\n<ul>\n<li>Learning rate: 0.001</li>\n<li>Optimizer: AdamW</li>\n<li>Batch size: 32</li>\n</ul>\n<p>Models were trained with a 10% validation set for a maximum of 20 epochs, with early stopping after not improving the test accuracy for 8 epochs. The best model is then used for inference.</p>\n<p>For every <code>n</code> samples in a single batch:</p>\n<ol>\n<li>A total of <code>n</code> random video's are selected.</li>\n<li>For each video we sample a frame together with the corresponding mask and flow.</li>\n</ol>\n<p>We also applied the following transformations to the training data:</p>\n<pre><code>transform=T.Compose([\n    T.Lambda(pad_to_square),\n    T.Normalize(mean=[, , ], std=[, , ]),\n    T.RandomHorizontalFlip(),\n    T.RandomAffine(degrees=, translate=(, ), scale=(, ), shear=),\n])\n</code></pre>\n<p>where <code>pad_to_square</code> is defined as:</p>\n<pre><code> ():\n    w, h = image.size\n    max_dim = (w, h)\n    pad_w = (max_dim - w) // \n    pad_h = (max_dim - h) // \n\n     T.Pad((pad_w, pad_h, max_dim - w - pad_w, max_dim - h - pad_h), fill=)(image)\n</code></pre>\n<p><a href=\"https://docs.pytorch.org/vision/main/models/generated/torchvision.models.resnet18.html\" target=\"_blank\">Resnet18</a> reshapes the input images to a square before the forward pass: <code>The images are resized to resize_size=[256] using interpolation=InterpolationMode.BILINEAR, followed by a central crop of crop_size=[224]</code>. By padding the image to a square first we prevent the shape of each vehicle to be warped.</p>\n<p>All training runs are logged to a WandB repository which can be found <a href=\"https://wandb.ai/maxzw/nexar-collision-prediction\" target=\"_blank\">here</a>.</p>\n<h2>Inference</h2>\n<p>We applied the following transformations to the test data:</p>\n<pre><code>test_transform=T.Compose([\n    T.Lambda(pad_to_square),\n    T.Normalize(mean=[, , ], std=[, , ]),\n])\n</code></pre>\n<p>The model prediction is a weighted average of the predictions of the last three frames:</p>\n<ul>\n<li><strong>Last frame</strong>: 50%</li>\n<li><strong>Second to last frame</strong>: 30%</li>\n<li><strong>Third to last frame</strong>: 20%</li>\n</ul>\n<p>The performance of individual models reached between 0.75 and 0.83 (see validation accuracy below - note that this is without a weighted average), but there was a lot of variation in the predictions. Therefore the final predictions are the averaged predictions of an ensemble of 2 models, which resulted in the final score.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7949671%2Fff86394430b189e77a2b81966fdcbf13%2Fvalidation%20accuracy.png?generation=1747120372267833&amp;alt=media\" alt=\"\"></p>\n<h2>Code</h2>\n<ul>\n<li>A github repository containing all code required to train and evaluate a model can be found <a href=\"https://github.com/maxzw/nexar-collision-prediction\" target=\"_blank\">here</a>.</li>\n<li>A Kaggle notebook to run inference on the test set can be found <a href=\"https://www.kaggle.com/code/mzwager/inference-notebook-lb-0-861\" target=\"_blank\">here</a>.</li>\n</ul>\n<h2>Hardware details</h2>\n<p>All code is run on a M2 Macbook Pro (16GB). You can find the specs <a href=\"https://support.apple.com/en-us/111838\" target=\"_blank\">here</a>. Using Pytorch Lightning helps with utilizing Apple's mps compute. </p>\n<h2>Limitations</h2>\n<ul>\n<li>Because the input images to both CNNs are multiplied with the vehicle mask, these images will be completely black (all values will be zero) if YOLO-v8 cannot find any vehicles. This means a good segmentation model is essential for this approach to work.</li>\n<li>Because both CNNs work independently it's not possible for a model to learn relationships between the positition, direction and velocity of the vehicles and the RGB values, for example breaking lights. I experimented with using a non pre-trained CNN that accepted 6-channel images with shape (B, 6, H, W), but this did not perform well. This highlights the benefit of using pre-trained computer vision models.</li>\n</ul>",
  "messages": [
    {
      "id": "3200600",
      "postDate": "05/12/2025 18:59:23",
      "content": "<p>Hi everyone! This post contains my 3rd place solution to this Kaggle competition. Thanks a lot for the organizers for setting up this competition, and congrats to all the winners!</p>\n<p><strong>Summary</strong>:</p>\n<ul>\n<li>Transformed the problem to image classification.</li>\n<li>Extracted vehicle mask using YOLO-v8 and optical flow.</li>\n<li>Used two CNN (ResNet18 pre-trained on ImageNet) as backbones with single classificication head.<ul>\n<li>One CNN used the masked image, the other one used the mask + masked optical flow.</li></ul></li>\n<li>Averaged the predictions of an ensemble of 2 models.</li>\n</ul>\n<h3>Two Frames Is All You Need</h3>\n<p>To prevent overfitting on only 1500 video's with large video-based deep learning models, I took a very simple CNN-based approach that was inspired by the Dutch drivers' licence theoretical exam (all my fellow Dutchies should recognize this!). In this exam the student is presented with an image of a road scene where potentially an accident is about to happen. The student has three options to choose from: 'break', 'let go of the gas' or 'do nothing'. I'm convinced that if driving students are able to predict if an accident will happen from a single image, than a computer vision model should be too!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7949671%2Fee5a073dce99e9ae5c3c12008e542f18%2Fgevaarherkenning.jpg?generation=1747074061294039&amp;alt=media\" alt=\"\"></p>\n<p>We further simplify the classification problem by only presenting the model with the position, direction and velocity of any vehicles in the frame. This ensures that the model will not learn wrong associations (e.g. predicting that an accident will happen based on the color of a car).</p>\n<p>For example, for video 00325 at 00:18 we have the following frame:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7949671%2F7234c1cf9c9fbafbca54ea6ef37d3109%2Fframe.png?generation=1747074492112612&amp;alt=media\" alt=\"\"></p>\n<p>From which we can extract the following vehicle mask:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7949671%2Fa4445e01e1a4bfd5707023aec82c43c4%2Fplot_mask.png?generation=1747074511114670&amp;alt=media\" alt=\"\"></p>\n<p>And for every vehicle add the direction and velocity using Farneback Optical Flow, resulting in a 3-channel image where the channels represent the <strong>position</strong>, <strong>direction</strong> and <strong>velocity</strong> of each vehicle (if any) respectively:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7949671%2F83504009c309745f1ee31f02eb279d1b%2Foptical_flow.png?generation=1747074568521329&amp;alt=media\" alt=\"\"><br>\n<strong>Note</strong>: These green arrows are not literally in the image that is passed to the CNN. I plotted the direction and velocity of the optical flow as arrows to make it easier to understand.</p>\n<p>This last image is all we should need to predict if an accident is about to happen. We clearly see that there is a vehicle right in front of us rapidly moving in our direction. And we only needed <strong>two frames</strong> to get this information! </p>\n<h2>Data Processing</h2>\n<p>For extracting the vehicle mask I used the YOLO-v8 segmentation model by <a href=\"https://docs.ultralytics.com/tasks/segment/\" target=\"_blank\">Ultralytics</a>. For calculating the optical flow I used the <code>calcOpticalFlowFarneback</code> method from OpenCV, where the difference between each target frame <em>t</em> and frame <em>t-3</em> was used (delta 0.1 sec at 30FPS).</p>\n<p>Different data processing strategies were used depending on the dataset and label:</p>\n<ul>\n<li><strong>Negative training samples</strong>: 40 frames (every 1 second at 30FPS) were sampled and the frame (RGB), mask and flows were saved.</li>\n<li><strong>Positive training samples</strong>: For all frames between <code>time_of_alert</code> and <code>time_of_event</code> the frame (RGB), mask and flows were saved.</li>\n<li><strong>Test</strong>: For the last three frames the frame (RGB), mask and flows were saved.</li>\n</ul>\n<p>The data is then stored as pytorch tensors in the following folder structure:</p>\n<pre><code>/\n├── flows/\n│   ├── \n│   ├── \n│   └── \n├── frames/\n│   ├── \n│   ├── \n│   └── \n└── masks/\n    ├── \n    ├── \n    └── \n\n/\n...\n</code></pre>\n<h2>Model architecture</h2>\n<p>The model consists of two CNN backbones that are fed to a single classifier:</p>\n<ul>\n<li><strong>CNN for mask and flow</strong>: ResNet18 pre-trained on imagenet. Last fully connected layer is replaced by Identity operation.</li>\n<li><strong>CNN for frame</strong>: ResNet18 pre-trained on imagenet (CNN weights frozen). Last fully connected layer is replaced by a trainable linear layer of 512 to 512.</li>\n<li><strong>Classifier</strong>: Linear layer of (512 + 512) to 1.</li>\n</ul>\n<p>Note that the <strong>CNN for frame</strong> model is only fed the frame where <code>mask &gt; 0</code>:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7949671%2F3329ea5bc464f75b2bdef4878c0bb089%2Fmasked_RGB.png?generation=1747120748458166&amp;alt=media\" alt=\"\"></p>\n<h2>Training</h2>\n<p>The following hyperparameters were used:</p>\n<ul>\n<li>Learning rate: 0.001</li>\n<li>Optimizer: AdamW</li>\n<li>Batch size: 32</li>\n</ul>\n<p>Models were trained with a 10% validation set for a maximum of 20 epochs, with early stopping after not improving the test accuracy for 8 epochs. The best model is then used for inference.</p>\n<p>For every <code>n</code> samples in a single batch:</p>\n<ol>\n<li>A total of <code>n</code> random video's are selected.</li>\n<li>For each video we sample a frame together with the corresponding mask and flow.</li>\n</ol>\n<p>We also applied the following transformations to the training data:</p>\n<pre><code>transform=T.Compose([\n    T.Lambda(pad_to_square),\n    T.Normalize(mean=[, , ], std=[, , ]),\n    T.RandomHorizontalFlip(),\n    T.RandomAffine(degrees=, translate=(, ), scale=(, ), shear=),\n])\n</code></pre>\n<p>where <code>pad_to_square</code> is defined as:</p>\n<pre><code> ():\n    w, h = image.size\n    max_dim = (w, h)\n    pad_w = (max_dim - w) // \n    pad_h = (max_dim - h) // \n\n     T.Pad((pad_w, pad_h, max_dim - w - pad_w, max_dim - h - pad_h), fill=)(image)\n</code></pre>\n<p><a href=\"https://docs.pytorch.org/vision/main/models/generated/torchvision.models.resnet18.html\" target=\"_blank\">Resnet18</a> reshapes the input images to a square before the forward pass: <code>The images are resized to resize_size=[256] using interpolation=InterpolationMode.BILINEAR, followed by a central crop of crop_size=[224]</code>. By padding the image to a square first we prevent the shape of each vehicle to be warped.</p>\n<p>All training runs are logged to a WandB repository which can be found <a href=\"https://wandb.ai/maxzw/nexar-collision-prediction\" target=\"_blank\">here</a>.</p>\n<h2>Inference</h2>\n<p>We applied the following transformations to the test data:</p>\n<pre><code>test_transform=T.Compose([\n    T.Lambda(pad_to_square),\n    T.Normalize(mean=[, , ], std=[, , ]),\n])\n</code></pre>\n<p>The model prediction is a weighted average of the predictions of the last three frames:</p>\n<ul>\n<li><strong>Last frame</strong>: 50%</li>\n<li><strong>Second to last frame</strong>: 30%</li>\n<li><strong>Third to last frame</strong>: 20%</li>\n</ul>\n<p>The performance of individual models reached between 0.75 and 0.83 (see validation accuracy below - note that this is without a weighted average), but there was a lot of variation in the predictions. Therefore the final predictions are the averaged predictions of an ensemble of 2 models, which resulted in the final score.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7949671%2Fff86394430b189e77a2b81966fdcbf13%2Fvalidation%20accuracy.png?generation=1747120372267833&amp;alt=media\" alt=\"\"></p>\n<h2>Code</h2>\n<ul>\n<li>A github repository containing all code required to train and evaluate a model can be found <a href=\"https://github.com/maxzw/nexar-collision-prediction\" target=\"_blank\">here</a>.</li>\n<li>A Kaggle notebook to run inference on the test set can be found <a href=\"https://www.kaggle.com/code/mzwager/inference-notebook-lb-0-861\" target=\"_blank\">here</a>.</li>\n</ul>\n<h2>Hardware details</h2>\n<p>All code is run on a M2 Macbook Pro (16GB). You can find the specs <a href=\"https://support.apple.com/en-us/111838\" target=\"_blank\">here</a>. Using Pytorch Lightning helps with utilizing Apple's mps compute. </p>\n<h2>Limitations</h2>\n<ul>\n<li>Because the input images to both CNNs are multiplied with the vehicle mask, these images will be completely black (all values will be zero) if YOLO-v8 cannot find any vehicles. This means a good segmentation model is essential for this approach to work.</li>\n<li>Because both CNNs work independently it's not possible for a model to learn relationships between the positition, direction and velocity of the vehicles and the RGB values, for example breaking lights. I experimented with using a non pre-trained CNN that accepted 6-channel images with shape (B, 6, H, W), but this did not perform well. This highlights the benefit of using pre-trained computer vision models.</li>\n</ul>",
      "rawMarkdown": "Hi everyone! This post contains my 3rd place solution to this Kaggle competition. Thanks a lot for the organizers for setting up this competition, and congrats to all the winners!\n\n**Summary**:\n- Transformed the problem to image classification.\n- Extracted vehicle mask using YOLO-v8 and optical flow.\n- Used two CNN (ResNet18 pre-trained on ImageNet) as backbones with single classificication head.\n    - One CNN used the masked image, the other one used the mask + masked optical flow.\n- Averaged the predictions of an ensemble of 2 models.\n\n### Two Frames Is All You Need\nTo prevent overfitting on only 1500 video's with large video-based deep learning models, I took a very simple CNN-based approach that was inspired by the Dutch drivers' licence theoretical exam (all my fellow Dutchies should recognize this!). In this exam the student is presented with an image of a road scene where potentially an accident is about to happen. The student has three options to choose from: 'break', 'let go of the gas' or 'do nothing'. I'm convinced that if driving students are able to predict if an accident will happen from a single image, than a computer vision model should be too!\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7949671%2Fee5a073dce99e9ae5c3c12008e542f18%2Fgevaarherkenning.jpg?generation=1747074061294039&alt=media)\n\nWe further simplify the classification problem by only presenting the model with the position, direction and velocity of any vehicles in the frame. This ensures that the model will not learn wrong associations (e.g. predicting that an accident will happen based on the color of a car).\n\nFor example, for video 00325 at 00:18 we have the following frame:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7949671%2F7234c1cf9c9fbafbca54ea6ef37d3109%2Fframe.png?generation=1747074492112612&alt=media)\n\nFrom which we can extract the following vehicle mask:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7949671%2Fa4445e01e1a4bfd5707023aec82c43c4%2Fplot_mask.png?generation=1747074511114670&alt=media)\n\nAnd for every vehicle add the direction and velocity using Farneback Optical Flow, resulting in a 3-channel image where the channels represent the **position**, **direction** and **velocity** of each vehicle (if any) respectively:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7949671%2F83504009c309745f1ee31f02eb279d1b%2Foptical_flow.png?generation=1747074568521329&alt=media)\n**Note**: These green arrows are not literally in the image that is passed to the CNN. I plotted the direction and velocity of the optical flow as arrows to make it easier to understand.\n\nThis last image is all we should need to predict if an accident is about to happen. We clearly see that there is a vehicle right in front of us rapidly moving in our direction. And we only needed **two frames** to get this information! \n\n## Data Processing\nFor extracting the vehicle mask I used the YOLO-v8 segmentation model by [Ultralytics](https://docs.ultralytics.com/tasks/segment/). For calculating the optical flow I used the `calcOpticalFlowFarneback` method from OpenCV, where the difference between each target frame *t* and frame *t-3* was used (delta 0.1 sec at 30FPS).\n\nDifferent data processing strategies were used depending on the dataset and label:\n- **Negative training samples**: 40 frames (every 1 second at 30FPS) were sampled and the frame (RGB), mask and flows were saved.\n- **Positive training samples**: For all frames between `time_of_alert` and `time_of_event` the frame (RGB), mask and flows were saved.\n- **Test**: For the last three frames the frame (RGB), mask and flows were saved.\n\nThe data is then stored as pytorch tensors in the following folder structure:\n```\n00001/\n├── flows/\n│   ├── 00.pt\n│   ├── 01.pt\n│   └── 02.pt\n├── frames/\n│   ├── 00.pt\n│   ├── 01.pt\n│   └── 02.pt\n└── masks/\n    ├── 00.pt\n    ├── 01.pt\n    └── 02.pt\n\n00002/\n...\n```\n\n## Model architecture\nThe model consists of two CNN backbones that are fed to a single classifier:\n- **CNN for mask and flow**: ResNet18 pre-trained on imagenet. Last fully connected layer is replaced by Identity operation.\n- **CNN for frame**: ResNet18 pre-trained on imagenet (CNN weights frozen). Last fully connected layer is replaced by a trainable linear layer of 512 to 512.\n- **Classifier**: Linear layer of (512 + 512) to 1.\n\nNote that the **CNN for frame** model is only fed the frame where `mask > 0`:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7949671%2F3329ea5bc464f75b2bdef4878c0bb089%2Fmasked_RGB.png?generation=1747120748458166&alt=media)\n\n## Training\nThe following hyperparameters were used:\n- Learning rate: 0.001\n- Optimizer: AdamW\n- Batch size: 32\n\nModels were trained with a 10% validation set for a maximum of 20 epochs, with early stopping after not improving the test accuracy for 8 epochs. The best model is then used for inference.\n\nFor every `n` samples in a single batch:\n1. A total of `n` random video's are selected.\n2. For each video we sample a frame together with the corresponding mask and flow.\n\nWe also applied the following transformations to the training data:\n```python\ntransform=T.Compose([\n    T.Lambda(pad_to_square),\n    T.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]),\n    T.RandomHorizontalFlip(),\n    T.RandomAffine(degrees=5, translate=(0.1, 0.1), scale=(0.9, 1.1), shear=5),\n])\n```\n\nwhere `pad_to_square` is defined as:\n```python\ndef pad_to_square(image):\n    w, h = image.size\n    max_dim = max(w, h)\n    pad_w = (max_dim - w) // 2\n    pad_h = (max_dim - h) // 2\n\n    return T.Pad((pad_w, pad_h, max_dim - w - pad_w, max_dim - h - pad_h), fill=0)(image)\n```\n[Resnet18](https://docs.pytorch.org/vision/main/models/generated/torchvision.models.resnet18.html) reshapes the input images to a square before the forward pass: `The images are resized to resize_size=[256] using interpolation=InterpolationMode.BILINEAR, followed by a central crop of crop_size=[224]`. By padding the image to a square first we prevent the shape of each vehicle to be warped.\n\nAll training runs are logged to a WandB repository which can be found [here](https://wandb.ai/maxzw/nexar-collision-prediction).\n\n## Inference\nWe applied the following transformations to the test data:\n```python\ntest_transform=T.Compose([\n    T.Lambda(pad_to_square),\n    T.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]),\n])\n```\n\nThe model prediction is a weighted average of the predictions of the last three frames:\n- **Last frame**: 50%\n- **Second to last frame**: 30%\n- **Third to last frame**: 20%\n\nThe performance of individual models reached between 0.75 and 0.83 (see validation accuracy below - note that this is without a weighted average), but there was a lot of variation in the predictions. Therefore the final predictions are the averaged predictions of an ensemble of 2 models, which resulted in the final score.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7949671%2Fff86394430b189e77a2b81966fdcbf13%2Fvalidation%20accuracy.png?generation=1747120372267833&alt=media)\n\n## Code\n- A github repository containing all code required to train and evaluate a model can be found [here](https://github.com/maxzw/nexar-collision-prediction).\n- A Kaggle notebook to run inference on the test set can be found [here](https://www.kaggle.com/code/mzwager/inference-notebook-lb-0-861).\n\n## Hardware details\nAll code is run on a M2 Macbook Pro (16GB). You can find the specs [here](https://support.apple.com/en-us/111838). Using Pytorch Lightning helps with utilizing Apple's mps compute. \n\n## Limitations\n- Because the input images to both CNNs are multiplied with the vehicle mask, these images will be completely black (all values will be zero) if YOLO-v8 cannot find any vehicles. This means a good segmentation model is essential for this approach to work.\n- Because both CNNs work independently it's not possible for a model to learn relationships between the positition, direction and velocity of the vehicles and the RGB values, for example breaking lights. I experimented with using a non pre-trained CNN that accepted 6-channel images with shape (B, 6, H, W), but this did not perform well. This highlights the benefit of using pre-trained computer vision models.",
      "votes": null
    },
    {
      "id": "3200636",
      "postDate": "05/12/2025 19:55:03",
      "content": "<p>Very nice!</p>",
      "rawMarkdown": "Very nice!",
      "votes": null
    },
    {
      "id": "3200863",
      "postDate": "05/13/2025 06:58:47",
      "content": "<p>Clever approach!</p>",
      "rawMarkdown": "Clever approach!",
      "votes": null
    },
    {
      "id": "3200891",
      "postDate": "05/13/2025 07:41:06",
      "content": "<p>Cool, well done!</p>",
      "rawMarkdown": "Cool, well done!",
      "votes": null
    },
    {
      "id": "3200911",
      "postDate": "05/13/2025 08:09:42",
      "content": "<p>Great work!</p>",
      "rawMarkdown": "Great work!",
      "votes": null
    },
    {
      "id": "3202501",
      "postDate": "05/15/2025 13:37:09",
      "content": "<p>Congratulations, Great job!</p>",
      "rawMarkdown": "Congratulations, Great job!",
      "votes": null
    },
    {
      "id": "3202715",
      "postDate": "05/15/2025 19:39:11",
      "content": "<p>Proof that smart pre-processing is better than large models !</p>",
      "rawMarkdown": "Proof that smart pre-processing is better than large models !",
      "votes": null
    },
    {
      "id": "3202785",
      "postDate": "05/15/2025 22:48:19",
      "content": "<p>Using <code>x</code>, <code>y,</code> and <code>heading</code> provides structured spatial information that complements image and lidar data. While images and lidar offer rich visual and depth cues, they require complex processing to extract positional information. In contrast, <code>x</code>, <code>y</code>, and <code>heading</code> directly convey an object's location and orientation, which is invaluable for tasks like trajectory prediction and collision estimation.</p>\n<p>Combining these structured features with unstructured sensor data can enhance model performance by providing both precise localization and contextual understanding.</p>",
      "rawMarkdown": "Using `x`, `y,` and `heading` provides structured spatial information that complements image and lidar data. While images and lidar offer rich visual and depth cues, they require complex processing to extract positional information. In contrast, `x`, `y`, and `heading` directly convey an object's location and orientation, which is invaluable for tasks like trajectory prediction and collision estimation.\n\nCombining these structured features with unstructured sensor data can enhance model performance by providing both precise localization and contextual understanding.",
      "votes": null
    },
    {
      "id": "3207979",
      "postDate": "05/23/2025 14:01:07",
      "content": "<p>Love you approach</p>",
      "rawMarkdown": "Love you approach",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3200636,
      "author_name": "paulendresen76",
      "author_url": "",
      "post_date": "05/12/2025 19:55:03",
      "content": "<p>Very nice!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3200863,
      "author_name": "rvdgeer",
      "author_url": "",
      "post_date": "05/13/2025 06:58:47",
      "content": "<p>Clever approach!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3200891,
      "author_name": "diederikp",
      "author_url": "",
      "post_date": "05/13/2025 07:41:06",
      "content": "<p>Cool, well done!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3200911,
      "author_name": "bejiafif",
      "author_url": "",
      "post_date": "05/13/2025 08:09:42",
      "content": "<p>Great work!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3202501,
      "author_name": "julesvanligtenberg",
      "author_url": "",
      "post_date": "05/15/2025 13:37:09",
      "content": "<p>Congratulations, Great job!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3202715,
      "author_name": "bratjay",
      "author_url": "",
      "post_date": "05/15/2025 19:39:11",
      "content": "<p>Proof that smart pre-processing is better than large models !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3202785,
      "author_name": "alifsathar",
      "author_url": "",
      "post_date": "05/15/2025 22:48:19",
      "content": "<p>Using <code>x</code>, <code>y,</code> and <code>heading</code> provides structured spatial information that complements image and lidar data. While images and lidar offer rich visual and depth cues, they require complex processing to extract positional information. In contrast, <code>x</code>, <code>y</code>, and <code>heading</code> directly convey an object's location and orientation, which is invaluable for tasks like trajectory prediction and collision estimation.</p>\n<p>Combining these structured features with unstructured sensor data can enhance model performance by providing both precise localization and contextual understanding.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3207979,
      "author_name": "lavrikovav",
      "author_url": "",
      "post_date": "05/23/2025 14:01:07",
      "content": "<p>Love you approach</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3200600": "Hi everyone! This post contains my 3rd place solution to this Kaggle competition. Thanks a lot for the organizers for setting up this competition, and congrats to all the winners!\n\n**Summary**:\n- Transformed the problem to image classification.\n- Extracted vehicle mask using YOLO-v8 and optical flow.\n- Used two CNN (ResNet18 pre-trained on ImageNet) as backbones with single classificication head.\n    - One CNN used the masked image, the other one used the mask + masked optical flow.\n- Averaged the predictions of an ensemble of 2 models.\n\n### Two Frames Is All You Need\nTo prevent overfitting on only 1500 video's with large video-based deep learning models, I took a very simple CNN-based approach that was inspired by the Dutch drivers' licence theoretical exam (all my fellow Dutchies should recognize this!). In this exam the student is presented with an image of a road scene where potentially an accident is about to happen. The student has three options to choose from: 'break', 'let go of the gas' or 'do nothing'. I'm convinced that if driving students are able to predict if an accident will happen from a single image, than a computer vision model should be too!\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7949671%2Fee5a073dce99e9ae5c3c12008e542f18%2Fgevaarherkenning.jpg?generation=1747074061294039&alt=media)\n\nWe further simplify the classification problem by only presenting the model with the position, direction and velocity of any vehicles in the frame. This ensures that the model will not learn wrong associations (e.g. predicting that an accident will happen based on the color of a car).\n\nFor example, for video 00325 at 00:18 we have the following frame:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7949671%2F7234c1cf9c9fbafbca54ea6ef37d3109%2Fframe.png?generation=1747074492112612&alt=media)\n\nFrom which we can extract the following vehicle mask:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7949671%2Fa4445e01e1a4bfd5707023aec82c43c4%2Fplot_mask.png?generation=1747074511114670&alt=media)\n\nAnd for every vehicle add the direction and velocity using Farneback Optical Flow, resulting in a 3-channel image where the channels represent the **position**, **direction** and **velocity** of each vehicle (if any) respectively:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7949671%2F83504009c309745f1ee31f02eb279d1b%2Foptical_flow.png?generation=1747074568521329&alt=media)\n**Note**: These green arrows are not literally in the image that is passed to the CNN. I plotted the direction and velocity of the optical flow as arrows to make it easier to understand.\n\nThis last image is all we should need to predict if an accident is about to happen. We clearly see that there is a vehicle right in front of us rapidly moving in our direction. And we only needed **two frames** to get this information! \n\n## Data Processing\nFor extracting the vehicle mask I used the YOLO-v8 segmentation model by [Ultralytics](https://docs.ultralytics.com/tasks/segment/). For calculating the optical flow I used the `calcOpticalFlowFarneback` method from OpenCV, where the difference between each target frame *t* and frame *t-3* was used (delta 0.1 sec at 30FPS).\n\nDifferent data processing strategies were used depending on the dataset and label:\n- **Negative training samples**: 40 frames (every 1 second at 30FPS) were sampled and the frame (RGB), mask and flows were saved.\n- **Positive training samples**: For all frames between `time_of_alert` and `time_of_event` the frame (RGB), mask and flows were saved.\n- **Test**: For the last three frames the frame (RGB), mask and flows were saved.\n\nThe data is then stored as pytorch tensors in the following folder structure:\n```\n00001/\n├── flows/\n│   ├── 00.pt\n│   ├── 01.pt\n│   └── 02.pt\n├── frames/\n│   ├── 00.pt\n│   ├── 01.pt\n│   └── 02.pt\n└── masks/\n    ├── 00.pt\n    ├── 01.pt\n    └── 02.pt\n\n00002/\n...\n```\n\n## Model architecture\nThe model consists of two CNN backbones that are fed to a single classifier:\n- **CNN for mask and flow**: ResNet18 pre-trained on imagenet. Last fully connected layer is replaced by Identity operation.\n- **CNN for frame**: ResNet18 pre-trained on imagenet (CNN weights frozen). Last fully connected layer is replaced by a trainable linear layer of 512 to 512.\n- **Classifier**: Linear layer of (512 + 512) to 1.\n\nNote that the **CNN for frame** model is only fed the frame where `mask > 0`:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7949671%2F3329ea5bc464f75b2bdef4878c0bb089%2Fmasked_RGB.png?generation=1747120748458166&alt=media)\n\n## Training\nThe following hyperparameters were used:\n- Learning rate: 0.001\n- Optimizer: AdamW\n- Batch size: 32\n\nModels were trained with a 10% validation set for a maximum of 20 epochs, with early stopping after not improving the test accuracy for 8 epochs. The best model is then used for inference.\n\nFor every `n` samples in a single batch:\n1. A total of `n` random video's are selected.\n2. For each video we sample a frame together with the corresponding mask and flow.\n\nWe also applied the following transformations to the training data:\n```python\ntransform=T.Compose([\n    T.Lambda(pad_to_square),\n    T.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]),\n    T.RandomHorizontalFlip(),\n    T.RandomAffine(degrees=5, translate=(0.1, 0.1), scale=(0.9, 1.1), shear=5),\n])\n```\n\nwhere `pad_to_square` is defined as:\n```python\ndef pad_to_square(image):\n    w, h = image.size\n    max_dim = max(w, h)\n    pad_w = (max_dim - w) // 2\n    pad_h = (max_dim - h) // 2\n\n    return T.Pad((pad_w, pad_h, max_dim - w - pad_w, max_dim - h - pad_h), fill=0)(image)\n```\n[Resnet18](https://docs.pytorch.org/vision/main/models/generated/torchvision.models.resnet18.html) reshapes the input images to a square before the forward pass: `The images are resized to resize_size=[256] using interpolation=InterpolationMode.BILINEAR, followed by a central crop of crop_size=[224]`. By padding the image to a square first we prevent the shape of each vehicle to be warped.\n\nAll training runs are logged to a WandB repository which can be found [here](https://wandb.ai/maxzw/nexar-collision-prediction).\n\n## Inference\nWe applied the following transformations to the test data:\n```python\ntest_transform=T.Compose([\n    T.Lambda(pad_to_square),\n    T.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225]),\n])\n```\n\nThe model prediction is a weighted average of the predictions of the last three frames:\n- **Last frame**: 50%\n- **Second to last frame**: 30%\n- **Third to last frame**: 20%\n\nThe performance of individual models reached between 0.75 and 0.83 (see validation accuracy below - note that this is without a weighted average), but there was a lot of variation in the predictions. Therefore the final predictions are the averaged predictions of an ensemble of 2 models, which resulted in the final score.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7949671%2Fff86394430b189e77a2b81966fdcbf13%2Fvalidation%20accuracy.png?generation=1747120372267833&alt=media)\n\n## Code\n- A github repository containing all code required to train and evaluate a model can be found [here](https://github.com/maxzw/nexar-collision-prediction).\n- A Kaggle notebook to run inference on the test set can be found [here](https://www.kaggle.com/code/mzwager/inference-notebook-lb-0-861).\n\n## Hardware details\nAll code is run on a M2 Macbook Pro (16GB). You can find the specs [here](https://support.apple.com/en-us/111838). Using Pytorch Lightning helps with utilizing Apple's mps compute. \n\n## Limitations\n- Because the input images to both CNNs are multiplied with the vehicle mask, these images will be completely black (all values will be zero) if YOLO-v8 cannot find any vehicles. This means a good segmentation model is essential for this approach to work.\n- Because both CNNs work independently it's not possible for a model to learn relationships between the positition, direction and velocity of the vehicles and the RGB values, for example breaking lights. I experimented with using a non pre-trained CNN that accepted 6-channel images with shape (B, 6, H, W), but this did not perform well. This highlights the benefit of using pre-trained computer vision models.",
    "3200636": "Very nice!",
    "3200863": "Clever approach!",
    "3200891": "Cool, well done!",
    "3200911": "Great work!",
    "3202501": "Congratulations, Great job!",
    "3202715": "Proof that smart pre-processing is better than large models !",
    "3202785": "Using `x`, `y,` and `heading` provides structured spatial information that complements image and lidar data. While images and lidar offer rich visual and depth cues, they require complex processing to extract positional information. In contrast, `x`, `y`, and `heading` directly convey an object's location and orientation, which is invaluable for tasks like trajectory prediction and collision estimation.\n\nCombining these structured features with unstructured sensor data can enhance model performance by providing both precise localization and contextual understanding.",
    "3207979": "Love you approach"
  },
  "source": "meta"
}