{
  "id": 578354,
  "title": "9th Solution: ResNet18 + GRU |  0.824",
  "url": "/competitions/nexar-collision-prediction/discussion/578354",
  "author_name": "Ángel Jacinto Sánchez Ruiz",
  "post_date": "2025-05-10T08:56:05.608000",
  "votes": 7,
  "comment_count": 0,
  "views": 0,
  "content": "<h1>🚗 Collision Prediction with CNN-GRU | 5-Fold CV + TTA</h1>\n<p><strong><a href=\"https://www.kaggle.com/code/sacuscreed/fork-of-nexar-cnngru?scriptVersionId=225218134\" target=\"_blank\">Public Notebook Link</a></strong>  <br>\n<em>Score: 0.824 | Training Time: 2h 8m 31s · GPU T4 x2</em></p>\n<hr>\n<p>Thanks to <a href=\"https://www.kaggle.com/SLi\" target=\"_blank\">@SLi</a> for <a href=\"https://www.kaggle.com/competitions/nexar-collision-prediction/discussion/564606\" target=\"_blank\"><strong>Script and dataset of resized videos</strong></a></p>\n<hr>\n<h2>🔍 Approach Overview</h2>\n<p>We predict vehicle collisions using <strong>spatiotemporal video analysis</strong> with a hybrid CNN-GRU architecture. </p>\n<ul>\n<li>🎯 <strong>Stratified 5-fold CV</strong> Although best score was from fold 2, full CV achieved <strong>0.819</strong></li>\n<li>🌀 <strong>Test-Time Augmentation</strong> Horizontal Flips</li>\n<li>🖼️ <strong>ResNet18</strong> for spatial features + <strong>BiGRU</strong> for temporal modeling  </li>\n</ul>\n<hr>\n<h2>🧠 Model Architecture</h2>\n<pre><code> (nn.Module):\n     ():\n        ().__init__()\n        \n        .cnn = resnet18(weights=ResNet18_Weights.IMAGENET1K_V1)\n        .cnn.fc = nn.Identity()  \n        \n        .gru = nn.GRU(input_size=, hidden_size=, bidirectional=)\n        .classifier = nn.Linear(, )  \n\n     ():\n        \n        batch_size, T = x.shape[:]\n        x = x.view(-, *x.shape[:])  \n        x = .cnn(x)  \n        x = x.view(batch_size, T, -)  \n        x, _ = .gru(x)  \n        x = x[:, -]  \n        x = .classifier(x)\n         x\n</code></pre>\n<hr>\n<h2>Training</h2>\n<p>During training we cut random clips of <strong>T=64</strong> frames each <strong>S=3</strong> of them ending <strong>.5-1.5 seconds</strong> before <strong>time_of_event</strong> for positives. Leading to about <strong>6.4 seconds</strong> long at about <strong>10 fps</strong> clips. For negatives we just random cut at any possible time. Afterwards, a random <strong>30 degrees</strong> rotation is applied (to avoid too unrealistic images). Next, we cut 224x224 random crops from initial 256x256 resized clips. Finally we apply <strong>Gaussian Noise</strong>, <strong>Brightness</strong> and <strong>Contrast</strong> augmentations. We experimented with <strong>Mixup</strong> too but without clear improvements.</p>\n<p>Plane <strong>Cross Entropy Loss</strong> have been used.</p>",
  "messages": [
    {
      "id": 3198952,
      "postDate": "2025-05-10T08:56:05.610Z",
      "content": "<h1>🚗 Collision Prediction with CNN-GRU | 5-Fold CV + TTA</h1>\n<p><strong><a href=\"https://www.kaggle.com/code/sacuscreed/fork-of-nexar-cnngru?scriptVersionId=225218134\" target=\"_blank\">Public Notebook Link</a></strong>  <br>\n<em>Score: 0.824 | Training Time: 2h 8m 31s · GPU T4 x2</em></p>\n<hr>\n<p>Thanks to <a href=\"https://www.kaggle.com/SLi\" target=\"_blank\">@SLi</a> for <a href=\"https://www.kaggle.com/competitions/nexar-collision-prediction/discussion/564606\" target=\"_blank\"><strong>Script and dataset of resized videos</strong></a></p>\n<hr>\n<h2>🔍 Approach Overview</h2>\n<p>We predict vehicle collisions using <strong>spatiotemporal video analysis</strong> with a hybrid CNN-GRU architecture. </p>\n<ul>\n<li>🎯 <strong>Stratified 5-fold CV</strong> Although best score was from fold 2, full CV achieved <strong>0.819</strong></li>\n<li>🌀 <strong>Test-Time Augmentation</strong> Horizontal Flips</li>\n<li>🖼️ <strong>ResNet18</strong> for spatial features + <strong>BiGRU</strong> for temporal modeling  </li>\n</ul>\n<hr>\n<h2>🧠 Model Architecture</h2>\n<pre><code> (nn.Module):\n     ():\n        ().__init__()\n        \n        .cnn = resnet18(weights=ResNet18_Weights.IMAGENET1K_V1)\n        .cnn.fc = nn.Identity()  \n        \n        .gru = nn.GRU(input_size=, hidden_size=, bidirectional=)\n        .classifier = nn.Linear(, )  \n\n     ():\n        \n        batch_size, T = x.shape[:]\n        x = x.view(-, *x.shape[:])  \n        x = .cnn(x)  \n        x = x.view(batch_size, T, -)  \n        x, _ = .gru(x)  \n        x = x[:, -]  \n        x = .classifier(x)\n         x\n</code></pre>\n<hr>\n<h2>Training</h2>\n<p>During training we cut random clips of <strong>T=64</strong> frames each <strong>S=3</strong> of them ending <strong>.5-1.5 seconds</strong> before <strong>time_of_event</strong> for positives. Leading to about <strong>6.4 seconds</strong> long at about <strong>10 fps</strong> clips. For negatives we just random cut at any possible time. Afterwards, a random <strong>30 degrees</strong> rotation is applied (to avoid too unrealistic images). Next, we cut 224x224 random crops from initial 256x256 resized clips. Finally we apply <strong>Gaussian Noise</strong>, <strong>Brightness</strong> and <strong>Contrast</strong> augmentations. We experimented with <strong>Mixup</strong> too but without clear improvements.</p>\n<p>Plane <strong>Cross Entropy Loss</strong> have been used.</p>",
      "rawMarkdown": "# 🚗 Collision Prediction with CNN-GRU | 5-Fold CV + TTA\n\n**[Public Notebook Link](https://www.kaggle.com/code/sacuscreed/fork-of-nexar-cnngru?scriptVersionId=225218134)**  \n*Score: 0.824 | Training Time: 2h 8m 31s · GPU T4 x2*\n\n---\n\nThanks to @SLi for [**Script and dataset of resized videos**](https://www.kaggle.com/competitions/nexar-collision-prediction/discussion/564606)\n\n---\n\n## 🔍 Approach Overview  \nWe predict vehicle collisions using **spatiotemporal video analysis** with a hybrid CNN-GRU architecture. \n- 🎯 **Stratified 5-fold CV** Although best score was from fold 2, full CV achieved **0.819**\n- 🌀 **Test-Time Augmentation** Horizontal Flips\n- 🖼️ **ResNet18** for spatial features + **BiGRU** for temporal modeling  \n\n---\n\n## 🧠 Model Architecture  \n\n   \n    class myCNNGRU(nn.Module):\n        def __init__(self):\n            super().__init__()\n            # 2D Backbone\n            self.cnn = resnet18(weights=ResNet18_Weights.IMAGENET1K_V1)\n            self.cnn.fc = nn.Identity()  # Remove classification head\n            # Temporal Model\n            self.gru = nn.GRU(input_size=512, hidden_size=128, bidirectional=True)\n            self.classifier = nn.Linear(256, 2)  # Bidirectional → 256\n\n        def forward(self, x):\n            # x: [batch, frames, C, H, W]\n            batch_size, T = x.shape[:2]\n            x = x.view(-1, *x.shape[2:])  # Merge batch + time\n            x = self.cnn(x)  # [batch*T, 512]\n            x = x.view(batch_size, T, -1)  # Unmerge\n            x, _ = self.gru(x)  # [batch, T, 256]\n            x = x[:, -1]  # Last timestep\n            x = self.classifier(x)\n            return x\n\n---\n\n## Training\n\nDuring training we cut random clips of **T=64** frames each **S=3** of them ending **.5-1.5 seconds** before **time_of_event** for positives. Leading to about **6.4 seconds** long at about **10 fps** clips. For negatives we just random cut at any possible time. Afterwards, a random **30 degrees** rotation is applied (to avoid too unrealistic images). Next, we cut 224x224 random crops from initial 256x256 resized clips. Finally we apply **Gaussian Noise**, **Brightness** and **Contrast** augmentations. We experimented with **Mixup** too but without clear improvements.\n\nPlane **Cross Entropy Loss** have been used.",
      "votes": 7
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3198952": "# 🚗 Collision Prediction with CNN-GRU | 5-Fold CV + TTA\n\n**[Public Notebook Link](https://www.kaggle.com/code/sacuscreed/fork-of-nexar-cnngru?scriptVersionId=225218134)**  \n*Score: 0.824 | Training Time: 2h 8m 31s · GPU T4 x2*\n\n---\n\nThanks to @SLi for [**Script and dataset of resized videos**](https://www.kaggle.com/competitions/nexar-collision-prediction/discussion/564606)\n\n---\n\n## 🔍 Approach Overview  \nWe predict vehicle collisions using **spatiotemporal video analysis** with a hybrid CNN-GRU architecture. \n- 🎯 **Stratified 5-fold CV** Although best score was from fold 2, full CV achieved **0.819**\n- 🌀 **Test-Time Augmentation** Horizontal Flips\n- 🖼️ **ResNet18** for spatial features + **BiGRU** for temporal modeling  \n\n---\n\n## 🧠 Model Architecture  \n\n   \n    class myCNNGRU(nn.Module):\n        def __init__(self):\n            super().__init__()\n            # 2D Backbone\n            self.cnn = resnet18(weights=ResNet18_Weights.IMAGENET1K_V1)\n            self.cnn.fc = nn.Identity()  # Remove classification head\n            # Temporal Model\n            self.gru = nn.GRU(input_size=512, hidden_size=128, bidirectional=True)\n            self.classifier = nn.Linear(256, 2)  # Bidirectional → 256\n\n        def forward(self, x):\n            # x: [batch, frames, C, H, W]\n            batch_size, T = x.shape[:2]\n            x = x.view(-1, *x.shape[2:])  # Merge batch + time\n            x = self.cnn(x)  # [batch*T, 512]\n            x = x.view(batch_size, T, -1)  # Unmerge\n            x, _ = self.gru(x)  # [batch, T, 256]\n            x = x[:, -1]  # Last timestep\n            x = self.classifier(x)\n            return x\n\n---\n\n## Training\n\nDuring training we cut random clips of **T=64** frames each **S=3** of them ending **.5-1.5 seconds** before **time_of_event** for positives. Leading to about **6.4 seconds** long at about **10 fps** clips. For negatives we just random cut at any possible time. Afterwards, a random **30 degrees** rotation is applied (to avoid too unrealistic images). Next, we cut 224x224 random crops from initial 256x256 resized clips. Finally we apply **Gaussian Noise**, **Brightness** and **Contrast** augmentations. We experimented with **Mixup** too but without clear improvements.\n\nPlane **Cross Entropy Loss** have been used."
  }
}