{
  "id": 578673,
  "title": "[Private 0.853]My Solution",
  "url": "/competitions/nexar-collision-prediction/writeups/hi-f-private-0-853-my-solution",
  "author_name": "",
  "post_date": "2025-05-12T14:32:24.370Z",
  "votes": 4,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Thank you for this incredibly meaningful competition. I also wish to express my deepest respect and congratulations to the winners for their remarkable efforts. This was my first competition dealing with video classification, and I struggled a great deal, but I would like to share an approach.</p>\n<p>I began by thinking about how to help the model understand the concept of “what constitutes a collision.” One idea that we, as humans, understand but the model might not is that for a collision to occur, cars must come close to each other, and we also need to consider three-dimensional structure (i.e., depth). Therefore, I decided to use a model called “deepanything” to estimate depth from the images and then feed those depth estimates into the model.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4744935%2Fce56febf4fb8bbecec783b03f59078e0%2F2025-05-12%2023.22.26.png?generation=1747059781844780&amp;alt=media\" alt=\"source+deepanything\"></p>\n<p>Having incorporated depth with deepanything, my next step was to help the model consider horizontal or lateral movement, so I introduced optical flow using OpenCV. Optical flow produces two channels of data, and when combined with the depth information, we get three channels in total. By stacking these channels, I created images that incorporate all three types of information.</p>\n<p>Feeding these into the model resulted in fairly high performance, but it still wasn’t enough to catch up to the top performers. That led me to revisit the question of “what exactly is a collision?”</p>\n<p>Upon reviewing the competition videos, I noticed that they primarily involved collisions between cars. I hypothesized that making the model’s task easier might simply involve extracting the cars and feeding those extracted images to the model, so I used YOLO for car segmentation.<br>\nAt this point, because I was already dealing with three-channel images (depth and the two-channel optical flow), I decided to stack the car-segmented images vertically, as shown below. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4744935%2Fe3e4bbe21102c6c6d07d83f357f4e762%2F2025-05-12%2023.28.58.png?generation=1747060153813246&amp;alt=media\" alt=\"with_yolo\"></p>\n<p>Training the model with this approach yielded a single model that scored 0.766 on the public leaderboard and 0.852 on the private leaderboard.</p>\n<p>Finally, I would like to express my sincere gratitude to the host for organizing such a valuable competition, and I offer my deepest respect to all the fellow Kagglers who competed alongside me.</p>",
  "messages": [
    {
      "id": "3200405",
      "postDate": "05/12/2025 14:29:34",
      "content": "<p>Thank you for this incredibly meaningful competition. I also wish to express my deepest respect and congratulations to the winners for their remarkable efforts. This was my first competition dealing with video classification, and I struggled a great deal, but I would like to share an approach.</p>\n<p>I began by thinking about how to help the model understand the concept of “what constitutes a collision.” One idea that we, as humans, understand but the model might not is that for a collision to occur, cars must come close to each other, and we also need to consider three-dimensional structure (i.e., depth). Therefore, I decided to use a model called “deepanything” to estimate depth from the images and then feed those depth estimates into the model.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4744935%2Fce56febf4fb8bbecec783b03f59078e0%2F2025-05-12%2023.22.26.png?generation=1747059781844780&amp;alt=media\" alt=\"source+deepanything\"></p>\n<p>Having incorporated depth with deepanything, my next step was to help the model consider horizontal or lateral movement, so I introduced optical flow using OpenCV. Optical flow produces two channels of data, and when combined with the depth information, we get three channels in total. By stacking these channels, I created images that incorporate all three types of information.</p>\n<p>Feeding these into the model resulted in fairly high performance, but it still wasn’t enough to catch up to the top performers. That led me to revisit the question of “what exactly is a collision?”</p>\n<p>Upon reviewing the competition videos, I noticed that they primarily involved collisions between cars. I hypothesized that making the model’s task easier might simply involve extracting the cars and feeding those extracted images to the model, so I used YOLO for car segmentation.<br>\nAt this point, because I was already dealing with three-channel images (depth and the two-channel optical flow), I decided to stack the car-segmented images vertically, as shown below. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4744935%2Fe3e4bbe21102c6c6d07d83f357f4e762%2F2025-05-12%2023.28.58.png?generation=1747060153813246&amp;alt=media\" alt=\"with_yolo\"></p>\n<p>Training the model with this approach yielded a single model that scored 0.766 on the public leaderboard and 0.852 on the private leaderboard.</p>\n<p>Finally, I would like to express my sincere gratitude to the host for organizing such a valuable competition, and I offer my deepest respect to all the fellow Kagglers who competed alongside me.</p>",
      "rawMarkdown": "Thank you for this incredibly meaningful competition. I also wish to express my deepest respect and congratulations to the winners for their remarkable efforts. This was my first competition dealing with video classification, and I struggled a great deal, but I would like to share an approach.\n\nI began by thinking about how to help the model understand the concept of “what constitutes a collision.” One idea that we, as humans, understand but the model might not is that for a collision to occur, cars must come close to each other, and we also need to consider three-dimensional structure (i.e., depth). Therefore, I decided to use a model called “deepanything” to estimate depth from the images and then feed those depth estimates into the model.\n\n![source+deepanything](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4744935%2Fce56febf4fb8bbecec783b03f59078e0%2F2025-05-12%2023.22.26.png?generation=1747059781844780&alt=media)\n\nHaving incorporated depth with deepanything, my next step was to help the model consider horizontal or lateral movement, so I introduced optical flow using OpenCV. Optical flow produces two channels of data, and when combined with the depth information, we get three channels in total. By stacking these channels, I created images that incorporate all three types of information.\n\nFeeding these into the model resulted in fairly high performance, but it still wasn’t enough to catch up to the top performers. That led me to revisit the question of “what exactly is a collision?”\n\nUpon reviewing the competition videos, I noticed that they primarily involved collisions between cars. I hypothesized that making the model’s task easier might simply involve extracting the cars and feeding those extracted images to the model, so I used YOLO for car segmentation.\nAt this point, because I was already dealing with three-channel images (depth and the two-channel optical flow), I decided to stack the car-segmented images vertically, as shown below. \n\n![with_yolo](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4744935%2Fe3e4bbe21102c6c6d07d83f357f4e762%2F2025-05-12%2023.28.58.png?generation=1747060153813246&alt=media)\n\nTraining the model with this approach yielded a single model that scored 0.766 on the public leaderboard and 0.852 on the private leaderboard.\n\nFinally, I would like to express my sincere gratitude to the host for organizing such a valuable competition, and I offer my deepest respect to all the fellow Kagglers who competed alongside me.",
      "votes": null
    },
    {
      "id": "3200432",
      "postDate": "05/12/2025 15:05:59",
      "content": "<p>Wow! This is a really cool approach!</p>",
      "rawMarkdown": "Wow! This is a really cool approach!",
      "votes": null
    },
    {
      "id": "3200445",
      "postDate": "05/12/2025 15:23:53",
      "content": "<p>Thank you! I read your explanation of your approach, and it taught me a lot in ways I had never considered before. Congratulations on your victory!</p>",
      "rawMarkdown": "Thank you! I read your explanation of your approach, and it taught me a lot in ways I had never considered before. Congratulations on your victory!",
      "votes": null
    },
    {
      "id": "3200634",
      "postDate": "05/12/2025 19:51:11",
      "content": "<p>Congratulations!</p>\n<p>I followed a similar approach — my input consists of RGB, 2 optical flow channels, 1 depth channel (from depth Anything V2), and 1 mask channel (Vehicle+Lane) ( 7 channel ).</p>\n<p>May I ask which backbone model you used?</p>",
      "rawMarkdown": "Congratulations!\n\nI followed a similar approach — my input consists of RGB, 2 optical flow channels, 1 depth channel (from depth Anything V2), and 1 mask channel (Vehicle+Lane) ( 7 channel ).\n\nMay I ask which backbone model you used?",
      "votes": null
    },
    {
      "id": "3200955",
      "postDate": "05/13/2025 09:02:47",
      "content": "<p>I used CNN+RNN approach and employed ResNet18 for the CNN.</p>",
      "rawMarkdown": "I used CNN+RNN approach and employed ResNet18 for the CNN.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3200432,
      "author_name": "paulendresen76",
      "author_url": "",
      "post_date": "05/12/2025 15:05:59",
      "content": "<p>Wow! This is a really cool approach!</p>",
      "votes": null,
      "replies": [
        {
          "id": 3200445,
          "author_name": "hiroakifukuse",
          "author_url": "",
          "post_date": "05/12/2025 15:23:53",
          "content": "<p>Thank you! I read your explanation of your approach, and it taught me a lot in ways I had never considered before. Congratulations on your victory!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3200634,
      "author_name": "gopimali",
      "author_url": "",
      "post_date": "05/12/2025 19:51:11",
      "content": "<p>Congratulations!</p>\n<p>I followed a similar approach — my input consists of RGB, 2 optical flow channels, 1 depth channel (from depth Anything V2), and 1 mask channel (Vehicle+Lane) ( 7 channel ).</p>\n<p>May I ask which backbone model you used?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3200955,
          "author_name": "hiroakifukuse",
          "author_url": "",
          "post_date": "05/13/2025 09:02:47",
          "content": "<p>I used CNN+RNN approach and employed ResNet18 for the CNN.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3200405": "Thank you for this incredibly meaningful competition. I also wish to express my deepest respect and congratulations to the winners for their remarkable efforts. This was my first competition dealing with video classification, and I struggled a great deal, but I would like to share an approach.\n\nI began by thinking about how to help the model understand the concept of “what constitutes a collision.” One idea that we, as humans, understand but the model might not is that for a collision to occur, cars must come close to each other, and we also need to consider three-dimensional structure (i.e., depth). Therefore, I decided to use a model called “deepanything” to estimate depth from the images and then feed those depth estimates into the model.\n\n![source+deepanything](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4744935%2Fce56febf4fb8bbecec783b03f59078e0%2F2025-05-12%2023.22.26.png?generation=1747059781844780&alt=media)\n\nHaving incorporated depth with deepanything, my next step was to help the model consider horizontal or lateral movement, so I introduced optical flow using OpenCV. Optical flow produces two channels of data, and when combined with the depth information, we get three channels in total. By stacking these channels, I created images that incorporate all three types of information.\n\nFeeding these into the model resulted in fairly high performance, but it still wasn’t enough to catch up to the top performers. That led me to revisit the question of “what exactly is a collision?”\n\nUpon reviewing the competition videos, I noticed that they primarily involved collisions between cars. I hypothesized that making the model’s task easier might simply involve extracting the cars and feeding those extracted images to the model, so I used YOLO for car segmentation.\nAt this point, because I was already dealing with three-channel images (depth and the two-channel optical flow), I decided to stack the car-segmented images vertically, as shown below. \n\n![with_yolo](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4744935%2Fe3e4bbe21102c6c6d07d83f357f4e762%2F2025-05-12%2023.28.58.png?generation=1747060153813246&alt=media)\n\nTraining the model with this approach yielded a single model that scored 0.766 on the public leaderboard and 0.852 on the private leaderboard.\n\nFinally, I would like to express my sincere gratitude to the host for organizing such a valuable competition, and I offer my deepest respect to all the fellow Kagglers who competed alongside me.",
    "3200432": "Wow! This is a really cool approach!",
    "3200445": "Thank you! I read your explanation of your approach, and it taught me a lot in ways I had never considered before. Congratulations on your victory!",
    "3200634": "Congratulations!\n\nI followed a similar approach — my input consists of RGB, 2 optical flow channels, 1 depth channel (from depth Anything V2), and 1 mask channel (Vehicle+Lane) ( 7 channel ).\n\nMay I ask which backbone model you used?",
    "3200955": "I used CNN+RNN approach and employed ResNet18 for the CNN."
  },
  "source": "meta"
}