{
  "id": 563503,
  "title": "Are there missing alerts in the data?",
  "url": "/competitions/nexar-collision-prediction/discussion/563503",
  "author_name": "PaulEE",
  "post_date": "2025-02-17T15:44:57.388000",
  "votes": 6,
  "comment_count": 25,
  "views": 0,
  "content": "<p>I am debugging the False Positives I got on the training set, where my model thinks it should be an alert - in very many of those cases it seems to me an alert \"should\" have been thrown, even though it did not end up in a near-collision or collision afterwards - for example see video 01390, 18 seconds in. Car coming in from the left in front of the car, and driver brakes hard. Question is, are alerts that car/dashcam maybe threw removed if it didn't turn into a near collision or accident? (screenshot: <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262906%2F036c04833de4c30209c57a84cfc31927%2FScreenshot%202025-02-17%20084055.png?generation=1739807090569402&amp;alt=media\" alt=\"\">each vertical line is a separate video, and green is the highlighted frames that are False Positives).</p>",
  "messages": [
    {
      "id": 3126732,
      "postDate": "2025-02-17T15:44:57.390Z",
      "content": "<p>I am debugging the False Positives I got on the training set, where my model thinks it should be an alert - in very many of those cases it seems to me an alert \"should\" have been thrown, even though it did not end up in a near-collision or collision afterwards - for example see video 01390, 18 seconds in. Car coming in from the left in front of the car, and driver brakes hard. Question is, are alerts that car/dashcam maybe threw removed if it didn't turn into a near collision or accident? (screenshot: <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262906%2F036c04833de4c30209c57a84cfc31927%2FScreenshot%202025-02-17%20084055.png?generation=1739807090569402&amp;alt=media\" alt=\"\">each vertical line is a separate video, and green is the highlighted frames that are False Positives).</p>",
      "rawMarkdown": "I am debugging the False Positives I got on the training set, where my model thinks it should be an alert - in very many of those cases it seems to me an alert \"should\" have been thrown, even though it did not end up in a near-collision or collision afterwards - for example see video 01390, 18 seconds in. Car coming in from the left in front of the car, and driver brakes hard. Question is, are alerts that car/dashcam maybe threw removed if it didn't turn into a near collision or accident? (screenshot: ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262906%2F036c04833de4c30209c57a84cfc31927%2FScreenshot%202025-02-17%20084055.png?generation=1739807090569402&alt=media)each vertical line is a separate video, and green is the highlighted frames that are False Positives).",
      "votes": 6
    },
    {
      "id": 3126796,
      "postDate": "2025-02-17T17:17:48.240Z",
      "content": "<p>Alert times are only available for positive cases (collisions and near-collisions), and there is only one alert time per video (the closest to the event). Alert times are helpful to understand how ahead in time it is expected to anticipate that an event is about to happen (based on annotations from humans); e.g. if <code>event_time - alert_time = 0.7</code> it is not expected that a model can detect a potential accident 1 second before it happens. </p>",
      "rawMarkdown": "Alert times are only available for positive cases (collisions and near-collisions), and there is only one alert time per video (the closest to the event). Alert times are helpful to understand how ahead in time it is expected to anticipate that an event is about to happen (based on annotations from humans); e.g. if `event_time - alert_time = 0.7` it is not expected that a model can detect a potential accident 1 second before it happens. ",
      "votes": 1,
      "replies": [
        {
          "id": 3126807,
          "postDate": "2025-02-17T17:36:11.887Z",
          "content": "<p>I understand, but for example video 01390 (and some other videos that has no events), 18 seconds in, it seems like an alert should have been there comparing to videos that has an alert - driver hits the brakes hard to not hit car that suddenly appeared in front, could be classified as near collision - Question is if there are room/possibility there could be \"missing\" labels?</p>",
          "rawMarkdown": "I understand, but for example video 01390 (and some other videos that has no events), 18 seconds in, it seems like an alert should have been there comparing to videos that has an alert - driver hits the brakes hard to not hit car that suddenly appeared in front, could be classified as near collision - Question is if there are room/possibility there could be \"missing\" labels?",
          "votes": 1,
          "replies": [
            {
              "id": 3126884,
              "postDate": "2025-02-17T18:51:57.530Z",
              "content": "<p>I agree that video 01390 could be classified as a near-collision. This is an annotation error. While we did our best to minimize annotation errors, they can still happen. Please, consider this as noise in the dataset. If you know about more cases, please post them here. Thank you 🙏  </p>",
              "rawMarkdown": "I agree that video 01390 could be classified as a near-collision. This is an annotation error. While we did our best to minimize annotation errors, they can still happen. Please, consider this as noise in the dataset. If you know about more cases, please post them here. Thank you 🙏  ",
              "votes": 1
            },
            {
              "id": 3126891,
              "postDate": "2025-02-17T19:00:52.820Z",
              "content": "<p>Ok, thanks!</p>",
              "rawMarkdown": "Ok, thanks!"
            },
            {
              "id": 3135050,
              "postDate": "2025-02-27T02:09:33.650Z",
              "content": "<p><a href=\"https://www.kaggle.com/danielcmoura\" target=\"_blank\">@danielcmoura</a> I am finding many others - for example train video 00573: 00573,10.124,8.293,1 - there is no alert that should be in 10 seconds as I can see? Or is it a car that comes in under the line that is blurred? (at 7 seconds). <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262906%2F36ae2f65dffb015f43f4c18e7bbfe0f2%2FScreenshot%202025-02-26%20190903.png?generation=1740622170111724&amp;alt=media\" alt=\"\"></p>",
              "rawMarkdown": "@danielcmoura I am finding many others - for example train video 00573: 00573,10.124,8.293,1 - there is no alert that should be in 10 seconds as I can see? Or is it a car that comes in under the line that is blurred? (at 7 seconds). ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262906%2F36ae2f65dffb015f43f4c18e7bbfe0f2%2FScreenshot%202025-02-26%20190903.png?generation=1740622170111724&alt=media)"
            },
            {
              "id": 3135052,
              "postDate": "2025-02-27T02:11:28.980Z",
              "content": "<p>I am looking at all the high confident but wrong cases that my model outputs, so many of these are easy to find. I'll just either delete or relabel those.</p>",
              "rawMarkdown": "I am looking at all the high confident but wrong cases that my model outputs, so many of these are easy to find. I'll just either delete or relabel those."
            },
            {
              "id": 3141197,
              "postDate": "2025-03-05T10:10:19.257Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/paulendresen76\" target=\"_blank\">@paulendresen76</a> . <code>00573</code> is an edge-case… there is a near-collision on that video, but it is obfuscated by the blurring we do for privacy protection. It would make sense to me to exclude this video from the training data since the blurring is making extremely difficult to detect the event. </p>",
              "rawMarkdown": "Hi @paulendresen76 . `00573` is an edge-case... there is a near-collision on that video, but it is obfuscated by the blurring we do for privacy protection. It would make sense to me to exclude this video from the training data since the blurring is making extremely difficult to detect the event. "
            },
            {
              "id": 3141455,
              "postDate": "2025-03-05T15:42:28.130Z",
              "content": "<p>Thanks for confirming! That is what I am doing in these cases :-)</p>",
              "rawMarkdown": "Thanks for confirming! That is what I am doing in these cases :-)"
            },
            {
              "id": 3163104,
              "postDate": "2025-03-30T10:25:31.473Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 3179755,
      "postDate": "2025-04-15T16:42:10.790Z",
      "content": "<p>Hi PaulEEE,</p>\n<p>Thanks for sharing your great work! Would you mind uploading the clean and preprocessed dataset you used? It would be really helpful for everyone trying to learn from your approach.</p>\n<p>Also, if possible, could you briefly explain the preprocessing steps you applied to the data? That context would make it easier for others to follow along and replicate your results.</p>\n<p>Appreciate your help!</p>",
      "rawMarkdown": "Hi PaulEEE,\n\nThanks for sharing your great work! Would you mind uploading the clean and preprocessed dataset you used? It would be really helpful for everyone trying to learn from your approach.\n\nAlso, if possible, could you briefly explain the preprocessing steps you applied to the data? That context would make it easier for others to follow along and replicate your results.\n\nAppreciate your help!",
      "replies": [
        {
          "id": 3181184,
          "postDate": "2025-04-17T14:52:30.633Z",
          "content": "<p>Hi, I will publish everything after the competition is over as an article and on github -  I converted the videos into a 3LC table, where each frame represents a training sample, so I can use the tools we build. In the pytorch dataset that reads a 3LC table I load N frames leading up to the training sample frame, such that the actual training sample is N frames. Then I used 3LC to see how the model learned/struggled for individual frames(training sample) for my training runs (3LC capture metrics per training sample during training), and I really just worked with deleting ambigous individual ranges of frames(training samples) and weighting others that the model needed to focus on since it struggled to learn those scenarios. Only working with the data. It's not that easy to \"export\" back a dataset, as weight is 0 for all frames that I didnt want to train on, and varying weight for the other frames that needs more attention. I could of course export that as a value per frame per video. </p>",
          "rawMarkdown": "Hi, I will publish everything after the competition is over as an article and on github -  I converted the videos into a 3LC table, where each frame represents a training sample, so I can use the tools we build. In the pytorch dataset that reads a 3LC table I load N frames leading up to the training sample frame, such that the actual training sample is N frames. Then I used 3LC to see how the model learned/struggled for individual frames(training sample) for my training runs (3LC capture metrics per training sample during training), and I really just worked with deleting ambigous individual ranges of frames(training samples) and weighting others that the model needed to focus on since it struggled to learn those scenarios. Only working with the data. It's not that easy to \"export\" back a dataset, as weight is 0 for all frames that I didnt want to train on, and varying weight for the other frames that needs more attention. I could of course export that as a value per frame per video. ",
          "replies": [
            {
              "id": 3181187,
              "postDate": "2025-04-17T14:56:45.783Z",
              "content": "<p>It's worth mentioning that I didn't \"preprocess\", as I started training on all data, capture metrics, refine data, train again - iterating to understand how to delete and weight data, based on how the model reacted to the data it was fed. So I processed the data during the training.</p>",
              "rawMarkdown": "It's worth mentioning that I didn't \"preprocess\", as I started training on all data, capture metrics, refine data, train again - iterating to understand how to delete and weight data, based on how the model reacted to the data it was fed. So I processed the data during the training."
            }
          ]
        }
      ]
    },
    {
      "id": 3157107,
      "postDate": "2025-03-23T01:20:33.967Z",
      "content": "<p>Hi,I noticed that your image display interface is very elegant and straightforward. It doesn’t seem to be directly from any open-source library I’ve come across. I was wondering, by any chance, could you let me know what software or tool you used to create it?</p>",
      "rawMarkdown": "Hi,I noticed that your image display interface is very elegant and straightforward. It doesn’t seem to be directly from any open-source library I’ve come across. I was wondering, by any chance, could you let me know what software or tool you used to create it?",
      "replies": [
        {
          "id": 3157550,
          "postDate": "2025-03-23T15:16:19.793Z",
          "content": "<p>Hi it is free for non-commercial and Academia / Research - and also Kaggle competitions :-)<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262906%2F1c1d69aea4a2e50515c96efbdad0cbca%2FScreenshot%202025-03-23%20091436.png?generation=1742742925046909&amp;alt=media\" alt=\"\"></p>\n<p><a href=\"https://3lc.ai\" target=\"_blank\">https://3lc.ai</a> and https:/docs.3lc.ai</p>",
          "rawMarkdown": "Hi it is free for non-commercial and Academia / Research - and also Kaggle competitions :-)![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262906%2F1c1d69aea4a2e50515c96efbdad0cbca%2FScreenshot%202025-03-23%20091436.png?generation=1742742925046909&alt=media)\n\nhttps://3lc.ai and https:/docs.3lc.ai",
          "votes": 1,
          "replies": [
            {
              "id": 3157558,
              "postDate": "2025-03-23T15:27:07.370Z",
              "content": "<p>Disclaimer, I am the founder of 3LC :-)</p>",
              "rawMarkdown": "Disclaimer, I am the founder of 3LC :-)"
            },
            {
              "id": 3157850,
              "postDate": "2025-03-24T00:03:11.223Z",
              "content": "<p>LOL, thank you so much for information =)</p>",
              "rawMarkdown": "LOL, thank you so much for information =)"
            },
            {
              "id": 3161880,
              "postDate": "2025-03-28T14:40:59.197Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 3161881,
              "postDate": "2025-03-28T14:41:27.563Z",
              "content": "<p>Awesome! Your excellent skills amaze me. I have become the 45th fan（youtube) of 3lc and your fan.</p>",
              "rawMarkdown": "Awesome! Your excellent skills amaze me. I have become the 45th fan（youtube) of 3lc and your fan."
            }
          ]
        },
        {
          "id": 3157559,
          "postDate": "2025-03-23T15:28:35.477Z",
          "content": "<p>I register my data (URL's to image path) and additional data per frame (Time to Event etc.) as 3LC Tables with the 3LC python pacakge, then I can plot anything against each other in the dashboard. Everything runs locally. (In the training script I open the 3LC tables as torch datasets and just add a map function to return the relevant data/columns for the model. What is nice I can also capture runs/metrics per sample during training such that I can investigate my data together with how the model performs on each frame. I do use more than one frame for my solution, but I just load previous images in the map function. Here for example I plot which frames are correctly predicted vs wrong in models embedding space.<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262906%2F0f96f281552e348082a869fe9fb85545%2FScreenshot%202025-03-23%20092249.jpg?generation=1742743586652254&amp;alt=media\" alt=\"\"></p>\n<p>I would be happy to help out - but everything shared has to be shared with everybody according to the rules, so please ask questions here :-) </p>",
          "rawMarkdown": "I register my data (URL's to image path) and additional data per frame (Time to Event etc.) as 3LC Tables with the 3LC python pacakge, then I can plot anything against each other in the dashboard. Everything runs locally. (In the training script I open the 3LC tables as torch datasets and just add a map function to return the relevant data/columns for the model. What is nice I can also capture runs/metrics per sample during training such that I can investigate my data together with how the model performs on each frame. I do use more than one frame for my solution, but I just load previous images in the map function. Here for example I plot which frames are correctly predicted vs wrong in models embedding space.![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262906%2F0f96f281552e348082a869fe9fb85545%2FScreenshot%202025-03-23%20092249.jpg?generation=1742743586652254&alt=media)\n\nI would be happy to help out - but everything shared has to be shared with everybody according to the rules, so please ask questions here :-) ",
          "replies": [
            {
              "id": 3157636,
              "postDate": "2025-03-23T16:28:56.833Z",
              "content": "<p>So I see that in the area my model struggles I can investigate further and color frames with if the frames are between alert and event - and interestingly I see that the models embeddings in the area I have most error is 50/50 event/no event frames… and I can add loss as a 3rd axis to understand more. Not good that they are that close in embedding space, so what I've been doing is understand the training samples by selecting and looking at that data and in the UI and weight areas of samples I think the model should focus more on (in batch) and delete samples / ambigous cases where labelling might not be perfect. Then I use the new table revision for next training run, I am taking a 100% iterative data-centric approach to help the model understand the difference between when it should detect an event or not, emphasizing the hard data (for example when a car is close but drives straight(no event) vs close and its trajectory is bad), the model I am using is more or less off-the-shelf, simple and standard. If I had more data I would of course grow my dataset with frames in the area where it struggles the most in an automated fashion after understanding what it has a hard time to learn. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262906%2Fea4851c41220e9aec93699940e9c52eb%2FScreenshot%202025-03-23%20101126.png?generation=1742747240219431&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262906%2F217783d39fd0abfd89dfb4ebee102157%2FScreenshot%202025-03-23%20102234.png?generation=1742747250435877&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262906%2F6b4642cb068d9ca1aa6345d057995042%2FScreenshot%202025-03-23%20102246.png?generation=1742747259993111&amp;alt=media\" alt=\"\"></p>",
              "rawMarkdown": "So I see that in the area my model struggles I can investigate further and color frames with if the frames are between alert and event - and interestingly I see that the models embeddings in the area I have most error is 50/50 event/no event frames... and I can add loss as a 3rd axis to understand more. Not good that they are that close in embedding space, so what I've been doing is understand the training samples by selecting and looking at that data and in the UI and weight areas of samples I think the model should focus more on (in batch) and delete samples / ambigous cases where labelling might not be perfect. Then I use the new table revision for next training run, I am taking a 100% iterative data-centric approach to help the model understand the difference between when it should detect an event or not, emphasizing the hard data (for example when a car is close but drives straight(no event) vs close and its trajectory is bad), the model I am using is more or less off-the-shelf, simple and standard. If I had more data I would of course grow my dataset with frames in the area where it struggles the most in an automated fashion after understanding what it has a hard time to learn. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262906%2Fea4851c41220e9aec93699940e9c52eb%2FScreenshot%202025-03-23%20101126.png?generation=1742747240219431&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262906%2F217783d39fd0abfd89dfb4ebee102157%2FScreenshot%202025-03-23%20102234.png?generation=1742747250435877&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262906%2F6b4642cb068d9ca1aa6345d057995042%2FScreenshot%202025-03-23%20102246.png?generation=1742747259993111&alt=media)"
            },
            {
              "id": 3157743,
              "postDate": "2025-03-23T19:00:59.710Z",
              "content": "<p>Hi,<br>\nThank you for sharing these insights!</p>\n<p>If I may ask for some clarification—are you suggesting that removing ambiguous samples is the key factor in improving performance?</p>\n<p>Would this involve initially training the model, identifying areas where it struggles, removing those samples, and then retraining?</p>\n<p>Additionally, if you’re comfortable sharing, did you find that model selection had a significant impact, or was the iterative process of refining the dataset the more crucial factor?</p>\n<p>I truly appreciate your time and insights. Thanks again!</p>",
              "rawMarkdown": "Hi,\nThank you for sharing these insights!\n\nIf I may ask for some clarification—are you suggesting that removing ambiguous samples is the key factor in improving performance?\n\nWould this involve initially training the model, identifying areas where it struggles, removing those samples, and then retraining?\n\nAdditionally, if you’re comfortable sharing, did you find that model selection had a significant impact, or was the iterative process of refining the dataset the more crucial factor?\n\nI truly appreciate your time and insights. Thanks again!"
            },
            {
              "id": 3157786,
              "postDate": "2025-03-23T20:00:35.753Z",
              "content": "<p>Hi, <a href=\"https://www.kaggle.com/valentinarizzello\" target=\"_blank\">@valentinarizzello</a> </p>\n<p>Not only ambigious samples, but that as well. They main impact is weighting each sample based on importance, so I added a weight column to the 3LC table so I can assign weights in the UI (on many samples at once), and then use torch WeightedRandomSampler that takes that column in, so it's quick to experiment, for example setting weights to 10 will show that sample 10 times as often. But also reducing the dataset size helped a lot -  I found that there are a lot of correctly labelled samples that does not contribute to the model getting better and is rather confusing the model, but I also removed ambigious samples. I have now 1172007 frames left of the 1701166 in the train data, so deleted around 32% of the frames. (However when a frame/sample is loaded, I load the existing N frames before that frame as well as input to my model, so some of these frames might be the deleted ones as training samples, but can still be used).</p>\n<p>Training is relatively quick, each epoch I train on 6000 samples (which loads N previous frames for each sample), and it usually converges after 10 epochs, so I am not even using the 1 million samples.. My biggest struggle has been to optimize for which validation samples I validate against, because it takes 4 hours to run through all 231.105 validation samples I have (which gives the most accurate result), but I've been decimating that as well and trying to find the best representation that represents the full val set, so I can validate each epoch in a minute instead of 4 hours.</p>\n<p>To this: \"Would this involve initially training the model, identifying areas where it struggles, removing those samples, and then retraining?\" - yes exactly, except not only removing, but also weighting samples. Rapid experimentation and very many data iterations :-)</p>\n<p>To this: \"Additionally, if you’re comfortable sharing, did you find that model selection had a significant impact, or was the iterative process of refining the dataset the more crucial factor?\" <br>\nI used a pretty standard Multiscale Vision Transformer first time, and didn't try anything else, and manipulating the training data took score so far from 0.72 to 0.89.</p>\n<p>Now I am working on how to reduce training samples to under 100.000 frames, and understanding how I can get the model to separate even better where it still has challenges/hard time separating data in embedding space.</p>",
              "rawMarkdown": "Hi, @valentinarizzello \n\nNot only ambigious samples, but that as well. They main impact is weighting each sample based on importance, so I added a weight column to the 3LC table so I can assign weights in the UI (on many samples at once), and then use torch WeightedRandomSampler that takes that column in, so it's quick to experiment, for example setting weights to 10 will show that sample 10 times as often. But also reducing the dataset size helped a lot -  I found that there are a lot of correctly labelled samples that does not contribute to the model getting better and is rather confusing the model, but I also removed ambigious samples. I have now 1172007 frames left of the 1701166 in the train data, so deleted around 32% of the frames. (However when a frame/sample is loaded, I load the existing N frames before that frame as well as input to my model, so some of these frames might be the deleted ones as training samples, but can still be used).\n\nTraining is relatively quick, each epoch I train on 6000 samples (which loads N previous frames for each sample), and it usually converges after 10 epochs, so I am not even using the 1 million samples.. My biggest struggle has been to optimize for which validation samples I validate against, because it takes 4 hours to run through all 231.105 validation samples I have (which gives the most accurate result), but I've been decimating that as well and trying to find the best representation that represents the full val set, so I can validate each epoch in a minute instead of 4 hours.\n\nTo this: \"Would this involve initially training the model, identifying areas where it struggles, removing those samples, and then retraining?\" - yes exactly, except not only removing, but also weighting samples. Rapid experimentation and very many data iterations :-)\n\nTo this: \"Additionally, if you’re comfortable sharing, did you find that model selection had a significant impact, or was the iterative process of refining the dataset the more crucial factor?\" \nI used a pretty standard Multiscale Vision Transformer first time, and didn't try anything else, and manipulating the training data took score so far from 0.72 to 0.89.\n\nNow I am working on how to reduce training samples to under 100.000 frames, and understanding how I can get the model to separate even better where it still has challenges/hard time separating data in embedding space."
            },
            {
              "id": 3157795,
              "postDate": "2025-03-23T20:16:50.283Z",
              "content": "<p>I am sure a better model/experimenting with that, would help also! Hope sharing all this doesn't come and bite me… haha</p>",
              "rawMarkdown": "I am sure a better model/experimenting with that, would help also! Hope sharing all this doesn't come and bite me... haha"
            },
            {
              "id": 3158174,
              "postDate": "2025-03-24T09:34:33.860Z",
              "content": "<p>Hi Paul,<br>\nThanks for sharing these details!<br>\nIt's impressive how much of an impact strategic training data manipulation can have on model performance.</p>",
              "rawMarkdown": "Hi Paul,\nThanks for sharing these details!\nIt's impressive how much of an impact strategic training data manipulation can have on model performance."
            },
            {
              "id": 3158880,
              "postDate": "2025-03-25T02:47:16.400Z",
              "content": "<p>Yes - that's what I learned during my last 9 years working with AI/ML - it is in most cases disregarded, Andrew Ng tried 4 years ago to say how important it was to iterate on the data, but I feel it never catched on - and improving models can improve your score somewhat, but not much compared to improving your input data - what data you present to the model is everthing. More data is not better if it doesnt add more crucial information / helps the model separate, then it just adds more of the same.</p>",
              "rawMarkdown": "Yes - that's what I learned during my last 9 years working with AI/ML - it is in most cases disregarded, Andrew Ng tried 4 years ago to say how important it was to iterate on the data, but I feel it never catched on - and improving models can improve your score somewhat, but not much compared to improving your input data - what data you present to the model is everthing. More data is not better if it doesnt add more crucial information / helps the model separate, then it just adds more of the same."
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3126796,
      "author_name": "Daniel Moura",
      "author_url": "",
      "post_date": "2025-02-17T17:17:48.240000",
      "content": "<p>Alert times are only available for positive cases (collisions and near-collisions), and there is only one alert time per video (the closest to the event). Alert times are helpful to understand how ahead in time it is expected to anticipate that an event is about to happen (based on annotations from humans); e.g. if <code>event_time - alert_time = 0.7</code> it is not expected that a model can detect a potential accident 1 second before it happens. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 3126807,
          "author_name": "PaulEE",
          "author_url": "",
          "post_date": "2025-02-17T17:36:11.887000",
          "content": "<p>I understand, but for example video 01390 (and some other videos that has no events), 18 seconds in, it seems like an alert should have been there comparing to videos that has an alert - driver hits the brakes hard to not hit car that suddenly appeared in front, could be classified as near collision - Question is if there are room/possibility there could be \"missing\" labels?</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3126884,
              "author_name": "Daniel Moura",
              "author_url": "",
              "post_date": "2025-02-17T18:51:57.530000",
              "content": "<p>I agree that video 01390 could be classified as a near-collision. This is an annotation error. While we did our best to minimize annotation errors, they can still happen. Please, consider this as noise in the dataset. If you know about more cases, please post them here. Thank you 🙏  </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3126891,
              "author_name": "PaulEE",
              "author_url": "",
              "post_date": "2025-02-17T19:00:52.820000",
              "content": "<p>Ok, thanks!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3135050,
              "author_name": "PaulEE",
              "author_url": "",
              "post_date": "2025-02-27T02:09:33.650000",
              "content": "<p><a href=\"https://www.kaggle.com/danielcmoura\" target=\"_blank\">@danielcmoura</a> I am finding many others - for example train video 00573: 00573,10.124,8.293,1 - there is no alert that should be in 10 seconds as I can see? Or is it a car that comes in under the line that is blurred? (at 7 seconds). <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262906%2F36ae2f65dffb015f43f4c18e7bbfe0f2%2FScreenshot%202025-02-26%20190903.png?generation=1740622170111724&amp;alt=media\" alt=\"\"></p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3135052,
              "author_name": "PaulEE",
              "author_url": "",
              "post_date": "2025-02-27T02:11:28.980000",
              "content": "<p>I am looking at all the high confident but wrong cases that my model outputs, so many of these are easy to find. I'll just either delete or relabel those.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3141197,
              "author_name": "Daniel Moura",
              "author_url": "",
              "post_date": "2025-03-05T10:10:19.257000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/paulendresen76\" target=\"_blank\">@paulendresen76</a> . <code>00573</code> is an edge-case… there is a near-collision on that video, but it is obfuscated by the blurring we do for privacy protection. It would make sense to me to exclude this video from the training data since the blurring is making extremely difficult to detect the event. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3141455,
              "author_name": "PaulEE",
              "author_url": "",
              "post_date": "2025-03-05T15:42:28.130000",
              "content": "<p>Thanks for confirming! That is what I am doing in these cases :-)</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3163104,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-03-30T10:25:31.473000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3179755,
      "author_name": "Sarthak",
      "author_url": "",
      "post_date": "2025-04-15T16:42:10.790000",
      "content": "<p>Hi PaulEEE,</p>\n<p>Thanks for sharing your great work! Would you mind uploading the clean and preprocessed dataset you used? It would be really helpful for everyone trying to learn from your approach.</p>\n<p>Also, if possible, could you briefly explain the preprocessing steps you applied to the data? That context would make it easier for others to follow along and replicate your results.</p>\n<p>Appreciate your help!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3181184,
          "author_name": "PaulEE",
          "author_url": "",
          "post_date": "2025-04-17T14:52:30.633000",
          "content": "<p>Hi, I will publish everything after the competition is over as an article and on github -  I converted the videos into a 3LC table, where each frame represents a training sample, so I can use the tools we build. In the pytorch dataset that reads a 3LC table I load N frames leading up to the training sample frame, such that the actual training sample is N frames. Then I used 3LC to see how the model learned/struggled for individual frames(training sample) for my training runs (3LC capture metrics per training sample during training), and I really just worked with deleting ambigous individual ranges of frames(training samples) and weighting others that the model needed to focus on since it struggled to learn those scenarios. Only working with the data. It's not that easy to \"export\" back a dataset, as weight is 0 for all frames that I didnt want to train on, and varying weight for the other frames that needs more attention. I could of course export that as a value per frame per video. </p>",
          "votes": 0,
          "replies": [
            {
              "id": 3181187,
              "author_name": "PaulEE",
              "author_url": "",
              "post_date": "2025-04-17T14:56:45.783000",
              "content": "<p>It's worth mentioning that I didn't \"preprocess\", as I started training on all data, capture metrics, refine data, train again - iterating to understand how to delete and weight data, based on how the model reacted to the data it was fed. So I processed the data during the training.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3157107,
      "author_name": "SUHENG YANG",
      "author_url": "",
      "post_date": "2025-03-23T01:20:33.967000",
      "content": "<p>Hi,I noticed that your image display interface is very elegant and straightforward. It doesn’t seem to be directly from any open-source library I’ve come across. I was wondering, by any chance, could you let me know what software or tool you used to create it?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3157550,
          "author_name": "PaulEE",
          "author_url": "",
          "post_date": "2025-03-23T15:16:19.793000",
          "content": "<p>Hi it is free for non-commercial and Academia / Research - and also Kaggle competitions :-)<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262906%2F1c1d69aea4a2e50515c96efbdad0cbca%2FScreenshot%202025-03-23%20091436.png?generation=1742742925046909&amp;alt=media\" alt=\"\"></p>\n<p><a href=\"https://3lc.ai\" target=\"_blank\">https://3lc.ai</a> and https:/docs.3lc.ai</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3157558,
              "author_name": "PaulEE",
              "author_url": "",
              "post_date": "2025-03-23T15:27:07.370000",
              "content": "<p>Disclaimer, I am the founder of 3LC :-)</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3157850,
              "author_name": "SUHENG YANG",
              "author_url": "",
              "post_date": "2025-03-24T00:03:11.223000",
              "content": "<p>LOL, thank you so much for information =)</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3161880,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-03-28T14:40:59.197000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3161881,
              "author_name": "zhaoxuan lu",
              "author_url": "",
              "post_date": "2025-03-28T14:41:27.563000",
              "content": "<p>Awesome! Your excellent skills amaze me. I have become the 45th fan（youtube) of 3lc and your fan.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 3157559,
          "author_name": "PaulEE",
          "author_url": "",
          "post_date": "2025-03-23T15:28:35.477000",
          "content": "<p>I register my data (URL's to image path) and additional data per frame (Time to Event etc.) as 3LC Tables with the 3LC python pacakge, then I can plot anything against each other in the dashboard. Everything runs locally. (In the training script I open the 3LC tables as torch datasets and just add a map function to return the relevant data/columns for the model. What is nice I can also capture runs/metrics per sample during training such that I can investigate my data together with how the model performs on each frame. I do use more than one frame for my solution, but I just load previous images in the map function. Here for example I plot which frames are correctly predicted vs wrong in models embedding space.<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262906%2F0f96f281552e348082a869fe9fb85545%2FScreenshot%202025-03-23%20092249.jpg?generation=1742743586652254&amp;alt=media\" alt=\"\"></p>\n<p>I would be happy to help out - but everything shared has to be shared with everybody according to the rules, so please ask questions here :-) </p>",
          "votes": 0,
          "replies": [
            {
              "id": 3157636,
              "author_name": "PaulEE",
              "author_url": "",
              "post_date": "2025-03-23T16:28:56.833000",
              "content": "<p>So I see that in the area my model struggles I can investigate further and color frames with if the frames are between alert and event - and interestingly I see that the models embeddings in the area I have most error is 50/50 event/no event frames… and I can add loss as a 3rd axis to understand more. Not good that they are that close in embedding space, so what I've been doing is understand the training samples by selecting and looking at that data and in the UI and weight areas of samples I think the model should focus more on (in batch) and delete samples / ambigous cases where labelling might not be perfect. Then I use the new table revision for next training run, I am taking a 100% iterative data-centric approach to help the model understand the difference between when it should detect an event or not, emphasizing the hard data (for example when a car is close but drives straight(no event) vs close and its trajectory is bad), the model I am using is more or less off-the-shelf, simple and standard. If I had more data I would of course grow my dataset with frames in the area where it struggles the most in an automated fashion after understanding what it has a hard time to learn. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262906%2Fea4851c41220e9aec93699940e9c52eb%2FScreenshot%202025-03-23%20101126.png?generation=1742747240219431&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262906%2F217783d39fd0abfd89dfb4ebee102157%2FScreenshot%202025-03-23%20102234.png?generation=1742747250435877&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262906%2F6b4642cb068d9ca1aa6345d057995042%2FScreenshot%202025-03-23%20102246.png?generation=1742747259993111&amp;alt=media\" alt=\"\"></p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3157743,
              "author_name": "vale",
              "author_url": "",
              "post_date": "2025-03-23T19:00:59.710000",
              "content": "<p>Hi,<br>\nThank you for sharing these insights!</p>\n<p>If I may ask for some clarification—are you suggesting that removing ambiguous samples is the key factor in improving performance?</p>\n<p>Would this involve initially training the model, identifying areas where it struggles, removing those samples, and then retraining?</p>\n<p>Additionally, if you’re comfortable sharing, did you find that model selection had a significant impact, or was the iterative process of refining the dataset the more crucial factor?</p>\n<p>I truly appreciate your time and insights. Thanks again!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3157786,
              "author_name": "PaulEE",
              "author_url": "",
              "post_date": "2025-03-23T20:00:35.753000",
              "content": "<p>Hi, <a href=\"https://www.kaggle.com/valentinarizzello\" target=\"_blank\">@valentinarizzello</a> </p>\n<p>Not only ambigious samples, but that as well. They main impact is weighting each sample based on importance, so I added a weight column to the 3LC table so I can assign weights in the UI (on many samples at once), and then use torch WeightedRandomSampler that takes that column in, so it's quick to experiment, for example setting weights to 10 will show that sample 10 times as often. But also reducing the dataset size helped a lot -  I found that there are a lot of correctly labelled samples that does not contribute to the model getting better and is rather confusing the model, but I also removed ambigious samples. I have now 1172007 frames left of the 1701166 in the train data, so deleted around 32% of the frames. (However when a frame/sample is loaded, I load the existing N frames before that frame as well as input to my model, so some of these frames might be the deleted ones as training samples, but can still be used).</p>\n<p>Training is relatively quick, each epoch I train on 6000 samples (which loads N previous frames for each sample), and it usually converges after 10 epochs, so I am not even using the 1 million samples.. My biggest struggle has been to optimize for which validation samples I validate against, because it takes 4 hours to run through all 231.105 validation samples I have (which gives the most accurate result), but I've been decimating that as well and trying to find the best representation that represents the full val set, so I can validate each epoch in a minute instead of 4 hours.</p>\n<p>To this: \"Would this involve initially training the model, identifying areas where it struggles, removing those samples, and then retraining?\" - yes exactly, except not only removing, but also weighting samples. Rapid experimentation and very many data iterations :-)</p>\n<p>To this: \"Additionally, if you’re comfortable sharing, did you find that model selection had a significant impact, or was the iterative process of refining the dataset the more crucial factor?\" <br>\nI used a pretty standard Multiscale Vision Transformer first time, and didn't try anything else, and manipulating the training data took score so far from 0.72 to 0.89.</p>\n<p>Now I am working on how to reduce training samples to under 100.000 frames, and understanding how I can get the model to separate even better where it still has challenges/hard time separating data in embedding space.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3157795,
              "author_name": "PaulEE",
              "author_url": "",
              "post_date": "2025-03-23T20:16:50.283000",
              "content": "<p>I am sure a better model/experimenting with that, would help also! Hope sharing all this doesn't come and bite me… haha</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3158174,
              "author_name": "vale",
              "author_url": "",
              "post_date": "2025-03-24T09:34:33.860000",
              "content": "<p>Hi Paul,<br>\nThanks for sharing these details!<br>\nIt's impressive how much of an impact strategic training data manipulation can have on model performance.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3158880,
              "author_name": "PaulEE",
              "author_url": "",
              "post_date": "2025-03-25T02:47:16.400000",
              "content": "<p>Yes - that's what I learned during my last 9 years working with AI/ML - it is in most cases disregarded, Andrew Ng tried 4 years ago to say how important it was to iterate on the data, but I feel it never catched on - and improving models can improve your score somewhat, but not much compared to improving your input data - what data you present to the model is everthing. More data is not better if it doesnt add more crucial information / helps the model separate, then it just adds more of the same.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3126732": "I am debugging the False Positives I got on the training set, where my model thinks it should be an alert - in very many of those cases it seems to me an alert \"should\" have been thrown, even though it did not end up in a near-collision or collision afterwards - for example see video 01390, 18 seconds in. Car coming in from the left in front of the car, and driver brakes hard. Question is, are alerts that car/dashcam maybe threw removed if it didn't turn into a near collision or accident? (screenshot: ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2262906%2F036c04833de4c30209c57a84cfc31927%2FScreenshot%202025-02-17%20084055.png?generation=1739807090569402&alt=media)each vertical line is a separate video, and green is the highlighted frames that are False Positives).",
    "3126796": "Alert times are only available for positive cases (collisions and near-collisions), and there is only one alert time per video (the closest to the event). Alert times are helpful to understand how ahead in time it is expected to anticipate that an event is about to happen (based on annotations from humans); e.g. if `event_time - alert_time = 0.7` it is not expected that a model can detect a potential accident 1 second before it happens. ",
    "3179755": "Hi PaulEEE,\n\nThanks for sharing your great work! Would you mind uploading the clean and preprocessed dataset you used? It would be really helpful for everyone trying to learn from your approach.\n\nAlso, if possible, could you briefly explain the preprocessing steps you applied to the data? That context would make it easier for others to follow along and replicate your results.\n\nAppreciate your help!",
    "3157107": "Hi,I noticed that your image display interface is very elegant and straightforward. It doesn’t seem to be directly from any open-source library I’ve come across. I was wondering, by any chance, could you let me know what software or tool you used to create it?"
  }
}