{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":92399,"databundleVersionId":11038207,"sourceType":"competition"}],"dockerImageVersionId":30918,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# **Introduction**\n\nWelcome to this Exploratory Data Analysis (EDA) notebook for the Nexar Collision Prediction Challenge. \n\nThe goal of this challenge is to predict vehicle collisions (or near-misses) from dashcam videos before they occur. \n\nThis notebook covers:\n1. **CSV-based EDA** – exploring the training and test metadata (e.g., collision times, alert times, and targets).\n2. **Video-based EDA** – analyzing basic properties of the video files (e.g., duration, resolution, FPS, and approximate brightness).\n\nThrough this analysis, we aim to gain insights into data distribution, potential modeling challenges, and overall data quality.","metadata":{}},{"cell_type":"markdown","source":"# **Key Findings & Takeaways**\n\n1. **Balanced Dataset**  \n   - The training dataset is balanced with roughly equal numbers of positive (collision/near-miss) and negative (normal driving) cases, as mentioned in overview.\n\n2. **Event & Alert Times**  \n   - Collisions or near-misses often occur around **19–20 seconds** into the video.\n   - The alert times are typically close to, but slightly earlier than, the event times (**most often 1–2 seconds before**).\n\n3. **Time Gap**  \n   - For positive cases, the difference between `time_of_event` and `time_of_alert` peaks around **1 second**, indicating how quickly an accident might be predicted before it happens.\n\n4. **Video Durations**  \n   - Many training videos are around **40 seconds** long, matching the competition description. A few are shorter or longer, but **40 seconds is the main cluster**.\n\n5. **FPS & Resolution**  \n   - Most videos are at **~30 FPS** and have a **1280×720 resolution**. This is consistent and simplifies preprocessing since we have fewer variations to handle.\n\n6. **Brightness Variations**  \n   - **Average brightness ranges from about 20 to over 140**, indicating a variety of lighting conditions (e.g., daytime, dusk, nighttime, tunnels, etc.).  \n   - This could affect model performance if not handled properly.","metadata":{}},{"cell_type":"code","source":"import cv2\nimport matplotlib.pyplot as plt\nimport numpy as np\nimport os\nimport pandas as pd\nimport seaborn as sns\nimport sklearn.metrics\n\n# Set visualization style\nsns.set(style=\"whitegrid\")\nplt.rcParams[\"figure.figsize\"] = (10, 6)","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2025-02-23T15:23:57.126563Z","iopub.execute_input":"2025-02-23T15:23:57.126852Z","iopub.status.idle":"2025-02-23T15:23:58.879502Z","shell.execute_reply.started":"2025-02-23T15:23:57.126826Z","shell.execute_reply":"2025-02-23T15:23:58.878368Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_df = pd.read_csv('/kaggle/input/nexar-collision-prediction/train.csv')\nprint(\"Train CSV - First 5 Rows:\")\nprint(train_df.head())\nprint(\"\\nTrain CSV Info:\")\ntrain_df.info()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-02-23T15:24:11.810589Z","iopub.execute_input":"2025-02-23T15:24:11.810917Z","iopub.status.idle":"2025-02-23T15:24:11.862263Z","shell.execute_reply.started":"2025-02-23T15:24:11.810887Z","shell.execute_reply":"2025-02-23T15:24:11.861268Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(\"Train CSV - Summary Statistics:\")\nprint(train_df.describe())\nprint(\"\\nMissing Values in Train CSV:\")\nprint(train_df.isnull().sum())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-02-23T15:25:25.595893Z","iopub.execute_input":"2025-02-23T15:25:25.596259Z","iopub.status.idle":"2025-02-23T15:25:25.618809Z","shell.execute_reply.started":"2025-02-23T15:25:25.596226Z","shell.execute_reply":"2025-02-23T15:25:25.617869Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"* Some rows have NaN for time_of_event and time_of_alert (these correspond to negative/normal driving cases).\n* The id column gives a unique identifier for each video.\n* There are 1500 rows total.\n* Exactly half (750) have time_of_event and time_of_alert, corresponding to the positive (collision/near-miss) cases.\n* target is 0 or 1, indicating negative (normal driving) vs. positive (collision/near-miss) cases.\n* As expected, time_of_event and time_of_alert are missing in 750 rows each (negative cases).\n* The average collision event time is ~19.1 seconds, and the average alert time is ~17.5 seconds.* ","metadata":{}},{"cell_type":"code","source":"# 1. Distribution of Target Classes\nplt.figure()\nsns.countplot(x='target', data=train_df, palette=\"Set2\")\nplt.title(\"Target Class Distribution (0: Normal, 1: Collision/Near-miss)\")\nplt.xlabel(\"Target\")\nplt.ylabel(\"Count\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-02-23T15:25:35.776197Z","iopub.execute_input":"2025-02-23T15:25:35.776654Z","iopub.status.idle":"2025-02-23T15:25:36.051654Z","shell.execute_reply.started":"2025-02-23T15:25:35.776610Z","shell.execute_reply":"2025-02-23T15:25:36.050641Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 2. Distribution of Time of Event (for positive cases)\nif 'time_of_event' in train_df.columns:\n    plt.figure()\n    positive_events = train_df[train_df['target'] == 1]\n    sns.histplot(positive_events['time_of_event'].dropna(), bins=30, kde=True, color='salmon')\n    plt.title(\"Distribution of Time of Event (Positive Cases)\")\n    plt.xlabel(\"Time of Event (seconds)\")\n    plt.ylabel(\"Frequency\")\n    plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-02-23T15:25:42.981008Z","iopub.execute_input":"2025-02-23T15:25:42.981336Z","iopub.status.idle":"2025-02-23T15:25:43.320204Z","shell.execute_reply.started":"2025-02-23T15:25:42.981310Z","shell.execute_reply":"2025-02-23T15:25:43.319233Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 3. Distribution of Time of Alert (for positive cases)\nif 'time_of_alert' in train_df.columns:\n    plt.figure()\n    positive_alerts = train_df[train_df['target'] == 1]\n    sns.histplot(positive_alerts['time_of_alert'].dropna(), bins=30, kde=True, color='skyblue')\n    plt.title(\"Distribution of Time of Alert (Positive Cases)\")\n    plt.xlabel(\"Time of Alert (seconds)\")\n    plt.ylabel(\"Frequency\")\n    plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-02-23T15:25:50.990951Z","iopub.execute_input":"2025-02-23T15:25:50.991332Z","iopub.status.idle":"2025-02-23T15:25:51.412796Z","shell.execute_reply.started":"2025-02-23T15:25:50.991302Z","shell.execute_reply":"2025-02-23T15:25:51.411736Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 4. Distribution of Time Gap between Event and Alert (for positive cases)\nif 'time_of_event' in train_df.columns and 'time_of_alert' in train_df.columns:\n    train_df['time_gap'] = train_df['time_of_event'] - train_df['time_of_alert']\n    plt.figure()\n    sns.histplot(train_df.loc[train_df['target'] == 1, 'time_gap'].dropna(), bins=30, kde=True, color='green')\n    plt.title(\"Distribution of Time Gap (Event - Alert) for Positive Cases\")\n    plt.xlabel(\"Time Gap (seconds)\")\n    plt.ylabel(\"Frequency\")\n    plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-02-23T15:25:58.194156Z","iopub.execute_input":"2025-02-23T15:25:58.194535Z","iopub.status.idle":"2025-02-23T15:25:58.531595Z","shell.execute_reply.started":"2025-02-23T15:25:58.194496Z","shell.execute_reply":"2025-02-23T15:25:58.530611Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"test_df = pd.read_csv('/kaggle/input/nexar-collision-prediction/test.csv')\nprint(\"Test CSV - First 5 Rows:\")\nprint(test_df.head())\nprint(\"\\nTest CSV Shape:\", test_df.shape)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-02-23T15:26:36.149623Z","iopub.execute_input":"2025-02-23T15:26:36.149963Z","iopub.status.idle":"2025-02-23T15:26:36.162066Z","shell.execute_reply.started":"2025-02-23T15:26:36.149934Z","shell.execute_reply":"2025-02-23T15:26:36.161015Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"* Only video IDs are provided in the test set, as the event time remains private.\n* There are 1344 test videos in total.","metadata":{}},{"cell_type":"code","source":"def extract_video_info(video_path):\n    \"\"\"\n    Extract basic properties from a video file.\n    Returns a dictionary with FPS, frame count, resolution, and duration.\n    \"\"\"\n    cap = cv2.VideoCapture(video_path)\n    if not cap.isOpened():\n        print(f\"Failed to open video: {video_path}\")\n        return None\n    fps = cap.get(cv2.CAP_PROP_FPS)\n    frame_count = int(cap.get(cv2.CAP_PROP_FRAME_COUNT))\n    width = int(cap.get(cv2.CAP_PROP_FRAME_WIDTH))\n    height = int(cap.get(cv2.CAP_PROP_FRAME_HEIGHT))\n    duration = frame_count / fps if fps != 0 else 0\n    cap.release()\n    return {\n        \"video_path\": video_path,\n        \"fps\": fps,\n        \"frame_count\": frame_count,\n        \"width\": width,\n        \"height\": height,\n        \"duration\": duration\n    }\n\ndef compute_average_brightness(video_path, sample_frames=10):\n    \"\"\"\n    Compute an approximate average brightness by sampling a few frames.\n    Brightness is calculated as the mean grayscale pixel value.\n    \"\"\"\n    cap = cv2.VideoCapture(video_path)\n    if not cap.isOpened():\n        print(f\"Failed to open video: {video_path}\")\n        return None\n    brightness_vals = []\n    frame_count = int(cap.get(cv2.CAP_PROP_FRAME_COUNT))\n    if frame_count == 0:\n        cap.release()\n        return None\n    frames_to_sample = min(sample_frames, frame_count)\n    interval = max(1, frame_count // frames_to_sample)\n    for i in range(frames_to_sample):\n        cap.set(cv2.CAP_PROP_POS_FRAMES, i * interval)\n        ret, frame = cap.read()\n        if ret:\n            gray = cv2.cvtColor(frame, cv2.COLOR_BGR2GRAY)\n            brightness_vals.append(np.mean(gray))\n    cap.release()\n    return np.mean(brightness_vals) if brightness_vals else None","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-02-23T15:26:57.171363Z","iopub.execute_input":"2025-02-23T15:26:57.171730Z","iopub.status.idle":"2025-02-23T15:26:57.181899Z","shell.execute_reply.started":"2025-02-23T15:26:57.171698Z","shell.execute_reply":"2025-02-23T15:26:57.180899Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"video_folder = '/kaggle/input/nexar-collision-prediction/train' \n\n# List all mp4 files in the specified video folder\nvideo_files = [os.path.join(video_folder, f) for f in os.listdir(video_folder) if f.endswith('.mp4')]\n\n# For quick processing, sample only the first 50 videos (adjust as needed)\nsample_video_files = video_files[:50]\n\nvideo_data = []\nfor video_file in sample_video_files:\n    info = extract_video_info(video_file)\n    if info:\n        info['avg_brightness'] = compute_average_brightness(video_file, sample_frames=5)\n        video_data.append(info)\n\nvideo_df = pd.DataFrame(video_data)\nprint(\"Video EDA Data:\")\nvideo_df.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-02-23T15:50:00.648179Z","iopub.execute_input":"2025-02-23T15:50:00.648571Z","iopub.status.idle":"2025-02-23T15:51:41.665995Z","shell.execute_reply.started":"2025-02-23T15:50:00.648539Z","shell.execute_reply":"2025-02-23T15:51:41.665162Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"video_df.describe()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-02-23T15:52:04.137240Z","iopub.execute_input":"2025-02-23T15:52:04.137622Z","iopub.status.idle":"2025-02-23T15:52:04.168137Z","shell.execute_reply.started":"2025-02-23T15:52:04.137591Z","shell.execute_reply":"2025-02-23T15:52:04.167219Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 1. Video Duration Distribution\nplt.figure()\nsns.histplot(video_df['duration'], bins=20, kde=True, color='blue')\nplt.title(\"Distribution of Video Durations\")\nplt.xlabel(\"Duration (seconds)\")\nplt.ylabel(\"Count\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-02-23T15:29:19.282205Z","iopub.execute_input":"2025-02-23T15:29:19.282573Z","iopub.status.idle":"2025-02-23T15:29:19.580890Z","shell.execute_reply.started":"2025-02-23T15:29:19.282538Z","shell.execute_reply":"2025-02-23T15:29:19.579890Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 2. Video FPS Distribution\nplt.figure()\nsns.histplot(video_df['fps'], bins=10, kde=True, color='orange')\nplt.title(\"Distribution of Video FPS\")\nplt.xlabel(\"FPS\")\nplt.ylabel(\"Count\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-02-23T15:29:19.582805Z","iopub.execute_input":"2025-02-23T15:29:19.583100Z","iopub.status.idle":"2025-02-23T15:29:19.891502Z","shell.execute_reply.started":"2025-02-23T15:29:19.583073Z","shell.execute_reply":"2025-02-23T15:29:19.890607Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 3. Average Brightness Distribution\nplt.figure()\nsns.histplot(video_df['avg_brightness'].dropna(), bins=20, kde=True, color='green')\nplt.title(\"Distribution of Average Brightness\")\nplt.xlabel(\"Average Brightness\")\nplt.ylabel(\"Count\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-02-23T15:29:19.892498Z","iopub.execute_input":"2025-02-23T15:29:19.892767Z","iopub.status.idle":"2025-02-23T15:29:20.228538Z","shell.execute_reply.started":"2025-02-23T15:29:19.892745Z","shell.execute_reply":"2025-02-23T15:29:20.227523Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 4. Video Resolution: Width vs. Height\nplt.figure()\nsns.scatterplot(x='width', y='height', data=video_df, color='purple')\nplt.title(\"Video Resolution: Width vs. Height\")\nplt.xlabel(\"Width (pixels)\")\nplt.ylabel(\"Height (pixels)\")\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-02-23T15:29:20.229568Z","iopub.execute_input":"2025-02-23T15:29:20.229958Z","iopub.status.idle":"2025-02-23T15:29:20.526496Z","shell.execute_reply.started":"2025-02-23T15:29:20.229923Z","shell.execute_reply":"2025-02-23T15:29:20.525584Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"1. **Distribution of Video Durations**  \n   - Most videos last **~40 seconds**, with some variations (~20 or ~60 seconds).\n   \n2. **Distribution of FPS**  \n   - Mostly clustered around **30 FPS**, but a few outliers exist (e.g., 25.3 FPS, 31.0 FPS).\n\n3. **Average Brightness**  \n   - Ranges from **~20 to ~145**, reflecting **varied lighting conditions**.\n\n4. **Video Resolution**  \n   - Almost all videos are **1280×720**, simplifying preprocessing.","metadata":{}},{"cell_type":"markdown","source":"# **Conclusion**\n\nFrom this EDA, we can see:\n\n- The **training data is balanced**, with **750 positive and 750 negative cases**.\n- **Collisions typically occur around ~19–20 seconds**, with alerts happening **1–2 seconds earlier**.\n- **Most videos are 40 seconds long, 30 FPS, and 1280×720 resolution**, making preprocessing easier.\n- **Brightness variations** suggest the need for **light adaptation techniques (e.g., data augmentation, contrast normalization)**.\n- The **time gap between alert and event peaks around 1 second**, emphasizing the challenge of **early accident prediction**.\n\n### **Next Steps**\n- **Feature Engineering:** Handle **lighting conditions** and **time-to-accident constraints**.\n- **Model Training:** Use **consistent FPS/resolution**, but account for **varied brightness**.\n- **Temporal Modeling:** Optimize for **early warnings (1–2 sec before event)**.\n\nThis foundational understanding will guide **model development and preprocessing strategies** for accurate accident prediction.\n","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}