{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":92399,"databundleVersionId":11038207,"sourceType":"competition"}],"dockerImageVersionId":30918,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# 🚗💥 Nexar Dashcam Crash Prediction - Data Exploration\n\nThis notebook explores the Nexar Collision Prediction competition dataset. The goal is to predict whether a collision is about to occur based on dashcam video footage.\n\n## Dataset Overview\n\nThe dataset consists of road-facing videos captured by Nexar dashcams:\n- **Positive cases**: Videos where a collision occurs or is imminent (near-miss)\n- **Negative cases**: Videos showing normal driving with no collision or near-miss\n- Training set contains 1,500 videos (balanced between positive and negative cases)\n- Test set contains 1,344 videos\n\nFor positive cases in the training set, two timestamps are provided:\n- **Event time**: When the collision or near-miss occurs\n- **Alert time**: The earliest moment when action should be taken to prevent the accident","metadata":{}},{"cell_type":"markdown","source":"## Initial Setup\n\nFirst, we'll import the necessary libraries and explore what files are available in the input directory. We'll limit the output to avoid displaying too many files.","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\nimport os\n\n# Limit the number of files to display (e.g., 10)\nmax_files = 10\nfile_count = 0\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n        file_count += 1\n        if file_count >= max_files:  # Stop when reaching maximum count\n            print(\"... more files omitted ...\")\n            break\n    if file_count >= max_files:\n        break\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2025-04-04T05:07:47.601213Z","iopub.execute_input":"2025-04-04T05:07:47.601658Z","iopub.status.idle":"2025-04-04T05:07:48.062560Z","shell.execute_reply.started":"2025-04-04T05:07:47.601623Z","shell.execute_reply":"2025-04-04T05:07:48.061357Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Import libraries\nimport cv2\nfrom tqdm import tqdm\n\n# Suppress NaN warning messages\nimport warnings\nwarnings.filterwarnings('ignore', category=RuntimeWarning)\n\n# Exploratory Data Analysis (EDA)\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom sklearn.preprocessing import StandardScaler\nimport numpy as np","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-04T05:07:48.063719Z","iopub.execute_input":"2025-04-04T05:07:48.064158Z","iopub.status.idle":"2025-04-04T05:07:48.796099Z","shell.execute_reply.started":"2025-04-04T05:07:48.064124Z","shell.execute_reply":"2025-04-04T05:07:48.795025Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Data Loading and Exploration\n\nWe'll load the CSV files and check the basic properties of videos in the dataset. This includes:\n1. Examining the data structure and distribution\n2. Checking video properties (resolution, duration, frame rate)\n3. Exploring sample positive and negative examples","metadata":{}},{"cell_type":"code","source":"# 1. Load CSV files\ndef load_data(train_csv_path='train.csv', test_csv_path='test.csv'):\n    \"\"\"\n    Load training and testing data from CSV files.\n    \"\"\"\n    # Load CSV files\n    train_df = pd.read_csv(train_csv_path)\n    test_df = pd.read_csv(test_csv_path)\n    \n    # Format IDs as 5-digit numbers\n    train_df['id'] = train_df['id'].apply(lambda x: f\"{int(float(x)):05d}\")\n    test_df['id'] = test_df['id'].apply(lambda x: f\"{int(float(x)):05d}\")\n    \n    print(\"Training data shape:\", train_df.shape)\n    print(\"Testing data shape:\", test_df.shape)\n    \n    return train_df, test_df","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-04T05:07:48.798028Z","iopub.execute_input":"2025-04-04T05:07:48.798505Z","iopub.status.idle":"2025-04-04T05:07:48.806109Z","shell.execute_reply.started":"2025-04-04T05:07:48.798472Z","shell.execute_reply":"2025-04-04T05:07:48.804915Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Video Processing\n\nFor video analysis, we'll:\n1. Extract basic video information\n2. Sample frames from videos to understand their content\n3. Compare positive and negative examples","metadata":{}},{"cell_type":"code","source":"# 2. Get video file information\ndef check_video_info(video_path):\n    \"\"\"\n    Check basic information about a video file.\n    \"\"\"\n    cap = cv2.VideoCapture(video_path)\n    \n    if not cap.isOpened():\n        print(f\"Error: Could not open video {video_path}\")\n        return None\n    \n    # Get video properties\n    width = int(cap.get(cv2.CAP_PROP_FRAME_WIDTH))\n    height = int(cap.get(cv2.CAP_PROP_FRAME_HEIGHT))\n    fps = cap.get(cv2.CAP_PROP_FPS)\n    frame_count = int(cap.get(cv2.CAP_PROP_FRAME_COUNT))\n    duration = frame_count / fps if fps > 0 else 0\n    \n    cap.release()\n    \n    return {\n        'width': width,\n        'height': height,\n        'fps': fps,\n        'frame_count': frame_count,\n        'duration_seconds': duration\n    }","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-04T05:07:48.807906Z","iopub.execute_input":"2025-04-04T05:07:48.808340Z","iopub.status.idle":"2025-04-04T05:07:48.834229Z","shell.execute_reply.started":"2025-04-04T05:07:48.808296Z","shell.execute_reply":"2025-04-04T05:07:48.832981Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 3. Extract sample frames from videos\ndef extract_sample_frames(video_path, num_frames=5):\n    \"\"\"\n    Extract sample frames from a video for visualization.\n    \"\"\"\n    cap = cv2.VideoCapture(video_path)\n    \n    if not cap.isOpened():\n        print(f\"Error: Could not open video {video_path}\")\n        return None\n    \n    # Get total frame count\n    total_frames = int(cap.get(cv2.CAP_PROP_FRAME_COUNT))\n    \n    # Calculate evenly spaced frame indices\n    indices = np.linspace(0, total_frames - 1, num_frames, dtype=int)\n    \n    frames = []\n    for idx in indices:\n        cap.set(cv2.CAP_PROP_POS_FRAMES, idx)\n        ret, frame = cap.read()\n        if ret:\n            # Convert BGR to RGB\n            frame = cv2.cvtColor(frame, cv2.COLOR_BGR2RGB)\n            frames.append(frame)\n    \n    cap.release()\n    return frames","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-04T05:07:48.835350Z","iopub.execute_input":"2025-04-04T05:07:48.835626Z","iopub.status.idle":"2025-04-04T05:07:48.854363Z","shell.execute_reply.started":"2025-04-04T05:07:48.835603Z","shell.execute_reply":"2025-04-04T05:07:48.853300Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 4. Explore sample data (both positive and negative examples)\ndef explore_sample_data(train_df, video_dir='train', num_samples=2):\n    \"\"\"\n    Explore sample positive and negative examples.\n    \"\"\"\n    # Select positive samples\n    positive_samples = train_df[train_df['target'] == 1].sample(num_samples)\n    \n    # Select negative samples\n    negative_samples = train_df[train_df['target'] == 0].sample(num_samples)\n    \n    samples = pd.concat([positive_samples, negative_samples])\n    \n    results = []\n    for _, row in samples.iterrows():\n        video_path = os.path.join(video_dir, f\"{row['id']}.mp4\")\n        video_info = check_video_info(video_path)\n        \n        if video_info:\n            result = {\n                'id': row['id'],\n                'target': row['target'],\n                'time_of_event': row.get('time_of_event', 'N/A'),\n                'time_of_alert': row.get('time_of_alert', 'N/A'),\n                'video_info': video_info\n            }\n            results.append(result)\n    \n    return results","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-04T05:07:48.855174Z","iopub.execute_input":"2025-04-04T05:07:48.855538Z","iopub.status.idle":"2025-04-04T05:07:48.872825Z","shell.execute_reply.started":"2025-04-04T05:07:48.855501Z","shell.execute_reply":"2025-04-04T05:07:48.871896Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Main function\ndef explore_data(train_csv='train.csv', test_csv='test.csv', train_dir='train'):\n    \"\"\"\n    Main function to explore the dataset.\n    \"\"\"\n    # 1. Load CSV files\n    train_df, test_df = load_data(train_csv, test_csv)\n    \n    # 2. Print basic data information\n    print(\"\\nTraining data columns:\", train_df.columns.tolist())\n    print(\"\\nSample of training data:\")\n    print(train_df.head())\n    \n    # 3. Check label distribution\n    if 'target' in train_df.columns:\n        print(\"\\nTarget distribution:\")\n        print(train_df['target'].value_counts())\n    \n    # 4. Explore sample data\n    print(\"\\nExploring sample videos...\")\n    sample_results = explore_sample_data(train_df, train_dir)\n    \n    for i, result in enumerate(sample_results):\n        print(f\"\\nSample {i+1}:\")\n        print(f\"  ID: {result['id']}\")\n        print(f\"  Target: {result['target']} ({'Positive/Accident' if result['target'] == 1 else 'Negative/Normal'})\")\n        print(f\"  Event time: {result['time_of_event']}\")\n        print(f\"  Alert time: {result['time_of_alert']}\")\n        print(f\"  Video info: {result['video_info']}\")\n    \n    return train_df, test_df","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-04T05:07:48.874142Z","iopub.execute_input":"2025-04-04T05:07:48.874434Z","iopub.status.idle":"2025-04-04T05:07:48.891270Z","shell.execute_reply.started":"2025-04-04T05:07:48.874409Z","shell.execute_reply":"2025-04-04T05:07:48.890245Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Run the code\nif __name__ == \"__main__\":\n    # Find CSV file paths\n    train_csv_path = None\n    test_csv_path = None\n    \n    for dirname, _, filenames in os.walk('/kaggle/input'):\n        for filename in filenames:\n            if filename == 'train.csv':\n                train_csv_path = os.path.join(dirname, filename)\n            elif filename == 'test.csv':\n                test_csv_path = os.path.join(dirname, filename)\n    \n    # Find train video directory path\n    train_video_dir = None\n    for dirname, _, filenames in os.walk('/kaggle/input'):\n        if os.path.basename(dirname) == 'train' and any(f.endswith('.mp4') for f in filenames):\n            train_video_dir = dirname\n            break\n    \n    # Run the exploration if paths are found\n    if train_csv_path and test_csv_path and train_video_dir:\n        print(f\"CSV file paths:\\n- train.csv: {train_csv_path}\\n- test.csv: {test_csv_path}\")\n        print(f\"Video directory: {train_video_dir}\")\n        \n        train_df, test_df = explore_data(train_csv_path, test_csv_path, train_video_dir)\n    else:\n        print(\"Could not find required files or directories.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-04T05:07:48.892462Z","iopub.execute_input":"2025-04-04T05:07:48.892858Z","iopub.status.idle":"2025-04-04T05:07:49.528613Z","shell.execute_reply.started":"2025-04-04T05:07:48.892819Z","shell.execute_reply":"2025-04-04T05:07:49.527278Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Exploratory Data Analysis (EDA)\n\nAfter loading and inspecting the basic structure of our data, we'll perform some exploratory analysis to better understand the dataset's characteristics.\n\n### What we'll explore:\n\n1. **Time Feature Analysis**\n   - For positive cases (accidents), analyze the relationship between event time and alert time\n   - Calculate the reaction time window (difference between event and alert)\n   - Visualize the distribution of event timings within videos\n\n2. **Video Properties Analysis**\n   - Compare video duration, resolution, and frame rate between positive and negative cases\n   - Check for any technical differences that might affect model performance\n\n3. **Motion Pattern Analysis**\n   - Calculate optical flow to measure movement in videos\n   - Compare motion patterns between accident and normal driving videos\n   - Identify potential movement signatures that precede accidents","metadata":{}},{"cell_type":"code","source":"def perform_eda(train_df, train_video_dir):\n    \"\"\"\n    Perform exploratory data analysis on the dataset.\n    \n    Args:\n        train_df: Training dataframe\n        train_video_dir: Directory containing training videos\n    \"\"\"\n    print(\"\\n## Exploratory Data Analysis ##\")\n    \n    # 1. Analyze time-related features for accident cases\n    analyze_time_features(train_df)\n    \n    # 2. Analyze video properties\n    analyze_video_properties(train_df, train_video_dir)\n    \n    # 3. Analyze optical flow across samples (movement detection)\n    analyze_sample_optical_flow(train_df, train_video_dir)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-04T05:07:49.531346Z","iopub.execute_input":"2025-04-04T05:07:49.531676Z","iopub.status.idle":"2025-04-04T05:07:49.536417Z","shell.execute_reply.started":"2025-04-04T05:07:49.531645Z","shell.execute_reply":"2025-04-04T05:07:49.535274Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def analyze_time_features(train_df):\n    \"\"\"\n    Analyze time-related features for accident cases\n    \"\"\"\n    # Focus on positive examples (accident cases)\n    accident_df = train_df[train_df['target'] == 1].copy()\n    \n    # Calculate reaction time (difference between event and alert)\n    accident_df['reaction_time'] = accident_df['time_of_event'] - accident_df['time_of_alert']\n    \n    print(f\"\\n1. Time Feature Analysis (Accident Cases Only - {len(accident_df)} samples)\")\n    print(\"\\nReaction time statistics (seconds):\")\n    print(accident_df['reaction_time'].describe())\n    \n    # Create visualization for reaction time\n    plt.figure(figsize=(10, 6))\n    sns.histplot(accident_df['reaction_time'], bins=20, kde=True)\n    plt.title('Distribution of Reaction Time (Event Time - Alert Time)')\n    plt.xlabel('Reaction Time (seconds)')\n    plt.ylabel('Count')\n    plt.axvline(accident_df['reaction_time'].mean(), color='red', linestyle='--', \n               label=f'Mean: {accident_df[\"reaction_time\"].mean():.2f}s')\n    plt.legend()\n    plt.grid(True, alpha=0.3)\n    plt.show()\n    \n    # Analyze event timing distribution\n    plt.figure(figsize=(10, 6))\n    sns.histplot(accident_df['time_of_event'], bins=20, kde=True)\n    plt.title('Distribution of Event Times in Videos')\n    plt.xlabel('Event Time (seconds)')\n    plt.ylabel('Count')\n    plt.grid(True, alpha=0.3)\n    plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-04T05:07:49.537743Z","iopub.execute_input":"2025-04-04T05:07:49.538014Z","iopub.status.idle":"2025-04-04T05:07:49.558590Z","shell.execute_reply.started":"2025-04-04T05:07:49.537992Z","shell.execute_reply":"2025-04-04T05:07:49.557528Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def analyze_video_properties(train_df, train_video_dir, sample_size=100):\n    \"\"\"\n    Analyze properties of videos in the dataset\n    \n    Args:\n        train_df: Training dataframe\n        train_video_dir: Directory containing training videos\n        sample_size: Number of videos to sample for analysis\n    \"\"\"\n    print(f\"\\n2. Video Properties Analysis (Sample of {sample_size} videos)\")\n    \n    # Sample videos for analysis\n    sampled_df = train_df.sample(min(sample_size, len(train_df)), random_state=42)\n    \n    # Collect video properties\n    video_properties = []\n    for _, row in tqdm(sampled_df.iterrows(), total=len(sampled_df), desc=\"Analyzing videos\"):\n        video_path = os.path.join(train_video_dir, f\"{row['id']}.mp4\")\n        video_info = check_video_info(video_path)\n        \n        if video_info:\n            video_info['target'] = row['target']\n            video_properties.append(video_info)\n    \n    # Convert to DataFrame\n    video_df = pd.DataFrame(video_properties)\n    \n    if len(video_df) > 0:\n        # Display summary statistics\n        print(\"\\nVideo property statistics:\")\n        print(video_df.describe())\n        \n        # Compare duration between positive and negative cases\n        plt.figure(figsize=(12, 6))\n        sns.boxplot(x='target', y='duration_seconds', data=video_df)\n        plt.title('Video Duration by Class')\n        plt.xlabel('Class (0=Normal, 1=Accident)')\n        plt.ylabel('Duration (seconds)')\n        plt.grid(True, alpha=0.3)\n        plt.show()\n        \n        # Check resolution distribution\n        plt.figure(figsize=(10, 6))\n        video_df['resolution'] = video_df['width'].astype(str) + 'x' + video_df['height'].astype(str)\n        resolution_counts = video_df['resolution'].value_counts()\n        resolution_counts.plot(kind='bar')\n        plt.title('Video Resolution Distribution')\n        plt.xlabel('Resolution')\n        plt.ylabel('Count')\n        plt.xticks(rotation=45)\n        plt.grid(True, alpha=0.3)\n        plt.show()\n    else:\n        print(\"No valid video properties found in the sample.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-04T05:07:49.559982Z","iopub.execute_input":"2025-04-04T05:07:49.560311Z","iopub.status.idle":"2025-04-04T05:07:49.580420Z","shell.execute_reply.started":"2025-04-04T05:07:49.560284Z","shell.execute_reply":"2025-04-04T05:07:49.579343Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def analyze_sample_optical_flow(train_df, train_video_dir, num_samples=5):\n    \"\"\"\n    Analyze optical flow in sample videos to detect motion patterns\n    \n    Args:\n        train_df: Training dataframe\n        train_video_dir: Directory containing training videos\n        num_samples: Number of samples to analyze from each class\n    \"\"\"\n    print(f\"\\n3. Optical Flow Analysis (Sample of {num_samples*2} videos)\")\n    \n    # Select samples from each class\n    positive_samples = train_df[train_df['target'] == 1].sample(num_samples)\n    negative_samples = train_df[train_df['target'] == 0].sample(num_samples)\n    samples = pd.concat([positive_samples, negative_samples])\n    \n    # Function to calculate average optical flow magnitude\n    def get_optical_flow(video_path, num_frames=10):\n        cap = cv2.VideoCapture(video_path)\n        if not cap.isOpened():\n            return None\n            \n        # Get total frames and calculate frame indices\n        total_frames = int(cap.get(cv2.CAP_PROP_FRAME_COUNT))\n        indices = np.linspace(0, total_frames - num_frames - 1, num_frames, dtype=int)\n        \n        flows = []\n        prev_frame = None\n        \n        for idx in indices:\n            cap.set(cv2.CAP_PROP_POS_FRAMES, idx)\n            ret, frame = cap.read()\n            \n            if not ret:\n                continue\n                \n            gray = cv2.cvtColor(frame, cv2.COLOR_BGR2GRAY)\n            \n            if prev_frame is not None:\n                # Calculate optical flow\n                flow = cv2.calcOpticalFlowFarneback(prev_frame, gray, None, 0.5, 3, 15, 3, 5, 1.2, 0)\n                # Calculate magnitude\n                magnitude = np.sqrt(flow[..., 0]**2 + flow[..., 1]**2)\n                flows.append(np.mean(magnitude))\n            \n            prev_frame = gray\n            \n        cap.release()\n        return flows if flows else None\n    \n    # Calculate optical flow for samples\n    results = []\n    for _, row in tqdm(samples.iterrows(), total=len(samples), desc=\"Calculating optical flow\"):\n        video_path = os.path.join(train_video_dir, f\"{row['id']}.mp4\")\n        flows = get_optical_flow(video_path)\n        \n        if flows:\n            results.append({\n                'id': row['id'],\n                'target': row['target'],\n                'mean_flow': np.mean(flows),\n                'max_flow': np.max(flows),\n                'flow_values': flows\n            })\n    \n    if results:\n        # Convert to DataFrame\n        flow_df = pd.DataFrame([\n            {'id': r['id'], 'target': r['target'], 'mean_flow': r['mean_flow'], 'max_flow': r['max_flow']} \n            for r in results\n        ])\n        \n        # Display results\n        print(\"\\nOptical flow statistics by class:\")\n        print(flow_df.groupby('target')[['mean_flow', 'max_flow']].describe())\n        \n        # Visualize mean flow by class\n        plt.figure(figsize=(10, 6))\n        sns.boxplot(x='target', y='mean_flow', data=flow_df)\n        plt.title('Mean Optical Flow Magnitude by Class')\n        plt.xlabel('Class (0=Normal, 1=Accident)')\n        plt.ylabel('Mean Flow Magnitude')\n        plt.grid(True, alpha=0.3)\n        plt.show()\n        \n        # Plot flow over time for a few samples\n        plt.figure(figsize=(12, 8))\n        for i, result in enumerate(results[:4]):  # Plot first 4 samples\n            plt.subplot(2, 2, i+1)\n            plt.plot(result['flow_values'])\n            plt.title(f\"Sample {result['id']} (Class {result['target']})\")\n            plt.xlabel('Frame Index')\n            plt.ylabel('Flow Magnitude')\n            plt.grid(True, alpha=0.3)\n        plt.tight_layout()\n        plt.show()\n    else:\n        print(\"No valid optical flow results found in the samples.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-04T05:07:49.581420Z","iopub.execute_input":"2025-04-04T05:07:49.581782Z","iopub.status.idle":"2025-04-04T05:07:49.604590Z","shell.execute_reply.started":"2025-04-04T05:07:49.581747Z","shell.execute_reply":"2025-04-04T05:07:49.603472Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Feature extraction for model building\ndef extract_basic_features(train_df, train_video_dir, sample_size=None):\n    \"\"\"\n    Extract basic features from videos for model building\n    \n    Args:\n        train_df: Training dataframe\n        train_video_dir: Directory containing training videos\n        sample_size: Number of videos to sample (None for all)\n        \n    Returns:\n        X: Feature matrix\n        y: Target labels\n        feature_names: Names of extracted features\n    \"\"\"\n    # Sample data if requested\n    if sample_size is not None:\n        df = train_df.sample(min(sample_size, len(train_df)), random_state=42)\n    else:\n        df = train_df\n    \n    features = []\n    labels = []\n    \n    print(\"Extracting features from videos...\")\n    for _, row in tqdm(df.iterrows(), total=len(df)):\n        video_path = os.path.join(train_video_dir, f\"{row['id']}.mp4\")\n        \n        # Check if video can be opened\n        cap = cv2.VideoCapture(video_path)\n        if not cap.isOpened():\n            continue\n        \n        # Basic video properties\n        width = int(cap.get(cv2.CAP_PROP_FRAME_WIDTH))\n        height = int(cap.get(cv2.CAP_PROP_FRAME_HEIGHT))\n        fps = cap.get(cv2.CAP_PROP_FPS)\n        frame_count = int(cap.get(cv2.CAP_PROP_FRAME_COUNT))\n        duration = frame_count / fps if fps > 0 else 0\n        \n        # Sample frames for analysis\n        num_sample_frames = min(10, frame_count)\n        indices = np.linspace(0, frame_count - 1, num_sample_frames, dtype=int)\n        \n        # Features to extract\n        avg_brightness = []\n        avg_motion = []\n        prev_gray = None\n        \n        for idx in indices:\n            cap.set(cv2.CAP_PROP_POS_FRAMES, idx)\n            ret, frame = cap.read()\n            if not ret:\n                continue\n                \n            # Brightness\n            gray = cv2.cvtColor(frame, cv2.COLOR_BGR2GRAY)\n            avg_brightness.append(np.mean(gray))\n            \n            # Motion (optical flow)\n            if prev_gray is not None:\n                flow = cv2.calcOpticalFlowFarneback(prev_gray, gray, None, 0.5, 3, 15, 3, 5, 1.2, 0)\n                magnitude = np.sqrt(flow[..., 0]**2 + flow[..., 1]**2)\n                avg_motion.append(np.mean(magnitude))\n                \n            prev_gray = gray\n        \n        cap.release()\n        \n        # Compile features\n        video_features = [\n            width, \n            height,\n            fps,\n            duration,\n            np.mean(avg_brightness) if avg_brightness else 0,\n            np.std(avg_brightness) if len(avg_brightness) > 1 else 0,\n            np.mean(avg_motion) if avg_motion else 0,\n            np.max(avg_motion) if avg_motion else 0,\n            np.std(avg_motion) if len(avg_motion) > 1 else 0\n        ]\n        \n        # For positive samples, add time features\n        if row['target'] == 1:\n            video_features.extend([\n                row['time_of_event'],\n                row['time_of_alert'],\n                row['time_of_event'] - row['time_of_alert']\n            ])\n        else:\n            # For negative samples, use zeros for time features\n            video_features.extend([0, 0, 0])\n        \n        features.append(video_features)\n        labels.append(row['target'])\n    \n    # Feature names for reference\n    feature_names = [\n        'width', 'height', 'fps', 'duration', \n        'avg_brightness', 'std_brightness',\n        'avg_motion', 'max_motion', 'std_motion',\n        'event_time', 'alert_time', 'reaction_time'\n    ]\n    \n    # Convert to numpy arrays\n    X = np.array(features)\n    y = np.array(labels)\n    \n    # Normalize features\n    scaler = StandardScaler()\n    X_scaled = scaler.fit_transform(X)\n    \n    print(f\"Extracted {X.shape[1]} features from {X.shape[0]} videos\")\n    \n    return X_scaled, y, feature_names","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-04T05:07:49.605625Z","iopub.execute_input":"2025-04-04T05:07:49.605995Z","iopub.status.idle":"2025-04-04T05:07:49.626413Z","shell.execute_reply.started":"2025-04-04T05:07:49.605959Z","shell.execute_reply":"2025-04-04T05:07:49.625298Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Add this to your main function\ndef analyze_and_extract_features(train_df, test_df, train_video_dir):\n    \"\"\"\n    Perform analysis and feature extraction\n    \n    Args:\n        train_df: Training dataframe\n        test_df: Testing dataframe\n        train_video_dir: Directory containing training videos\n        \n    Returns:\n        X: Feature matrix\n        y: Target labels\n        feature_names: Names of extracted features\n    \"\"\"\n    # Perform exploratory data analysis\n    perform_eda(train_df, train_video_dir)\n    \n    # Extract features for model building (optional)\n    # Uncomment to extract features\n    # X, y, feature_names = extract_basic_features(train_df, train_video_dir, sample_size=100)\n    # return X, y, feature_names\n    \n    return None, None, None","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-04T05:07:49.627461Z","iopub.execute_input":"2025-04-04T05:07:49.627818Z","iopub.status.idle":"2025-04-04T05:07:49.649037Z","shell.execute_reply.started":"2025-04-04T05:07:49.627790Z","shell.execute_reply":"2025-04-04T05:07:49.647711Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Run the code\nif __name__ == \"__main__\":\n    import os\n\n    # Find CSV file paths\n    train_csv_path = None\n    test_csv_path = None\n    train_video_dir = None\n    \n    for dirname, _, filenames in os.walk('/kaggle/input'):\n        for filename in filenames:\n            if filename == 'train.csv':\n                train_csv_path = os.path.join(dirname, filename)\n            elif filename == 'test.csv':\n                test_csv_path = os.path.join(dirname, filename)\n    \n    for dirname, _, filenames in os.walk('/kaggle/input'):\n        if os.path.basename(dirname) == 'train' and any(f.endswith('.mp4') for f in filenames):\n            train_video_dir = dirname\n            break\n    \n    # Run EDA functions only\n    if train_csv_path and test_csv_path and train_video_dir:\n        train_df, test_df = explore_data(train_csv_path, test_csv_path, train_video_dir)\n        \n        if train_df is not None:\n            print(\"\\nStarting additional data analysis...\")\n            \n            # Time features analysis\n            analyze_time_features(train_df)\n            \n            # Video properties analysis (reduced sample size)\n            analyze_video_properties(train_df, train_video_dir, sample_size=20)\n            \n            # Optical flow analysis (reduced sample size)\n            analyze_sample_optical_flow(train_df, train_video_dir, num_samples=2)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-04T05:07:49.650201Z","iopub.execute_input":"2025-04-04T05:07:49.650603Z","iopub.status.idle":"2025-04-04T05:08:16.361189Z","shell.execute_reply.started":"2025-04-04T05:07:49.650558Z","shell.execute_reply":"2025-04-04T05:08:16.360130Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## References\n- Inspired by the approach demonstrated in [💥Nexar DCP Challenge - Baseline 💡](https://www.kaggle.com/code/sandiledesmondmfazi/nexar-dcp-challenge-baseline)\n\n### Note\nBeing new to video data analysis, I developed this notebook with assistance from Claude AI to better understand the key concepts and approaches for analyzing dashcam footage. I welcome any feedback to improve this code.","metadata":{}}]}