{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"codemirror_mode":{"name":"ipython","version":2},"file_extension":".py","mimetype":"text/x-python","name":"python","nbconvert_exporter":"python","pygments_lexer":"ipython2","version":"2.7.6"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":127283,"databundleVersionId":15634477,"sourceType":"competition"}],"dockerImageVersionId":31259,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# 🚗 Accident Prediction – Baseline & Rule-Based Type Modeling\n\n### Overview\n\nThis notebook builds a **naive baseline submission** and a **rule-based accident type prediction submission** using the provided test_metadata.csv.\nThe goal is to generate predictions for the accident ```type``` using metadata only (no use of video clips).\nThe ```accident_time```, ```center_x```, and  ```center_y``` are generated naively using middle frame or frame's center.\n\n### 🎯 Purpose of This Notebook\n-  Establish a baseline benchmark.\n-  Provide simple and interpretable rule-based logic.\n-  Build a foundation for more advanced modeling approaches.","metadata":{}},{"cell_type":"markdown","source":"## Imports and Setup","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport os\nimport os.path as osp\n\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nfrom tqdm import tqdm","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-17T12:18:43.171470Z","iopub.execute_input":"2026-02-17T12:18:43.172582Z","iopub.status.idle":"2026-02-17T12:18:44.523918Z","shell.execute_reply.started":"2026-02-17T12:18:43.172540Z","shell.execute_reply":"2026-02-17T12:18:44.522996Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"METADATA_PATH = \"/kaggle/input/accident/test_metadata.csv\"  # Path to the test set .csv file\nmetadata_df = pd.read_csv(METADATA_PATH)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-17T12:18:44.525198Z","iopub.execute_input":"2026-02-17T12:18:44.525641Z","iopub.status.idle":"2026-02-17T12:18:44.561402Z","shell.execute_reply.started":"2026-02-17T12:18:44.525614Z","shell.execute_reply":"2026-02-17T12:18:44.560598Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 📊 Exploratory Data Analysis (EDA)\n\nTo better understand the dataset, we analyze categorical metadata features.\n\n","metadata":{}},{"cell_type":"code","source":"category_names = [\"region\", \"scene_layout\", \"weather\", \"day_time\", \"quality\"]\ncontinuous_names = [\"duration\", \"height\", \"width\"]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-17T12:18:44.562432Z","iopub.execute_input":"2026-02-17T12:18:44.562712Z","iopub.status.idle":"2026-02-17T12:18:44.566669Z","shell.execute_reply.started":"2026-02-17T12:18:44.562686Z","shell.execute_reply":"2026-02-17T12:18:44.565859Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"For each categorical feature:\n- Count distributions are visualized.\n- Percentages are annotated.\n- Categories are ordered by frequency\n\nThis provides insights into:\n- Most common road layouts\n- Weather distribution\n- Day vs night balance\n- Regional representation\n\nThese distributions guide the rule-based logic in the next section.","metadata":{}},{"cell_type":"code","source":"sns.set_style(\"whitegrid\")\nfig, axes = plt.subplots(len(category_names), 1, figsize=(10, len(category_names) * 4))\n\nfor ax, category_name in zip(axes, category_names):\n    order = metadata_df[category_name].value_counts().index\n    total = len(metadata_df)\n\n    sns.countplot(\n        x=category_name,\n        hue=category_name,\n        data=metadata_df,\n        ax=ax,\n        order=order,\n        hue_order=order,\n        palette=\"magma\"\n    )\n\n    ax.tick_params(axis='x', rotation=45)\n\n    for p in ax.patches:\n        height = p.get_height()\n        percentage = 100 * height / total\n        ax.annotate(f'{percentage:.1f}%',\n                    (p.get_x() + p.get_width() / 2., height),\n                    ha='center',\n                    va='bottom',\n                    fontsize=10,\n                    xytext=(0, 3),\n                    textcoords='offset points')\n\n    ax.set_title(f\"Distribution of `{category_name}`\")\n    ax.set_xlabel(\" \".join(category_name.capitalize().split(\"_\")))\n    ax.set_ylabel(\"Video Count\")\n\n\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-17T12:18:44.568757Z","iopub.execute_input":"2026-02-17T12:18:44.569015Z","iopub.status.idle":"2026-02-17T12:18:45.759311Z","shell.execute_reply.started":"2026-02-17T12:18:44.568992Z","shell.execute_reply":"2026-02-17T12:18:45.758411Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 🧩 Rule-Based Accident Type Prediction\n\nInstead of random guessing, we introduce a heuristic scoring system based on metadata features:\n\n- ```scene_layout```\n- ```weather```\n- ```day_time```\n\nEach feature contributes weighted scores to accident types.","metadata":{}},{"cell_type":"code","source":"def normalize_scores(scores: dict[str, float]) -> dict[str, float]:\n    \"\"\"Converts raw weights into probabilities to ensure same contribution by each predictor.\"\"\"\n    total = sum(scores.values())\n    scores = {key: val / total for key, val in scores.items()}\n    return scores","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-17T12:18:45.760392Z","iopub.execute_input":"2026-02-17T12:18:45.760666Z","iopub.status.idle":"2026-02-17T12:18:45.765345Z","shell.execute_reply.started":"2026-02-17T12:18:45.760640Z","shell.execute_reply":"2026-02-17T12:18:45.764539Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 🛣 Scene Layout-Based Prediction\n| Scene Layout                 | Higher Probability Accident Type |\n| ---------------------------- | -------------------------------- |\n| highway                      | single, rear-end                 |\n| signalized_intersection      | t-bone                           |\n| simple_intersection          | t-bone                           |\n| grade_separated_intersection | t-bone, single                   |\n| tunnel                       | single, sideswipe                |\n| parking_lot                  | single                           |\n| roundabout                   | t-bone                           |\n","metadata":{}},{"cell_type":"code","source":"def predict_scene_layout(scene_layout: str) -> dict[str, float]:\n    scores = {\n        \"rear-end\": 1,\n        \"t-bone\": 1,\n        \"single\": 1,\n        \"head-on\": 1,\n        \"sideswipe\": 1,\n    }\n\n    if scene_layout == \"highway\":\n        scores[\"single\"] = 5\n        scores[\"rear-end\"] = 3\n        scores[\"sideswipe\"] = 2\n    elif scene_layout == \"signalized_intersection\":\n        scores[\"t-bone\"] = 10\n    elif scene_layout == \"simple_intersection\":\n        scores[\"t-bone\"] = 10\n    elif scene_layout == \"grade_separated_intersection\":\n        scores[\"t-bone\"] = 5\n        scores[\"single\"] = 5\n    elif scene_layout == \"tunnel\":\n        scores[\"single\"] = 5\n        scores[\"sideswipe\"] = 3\n    elif scene_layout == \"parking_lot\":\n        scores[\"single\"] = 5\n    elif scene_layout == \"roundabout\":\n        scores[\"t-bone\"] = 5\n    scores = normalize_scores(scores)\n    return scores","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-17T12:18:45.766535Z","iopub.execute_input":"2026-02-17T12:18:45.766897Z","iopub.status.idle":"2026-02-17T12:18:45.772814Z","shell.execute_reply.started":"2026-02-17T12:18:45.766863Z","shell.execute_reply":"2026-02-17T12:18:45.772019Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 🌧 Weather-Based Prediction\n\n**Rules**:\n- ```rain``` → higher probability of single\n- ```snow``` → higher probability of single\n- ```normal``` → uniform distribution\n\n**Reasoning**:\nSlippery roads increase single-vehicle accidents.","metadata":{}},{"cell_type":"code","source":"def predict_weather(weather: str) -> dict[str, float]:\n    scores = {\n        \"rear-end\": 1,\n        \"t-bone\": 1,\n        \"single\": 1,\n        \"head-on\": 1,\n        \"sideswipe\": 1,\n    }\n\n    if weather == \"normal\":\n        pass\n    elif weather == \"rain\":\n        scores[\"single\"] = 5\n    elif weather == \"snow\":\n        scores[\"single\"] = 5\n\n    scores = normalize_scores(scores)\n    return scores","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-17T12:18:45.773860Z","iopub.execute_input":"2026-02-17T12:18:45.774155Z","iopub.status.idle":"2026-02-17T12:18:45.778773Z","shell.execute_reply.started":"2026-02-17T12:18:45.774122Z","shell.execute_reply":"2026-02-17T12:18:45.778081Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 🌙 Day/Night Prediction\n\n**Rules**:\n- ```night``` → increased probability of single\n- ```day``` → uniform distribution\n\n**Reasoning**:\nReduced visibility increases loss-of-control incidents.","metadata":{}},{"cell_type":"code","source":"def predict_day_time(day_time: str) -> dict[str, float]:\n    scores = {\n        \"rear-end\": 1,\n        \"t-bone\": 1,\n        \"single\": 1,\n        \"head-on\": 1,\n        \"sideswipe\": 1,\n    }\n    if day_time == \"day\":\n        pass\n    elif day_time == \"night\":\n        scores[\"single\"] = 3\n\n    scores = normalize_scores(scores)\n    return scores","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-17T12:18:45.779699Z","iopub.execute_input":"2026-02-17T12:18:45.780426Z","iopub.status.idle":"2026-02-17T12:18:45.784741Z","shell.execute_reply.started":"2026-02-17T12:18:45.780381Z","shell.execute_reply":"2026-02-17T12:18:45.783985Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🧮 Combined Prediction Logic\n\n**Steps**:\n1. Initialize zero scores for all accident types.\n2. Add normalized scores from:\n    - Scene layout\n    - Weather\n    - Day time\n3. Select the type with the highest total score.\n\nFinal prediction is achieved by taking the max scoring label, which effectively creates a **weighted voting system**.","metadata":{}},{"cell_type":"code","source":"def predict_accident(row):\n    results = {\n        \"rear-end\": 0,\n        \"t-bone\": 0,\n        \"single\": 0,\n        \"head-on\": 0,\n        \"sideswipe\": 0\n    }\n    scores = predict_scene_layout(row.scene_layout)\n    for key, value in scores.items():\n        results[key] += value\n\n    scores = predict_weather(row.weather)\n    for key, value in scores.items():\n        results[key] += value\n\n    scores = predict_day_time(row.day_time)\n    for key, value in scores.items():\n        results[key] += value\n\n    return max(results, key=results.get)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-17T12:18:45.785747Z","iopub.execute_input":"2026-02-17T12:18:45.785999Z","iopub.status.idle":"2026-02-17T12:18:45.790818Z","shell.execute_reply.started":"2026-02-17T12:18:45.785976Z","shell.execute_reply":"2026-02-17T12:18:45.790234Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"type_prediction = []\nfor i, row in tqdm(metadata_df.iterrows(), total=len(metadata_df)):\n    type_prediction.append(predict_accident(row))\ntype_prediction = np.array(type_prediction)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-17T12:18:45.792637Z","iopub.execute_input":"2026-02-17T12:18:45.792872Z","iopub.status.idle":"2026-02-17T12:18:45.944777Z","shell.execute_reply.started":"2026-02-17T12:18:45.792850Z","shell.execute_reply":"2026-02-17T12:18:45.943669Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Naive Temporal\n\n**Assumption**:\nThe accident occurs halfway through the video.\n\n**Rationale**:\nWith no temporal modeling, predicting the midpoint of the clip is a reasonable neutral baseline.","metadata":{}},{"cell_type":"code","source":"temporal_naive = np.array(metadata_df[\"duration\"] / 2)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-17T12:18:45.946011Z","iopub.execute_input":"2026-02-17T12:18:45.946753Z","iopub.status.idle":"2026-02-17T12:18:45.950548Z","shell.execute_reply.started":"2026-02-17T12:18:45.946723Z","shell.execute_reply":"2026-02-17T12:18:45.949745Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Naive Spatial\n\n**Assumption**: The accident occurs at the center of the frame (```center_x = 0.5```, ```center_y = 0.5```).\n\n**Rationale**:\nIn the absence of object detection or localization models, predicting the image center is a common baseline.\n","metadata":{}},{"cell_type":"code","source":"spatial_naive = np.ones((len(metadata_df), 2)) * 0.5","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-17T12:18:45.951548Z","iopub.execute_input":"2026-02-17T12:18:45.951784Z","iopub.status.idle":"2026-02-17T12:18:45.955496Z","shell.execute_reply.started":"2026-02-17T12:18:45.951762Z","shell.execute_reply":"2026-02-17T12:18:45.954773Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Generating the Improved Submission\n\n","metadata":{}},{"cell_type":"code","source":"accident_type_submission = pd.DataFrame.from_dict({\n    \"path\": metadata_df[\"path\"].values,\n    \"accident_time\": temporal_naive,\n    \"center_x\": spatial_naive[:, 0],\n    \"center_y\": spatial_naive[:, 1],\n    \"type\": type_prediction,\n})\naccident_type_submission.to_csv(\"./submission.csv\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-17T12:18:45.956437Z","iopub.execute_input":"2026-02-17T12:18:45.956757Z","iopub.status.idle":"2026-02-17T12:18:45.976051Z","shell.execute_reply.started":"2026-02-17T12:18:45.956723Z","shell.execute_reply":"2026-02-17T12:18:45.975290Z"}},"outputs":[],"execution_count":null}]}