{"cells":[{"metadata":{},"cell_type":"markdown","source":"# Visualizing the predictions"},{"metadata":{},"cell_type":"markdown","source":"In this notebook, I've shared my prediction visualization code. The code expects a csv file with predicted and ground truth boxes in `PredictionString` format (as used by `train.csv` and `sample_submission.csv`, see [here](https://www.kaggle.com/c/3d-object-detection-for-autonomous-vehicles/data) for more information)\n\nHow this works:\n\n* We convert the predicted and ground truth strings to [Box](https://github.com/lyft/nuscenes-devkit/blob/master/lyft_dataset_sdk/utils/data_classes.py#L478) format\n* Use matplotlib for BEV visualization\n* Use plotly for 3D interactive visualization\n\nI'm all open for new suggestions, drop a comment :)"},{"metadata":{"trusted":true,"_kg_hide-output":true},"cell_type":"code","source":"!pip install -U git+https://github.com/lyft/nuscenes-devkit","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"from copy import deepcopy\nimport pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport plotly.graph_objects as go\n\nfrom lyft_dataset_sdk.lyftdataset import LyftDataset, Box, Quaternion, view_points\n\n%matplotlib inline","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"I've uploaded a sample csv file for demo, here's what it looks like:"},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"df = pd.read_csv('../input/lyftpredictionvisualization/lyft.csv')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"df.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Here `GroundTruthString` is the ground truth string and `PredictionString` is the predicted string for the corresponding sample token `Id`, both the strings follow the format of `PredictionString` in `train.csv`"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"# some black magic, see https://www.kaggle.com/seshurajup/starter-lyft-level-5-av-dataset-from-github#625566\n!ln -s /kaggle/input/3d-object-detection-for-autonomous-vehicles/train_images images\n!ln -s /kaggle/input/3d-object-detection-for-autonomous-vehicles/train_maps maps\n!ln -s /kaggle/input/3d-object-detection-for-autonomous-vehicles/train_lidar lidar\n!ln -s /kaggle/input/3d-object-detection-for-autonomous-vehicles/train_data data","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"lyft = LyftDataset(data_path='.', json_path='data', verbose=False)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Some utility functions"},{"metadata":{"trusted":true},"cell_type":"code","source":"def get2Box(boxes_list, names_list, token, scores=None):\n    '''Given a list of boxes in `x,y,z,w,l,h,yaw` format, returns them in `Box` format\n    \n    Args:\n    boxes_list: a list of boxes in [x, y, z, w, l, h, yaw] format\n    names_list: classes the boxes belong to\n    token: token of the sample the boxes belong to\n    scores: predicted confidence scores, only for predicted boxes, \n    '''\n    boxes = []\n    for idx in range(len(boxes_list)):\n        center = boxes_list[idx, :3] # x, y, z\n        yaw = boxes_list[idx, 6]\n        size = boxes_list[idx, 3:6] # w, l, h\n        name = names_list[idx]\n        detection_score = 1.0 # for ground truths \n        if scores is not None:\n            detection_score = scores[idx]\n        quat = Quaternion(axis=[0, 0, 1], radians=yaw)\n        box = Box(\n            center=center,\n            size=size,\n            orientation=quat,\n            score=detection_score,\n            name=name,\n            token=token\n        )\n        boxes.append(box)\n    return boxes\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def get_pred_gt(pred_df, idx): \n    '''Given an index `idx`, this function reads ground truth and predicted strings and returns\n    corresponding boxes in `Box` format'''\n    \n    sample_token = pred_df.iloc[idx]['Id']\n    \n    string = pred_df.iloc[idx]['GroundTruthString'].split()\n    gt_objects = [string[x:x+8] for x in range(0, len(string), 8)]\n    \n    string = pred_df.iloc[idx]['PredictionString'].split()\n    pred_objects = [string[x:x+9] for x in range(0, len(string), 9)]\n    \n    # str -> float, in x,y,z,w,l,h,yaw format\n    gt_boxes = np.array([list(map(float, x[0:7])) for x in gt_objects])\n    gt_class = np.array([x[7] for x in gt_objects])\n    \n    pred_scores = np.array([float(x[0]) for x in pred_objects])\n    pred_boxes = np.array([list(map(float, x[1:8])) for x in pred_objects])\n    pred_class = np.array([x[8] for x in pred_objects])\n    \n    # x,y,z,w,l,h,yaw -> Box instance\n    predBoxes = get2Box(pred_boxes, pred_class, sample_token, scores=pred_scores)\n    gtBoxes = get2Box(gt_boxes, gt_class, sample_token)\n    \n    return predBoxes, gtBoxes ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def glb_to_sensor(box, sample_data):\n    '''Get a box from global frame to sensor's frame of reference '''\n    \n    box = box.copy() # v.imp\n    cs_record = lyft.get('calibrated_sensor', sample_data['calibrated_sensor_token'])\n    pose_record = lyft.get('ego_pose', sample_data['ego_pose_token'])\n    \n    # global to ego \n    box.translate(-np.array(pose_record['translation']))\n    box.rotate(Quaternion(pose_record['rotation']).inverse)\n    \n    # ego to sensor\n    box.translate(-np.array(cs_record['translation']))\n    box.rotate(Quaternion(cs_record['rotation']).inverse)\n    \n    return box","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# `string` -> `Box`"},{"metadata":{"trusted":true},"cell_type":"code","source":"# get predicted and ground boxes for each sample in `Box` format\npred_boxes = []\ngt_boxes = []\nfor idx in range(len(df)):\n    pBoxes, gBoxes = get_pred_gt(df, idx)\n    pred_boxes.append(pBoxes)\n    gt_boxes.append(gBoxes)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# let's take a peek\nidx = 0\npred_boxes[idx][0]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"gt_boxes[idx][0]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# BEV Visualization"},{"metadata":{"trusted":true},"cell_type":"code","source":"idx = 0 # change this to visualize other samples\nsample_token = df.iloc[idx]['Id']\n\nsample = lyft.get('sample', sample_token)\nsample_data = lyft.get('sample_data', sample['data']['LIDAR_TOP'])\npath = sample_data['filename']\nlidar_points = np.fromfile(path, dtype=np.float32, count=-1).reshape([-1, 5])[:, :4]\n\n_, ax = plt.subplots(1, 1, figsize=(9, 9))\n\n# create colors based on the distance of the point from lidar\naxes_limit=40\n\ndists = np.sqrt(np.sum(lidar_points[:, :2] ** 2, axis=1))\ncolors = np.minimum(1, dists / axes_limit / np.sqrt(2))\nax.scatter(lidar_points[:, 0], lidar_points[:, 1], c=colors, s=0.2)\nax.plot(0, 0, \"x\", color=\"red\") # plot lidar location\n\n# Limit visible range.\nax.set_xlim(-axes_limit, axes_limit)\nax.set_ylim(-axes_limit, axes_limit)\n\n# plot the ground truths\nfor box in gt_boxes[idx]:\n    box = glb_to_sensor(box, sample_data)\n    c = np.array([255, 158, 0 ]) / 255.0 # Orange\n    box.render(ax, view=np.eye(4), colors=(c, c, c))\n\n# plot the predicted boxes\nfor box in pred_boxes[idx]:\n    box = glb_to_sensor(box, sample_data)\n    c = np.array([0, 0, 230]) / 255.0 # Blue\n    box.render(ax, view=np.eye(4), colors=(c, c, c))\n\n# gotta invert for consistency with lyft's inbuilt plots\nplt.gca().invert_yaxis()\nplt.gca().invert_xaxis()\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The Orange ones are the ground truths and theh blue ones are the predictions, we can plot ground truths using lyft's inbuilt render_sample_data function"},{"metadata":{"trusted":true},"cell_type":"code","source":"lyft.render_sample_data(sample_data['token'])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"That's all for the BEV visualization, let's do 3D interactive visualization"},{"metadata":{},"cell_type":"markdown","source":"# 3D Interactive visualization"},{"metadata":{"trusted":true},"cell_type":"code","source":"def get_lines(boxes, name):\n    '''Takes in boxes, extracts edges and returns `go.Scatter3d` object for those lines'''\n    \n    x_lines = []\n    y_lines = []\n    z_lines = []\n\n    def f_lines_add_nones():\n        x_lines.append(None)\n        y_lines.append(None)\n        z_lines.append(None)\n\n    ixs_box_0 = [0, 1, 2, 3, 0]\n    ixs_box_1 = [4, 5, 6, 7, 4]\n\n    for box in boxes:\n        box = glb_to_sensor(box, sample_data)\n        points = view_points(box.corners(), view=np.eye(3), normalize=False)\n        x_lines.extend(points[0, ixs_box_0])\n        y_lines.extend(points[1, ixs_box_0])\n        z_lines.extend(points[2, ixs_box_0])\n        f_lines_add_nones()\n        x_lines.extend(points[0, ixs_box_1])\n        y_lines.extend(points[1, ixs_box_1])\n        z_lines.extend(points[2, ixs_box_1])\n        f_lines_add_nones()\n        for i in range(4):\n            x_lines.extend(points[0, [ixs_box_0[i], ixs_box_1[i]]])\n            y_lines.extend(points[1, [ixs_box_0[i], ixs_box_1[i]]])\n            z_lines.extend(points[2, [ixs_box_0[i], ixs_box_1[i]]])\n            f_lines_add_nones()\n\n    lines = go.Scatter3d(x=x_lines, y=y_lines, z=z_lines, mode=\"lines\", name=name)\n    return lines","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"idx = 0 # change this to visualize other samples\nsample_token = df.iloc[idx]['Id']\n\nsample = lyft.get(\"sample\", sample_token)\nsample_data = lyft.get(\"sample_data\", sample[\"data\"][\"LIDAR_TOP\"])\npath = sample_data['filename']\nlidar_points = np.fromfile(path, dtype=np.float32, count=-1).reshape([-1, 5])[:, :4]\n\n# plot the points\ndf_tmp = pd.DataFrame(lidar_points[:, :3], columns=[\"x\", \"y\", \"z\"])\ndf_tmp[\"norm\"] = np.sqrt(np.power(df_tmp[[\"x\", \"y\", \"z\"]].values, 2).sum(axis=1))\nscatter = go.Scatter3d(\n    x=df_tmp[\"x\"],\n    y=df_tmp[\"y\"],\n    z=df_tmp[\"z\"],\n    mode=\"markers\",\n    marker=dict(size=1, color=df_tmp[\"norm\"], opacity=0.8),\n)\n\ngt_lines = get_lines(gt_boxes[idx], 'gt_boxes')\npred_lines = get_lines(pred_boxes[idx], 'pred_boxes')\nfig = go.Figure(data=[scatter, gt_lines, pred_lines])\nfig.update_layout(scene_aspectmode=\"data\")\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"When visualizing above 3d plot, for a mouse: you can use scroll to zoom in/out, use left button to rotate the plot, use right button to translate the plot. This plot works smoothly when kernel is in edit mode, the rendered version after kernel commit is a bit laggy."},{"metadata":{},"cell_type":"markdown","source":"<font color='red'>Do upvote if you liked this kernel :) </font> 😇"},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":1}