{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Introduction\n\nIn this notebook, we explore the dataset provided for the Image Matching Challenge 2023. The goal of this competition is to generate 3D reconstructions from a set of images and accurately estimate the camera poses (rotation and translation) for each image. In this initial exploration, we perform an Exploratory Data Analysis (EDA) to better understand the dataset and its structure before diving into the development of machine learning models.","metadata":{}},{"cell_type":"markdown","source":"# The Plan\nOur plan for this EDA consists of the following steps:\n\n1. Load and inspect the dataset\n2. Analyze the distribution of scenes and datasets\n3. Visualize some sample images\n4. Explore the 3D reconstructions\n5. Analyze the distribution of rotation and translation values\n6. Visualize rotation and translation as quaternions\n7. Calculate relative pose errors\n8. Plot the distribution of pose errors","metadata":{}},{"cell_type":"markdown","source":"The first thing I want to do is to download read_write_model.py to do some 3D exploration.","metadata":{}},{"cell_type":"code","source":"!wget https://raw.githubusercontent.com/colmap/colmap/dev/scripts/python/read_write_model.py","metadata":{"execution":{"iopub.status.busy":"2023-05-02T07:56:21.425317Z","iopub.execute_input":"2023-05-02T07:56:21.425797Z","iopub.status.idle":"2023-05-02T07:56:22.622527Z","shell.execute_reply.started":"2023-05-02T07:56:21.425758Z","shell.execute_reply":"2023-05-02T07:56:22.620953Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Import the libraries\nimport os\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport random\nfrom PIL import Image\nimport pycolmap\nimport read_write_model","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-05-02T08:07:29.708366Z","iopub.execute_input":"2023-05-02T08:07:29.708786Z","iopub.status.idle":"2023-05-02T08:07:29.714875Z","shell.execute_reply.started":"2023-05-02T08:07:29.708746Z","shell.execute_reply":"2023-05-02T08:07:29.713811Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load the Dataset\n\nWe start by loading the training labels CSV file, which contains information about the images, their corresponding datasets and scenes, and the ground truth camera poses (rotation and translation).","metadata":{}},{"cell_type":"code","source":"# Load the data\nroot_dir = '/kaggle/input/image-matching-challenge-2023'\ntrain_labels_file = os.path.join(root_dir, 'train/train_labels.csv')\ntrain_labels = pd.read_csv(train_labels_file)\n\ntrain_labels.head()","metadata":{"execution":{"iopub.status.busy":"2023-05-02T08:07:30.146474Z","iopub.execute_input":"2023-05-02T08:07:30.146857Z","iopub.status.idle":"2023-05-02T08:07:30.164314Z","shell.execute_reply.started":"2023-05-02T08:07:30.146826Z","shell.execute_reply":"2023-05-02T08:07:30.163230Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_labels.columns","metadata":{"execution":{"iopub.status.busy":"2023-05-02T08:07:30.413551Z","iopub.execute_input":"2023-05-02T08:07:30.413952Z","iopub.status.idle":"2023-05-02T08:07:30.421227Z","shell.execute_reply.started":"2023-05-02T08:07:30.413918Z","shell.execute_reply":"2023-05-02T08:07:30.419966Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Analyze the Distribution of Scenes and Datasets\n\nHere, we investigate the number of scenes per dataset and the number of images per scene. This helps us understand the dataset's structure and its complexity.","metadata":{}},{"cell_type":"code","source":"# Analyze labels distribution\ndataset_counts = train_labels['dataset'].value_counts()\nscene_counts = train_labels['scene'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2023-05-02T08:07:31.296987Z","iopub.execute_input":"2023-05-02T08:07:31.297374Z","iopub.status.idle":"2023-05-02T08:07:31.306478Z","shell.execute_reply.started":"2023-05-02T08:07:31.297343Z","shell.execute_reply":"2023-05-02T08:07:31.305350Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(10, 5))\nplt.bar(dataset_counts.index, dataset_counts.values)\nplt.xlabel('Dataset')\nplt.ylabel('Number of Images')\nplt.title('Distribution of Images per Dataset')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-05-02T08:07:31.579102Z","iopub.execute_input":"2023-05-02T08:07:31.579467Z","iopub.status.idle":"2023-05-02T08:07:31.827846Z","shell.execute_reply.started":"2023-05-02T08:07:31.579439Z","shell.execute_reply":"2023-05-02T08:07:31.826965Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The distribution of images per dataset shows that the dataset is somewhat imbalanced, with the majority of the images belonging to the \"heritage\" dataset (247 images), followed by \"haiper\" (54 images), and lastly \"urban\" (26 images). This imbalance may affect the model's performance, as it may become biased towards the more dominant dataset (heritage) during training. It is essential to be aware of this imbalance and consider strategies to address it, such as data augmentation, oversampling, or using a different loss function that takes into account class imbalance.","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(10, 5))\nplt.bar(scene_counts.index, scene_counts.values)\nplt.xlabel('Scene')\nplt.ylabel('Number of Images')\nplt.title('Distribution of Images per Scene')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-05-02T08:07:31.961254Z","iopub.execute_input":"2023-05-02T08:07:31.961664Z","iopub.status.idle":"2023-05-02T08:07:32.247435Z","shell.execute_reply.started":"2023-05-02T08:07:31.961632Z","shell.execute_reply":"2023-05-02T08:07:32.246449Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The distribution of images per scene also shows some imbalance, with the \"dioscuri\" scene containing the most images (174), followed by \"wall\" (43), \"cyprus\" (30), \"kyiv-puppet-theater\" (26), \"fountain\" (23), \"chairs\" (16), and \"bike\" (15). This distribution indicates that the complexity and the number of images per scene vary significantly. It is important to consider this variation when training the model, as it may be necessary to ensure that the model can handle different complexities and scales of scenes.","metadata":{}},{"cell_type":"markdown","source":"# Visualizes the Images\n\nTo make sure we know how the data looks like and to understand it even more, we need to visualize the image in our dataset.","metadata":{}},{"cell_type":"code","source":"# Visualize images\nnum_images_to_show = 5\nrandom_image_indices = random.sample(range(len(train_labels)), num_images_to_show)\n\nfig, axes = plt.subplots(1, num_images_to_show, figsize=(20, 5))\ntrain_folder = os.path.join(root_dir, 'train')\n\nfor i, image_index in enumerate(random_image_indices):\n    relative_image_path = train_labels.iloc[image_index]['image_path']\n    image_path = os.path.join(train_folder, relative_image_path)\n    img = Image.open(image_path)\n    axes[i].imshow(img)\n    axes[i].set_title(f\"Dataset: {train_labels.iloc[image_index]['dataset']} | Scene: {train_labels.iloc[image_index]['scene']}\")\n    axes[i].axis('off')\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-05-02T08:07:34.267775Z","iopub.execute_input":"2023-05-02T08:07:34.268139Z","iopub.status.idle":"2023-05-02T08:07:35.774772Z","shell.execute_reply.started":"2023-05-02T08:07:34.268110Z","shell.execute_reply":"2023-05-02T08:07:35.773818Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Explore 3D Reconstructions\n\nIn this step, we explore the 3D reconstructions provided for the training datasets. These reconstructions contain information about the 3D structure of the scene and the camera poses used to create the images. By visualizing these reconstructions, we can gain insights into the spatial arrangement of the cameras, which can help us understand the complexity of the problem and the relationships between the images and their corresponding camera poses.\n\nDue to computational limitations and the speed of processing in the Kaggle notebook, we reduced the number of points and camera poses for visualization purposes. This allows us to have a faster, albeit less detailed, exploration of the 3D reconstructions. Despite this reduction in complexity, the visualizations still provide useful information about the spatial distribution of the cameras and the overall structure of the scenes, which can be helpful in guiding the development of our model.","metadata":{}},{"cell_type":"code","source":"# Explore 3D reconstructions\ndef plot_sfm_3d_reconstruction(reconstruction, num_points=1000, num_cameras=50):\n    fig = plt.figure()\n    ax = fig.add_subplot(111, projection='3d')\n\n    # Plot 3D points\n    point3D_ids = list(reconstruction[2].keys())\n    selected_point3D_ids = random.sample(point3D_ids, min(num_points, len(point3D_ids)))\n\n    for point3D_id in selected_point3D_ids:\n        point3D = reconstruction[2][point3D_id]\n        ax.scatter(point3D.xyz[0], point3D.xyz[1], point3D.xyz[2], c='b', marker='o')\n\n    # Plot camera poses\n    image_ids = list(reconstruction[1].keys())\n    selected_image_ids = random.sample(image_ids, min(num_cameras, len(image_ids)))\n\n    for image_id in selected_image_ids:\n        image = reconstruction[1][image_id]\n        camera_center = -image.qvec2rotmat().T @ image.tvec\n        ax.scatter(camera_center[0], camera_center[1], camera_center[2], c='r', marker='^')\n\n    ax.set_xlabel('X')\n    ax.set_ylabel('Y')\n    ax.set_zlabel('Z')\n    plt.show()\n\n# You can replace the following path with any other scene's sfm folder path\nsample_sfm_folder = os.path.join(train_folder, 'urban/kyiv-puppet-theater/sfm')\ncameras, images, points3D = read_write_model.read_model(sample_sfm_folder)\nreconstruction = (cameras, images, points3D)\nplot_sfm_3d_reconstruction(reconstruction)","metadata":{"execution":{"iopub.status.busy":"2023-05-02T07:56:28.465833Z","iopub.execute_input":"2023-05-02T07:56:28.466205Z","iopub.status.idle":"2023-05-02T07:56:45.486125Z","shell.execute_reply.started":"2023-05-02T07:56:28.466171Z","shell.execute_reply":"2023-05-02T07:56:45.485030Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Analyze the distribution of rotation and translation values\n\nNow, what we want to do is, we want to create histograms of the ground truth rotation and translation values to understand their distribution. This can help us identify potential patterns or trends in the camera poses.","metadata":{}},{"cell_type":"code","source":"# Inspect rotation matrices and translation vectors\ndef string_to_matrix(matrix_string):\n    return np.array(list(map(float, matrix_string.split(';')))).reshape(3, 3)\n\ndef string_to_vector(vector_string):\n    return np.array(list(map(float, vector_string.split(';'))))","metadata":{"execution":{"iopub.status.busy":"2023-05-02T08:14:00.980392Z","iopub.execute_input":"2023-05-02T08:14:00.980812Z","iopub.status.idle":"2023-05-02T08:14:00.986874Z","shell.execute_reply.started":"2023-05-02T08:14:00.980777Z","shell.execute_reply":"2023-05-02T08:14:00.985765Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"rotation_matrices = train_labels['rotation_matrix'].apply(string_to_matrix)\ntranslation_vectors = train_labels['translation_vector'].apply(string_to_vector)\n\n# Calculate rotation angles (in degrees) for all images\nrotation_angles = [np.rad2deg(np.arccos((np.trace(R) - 1) / 2)) for R in rotation_matrices]","metadata":{"execution":{"iopub.status.busy":"2023-05-02T08:14:01.588361Z","iopub.execute_input":"2023-05-02T08:14:01.588775Z","iopub.status.idle":"2023-05-02T08:14:01.604224Z","shell.execute_reply.started":"2023-05-02T08:14:01.588739Z","shell.execute_reply":"2023-05-02T08:14:01.603149Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure()\nplt.hist(rotation_angles, bins=50)\nplt.xlabel('Rotation Angle (degrees)')\nplt.ylabel('Number of Images')\nplt.title('Distribution of Rotation Angles')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-05-02T07:56:45.513021Z","iopub.execute_input":"2023-05-02T07:56:45.513606Z","iopub.status.idle":"2023-05-02T07:56:45.868894Z","shell.execute_reply.started":"2023-05-02T07:56:45.513563Z","shell.execute_reply":"2023-05-02T07:56:45.868004Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The distribution of rotation angles and translation distances in the dataset may seem random and somewhat uniform at first glance. The large number of values close to 0 in the first bin suggests that many camera poses have small rotations or translations relative to each other. This could be due to the nature of the images and the way they are captured around the scenes.\n\nHowever, the seemingly random distribution of the rest of the values is not surprising, given the problem's nature. In the 3D reconstruction task, camera poses are not expected to follow a specific pattern or distribution, as the images can be captured from various angles and distances around the scenes. This randomness reflects the real-world nature of image capture, where photographers may take pictures from a wide range of perspectives and positions.","metadata":{}},{"cell_type":"code","source":"# Calculate the magnitude of the translation vectors for all images\ntranslation_magnitudes = [np.linalg.norm(tvec) for tvec in translation_vectors]\n\nplt.figure()\nplt.hist(translation_magnitudes, bins=50)\nplt.xlabel('Translation Magnitude (meters)')\nplt.ylabel('Number of Images')\nplt.title('Distribution of Translation Magnitudes')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-05-02T07:56:45.870202Z","iopub.execute_input":"2023-05-02T07:56:45.870753Z","iopub.status.idle":"2023-05-02T07:56:46.209409Z","shell.execute_reply.started":"2023-05-02T07:56:45.870720Z","shell.execute_reply":"2023-05-02T07:56:46.208197Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Analyze image resolution and aspect ratio\nimage_resolutions = [Image.open(os.path.join(train_folder, path)).size for path in train_labels['image_path']]\nimage_widths, image_heights = zip(*image_resolutions)\naspect_ratios = [w / h for w, h in image_resolutions]","metadata":{"execution":{"iopub.status.busy":"2023-05-02T07:56:46.210875Z","iopub.execute_input":"2023-05-02T07:56:46.211217Z","iopub.status.idle":"2023-05-02T07:56:52.070269Z","shell.execute_reply.started":"2023-05-02T07:56:46.211187Z","shell.execute_reply":"2023-05-02T07:56:52.069126Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure()\nplt.hist(aspect_ratios, bins=50)\nplt.xlabel('Aspect Ratio')\nplt.ylabel('Number of Images')\nplt.title('Distribution of Image Aspect Ratios')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-05-02T07:56:52.071640Z","iopub.execute_input":"2023-05-02T07:56:52.072101Z","iopub.status.idle":"2023-05-02T07:56:52.411707Z","shell.execute_reply.started":"2023-05-02T07:56:52.072059Z","shell.execute_reply":"2023-05-02T07:56:52.410655Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure()\nplt.hist2d(image_widths, image_heights, bins=50)\nplt.xlabel('Image Width')\nplt.ylabel('Image Height')\nplt.title('Distribution of Image Resolutions')\nplt.colorbar(label='Number of Images')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-05-02T07:56:52.413220Z","iopub.execute_input":"2023-05-02T07:56:52.413657Z","iopub.status.idle":"2023-05-02T07:56:53.103891Z","shell.execute_reply.started":"2023-05-02T07:56:52.413616Z","shell.execute_reply":"2023-05-02T07:56:53.102547Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Calculate Relative Pose Errors\n\nFinally, we want to calculate relative pose errors, since it's essential for understanding the impact of translation errors on the model's predictions.","metadata":{}},{"cell_type":"code","source":"# Calculate relative pose errors\ndef compute_relative_pose_errors(train_labels):\n    pose_errors = []\n\n    for i, row in train_labels.iterrows():\n        # Ground truth rotation matrix and translation vector\n        gt_rotation_matrix = np.array(row['rotation_matrix'].split(';')).astype(float).reshape(3, 3)\n        gt_translation_vector = np.array(row['translation_vector'].split(';')).astype(float)\n\n        # Simulated predicted rotation matrix and translation vector (in practice, you'll use your model's predictions)\n        pred_rotation_matrix = gt_rotation_matrix  # Replace with your model's predicted rotation matrix\n        pred_translation_vector = gt_translation_vector  # Replace with your model's predicted translation vector\n\n        # Compute relative pose error\n        trace_value = np.clip((np.trace(gt_rotation_matrix.T @ pred_rotation_matrix) - 1) / 2, -1, 1)\n        rotation_error = np.rad2deg(np.arccos(trace_value))\n        translation_error = np.linalg.norm(gt_translation_vector - pred_translation_vector)\n\n        pose_errors.append((rotation_error, translation_error))\n\n    return pose_errors\n\npose_errors = compute_relative_pose_errors(train_labels)","metadata":{"execution":{"iopub.status.busy":"2023-05-02T07:59:44.236607Z","iopub.execute_input":"2023-05-02T07:59:44.237592Z","iopub.status.idle":"2023-05-02T07:59:44.297778Z","shell.execute_reply.started":"2023-05-02T07:59:44.237549Z","shell.execute_reply":"2023-05-02T07:59:44.296929Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Visualize the distribution of relative pose errors\nfig, (ax1, ax2) = plt.subplots(1, 2, figsize=(15, 5))\n\nrotation_errors = [error[0] for error in pose_errors]\ntranslation_errors = [error[1] for error in pose_errors]\n\nax1.hist(rotation_errors, bins=50)\nax1.set_xlabel('Rotation Error (degrees)')\nax1.set_ylabel('Frequency')\nax1.set_title('Distribution of Rotation Errors')\n\nax2.hist(translation_errors, bins=50)\nax2.set_xlabel('Translation Error (meters)')\nax2.set_ylabel('Frequency')\nax2.set_title('Distribution of Translation Errors')\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-05-02T08:00:32.590831Z","iopub.execute_input":"2023-05-02T08:00:32.591231Z","iopub.status.idle":"2023-05-02T08:00:33.164717Z","shell.execute_reply.started":"2023-05-02T08:00:32.591196Z","shell.execute_reply":"2023-05-02T08:00:33.163639Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the distribution of Rotation Errors and Translation Errors values, it appears that there are mostly zero errors with a few non-zero values. However, this information alone is not sufficient to draw conclusions. Comparing these errors to the prediction errors and visualizing the distributions will provide more insights into their impact on the model's performance.","metadata":{}}]}