{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Import necessary Libraries","metadata":{}},{"cell_type":"code","source":"import subprocess\n\ntry:\n    import pandas as pd\n    import numpy as np\n    import matplotlib.pyplot as plt\n    import seaborn as sns\n\n    # Code to execute using the imported modules\n    pass\n\nexcept ImportError as e:\n    print(\"Failed to import a required module: \", e)\n    \n    # Prompt user to install the missing module\n    module_name = str(e).split()[-1]\n    print(\"Attempting to install the missing module: \", module_name)\n    subprocess.check_call(['pip', 'install', module_name])\n    \nexcept Exception as e:\n    print(\"An unexpected error occurred: \", e)\n    \nfinally:\n    # Code to execute regardless of whether an exception was raised\n    pass","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-05-03T08:13:49.829381Z","iopub.execute_input":"2023-05-03T08:13:49.829916Z","iopub.status.idle":"2023-05-03T08:13:51.184390Z","shell.execute_reply.started":"2023-05-03T08:13:49.829883Z","shell.execute_reply":"2023-05-03T08:13:51.183594Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"📝 This code reads a CSV file called train_labels.csv located in a folder named train in the Kaggle input directory 📂 /kaggle/input/image-matching-challenge-2023/train/ and stores it in a pandas dataframe called train_labels. 🐼\n\n👀 It then displays the first few rows of the train_labels dataframe using the head() method. 👓\n\n📊 The code then displays some information about the train_labels dataframe using the info() method, including the number of rows, columns, data types, and the number of non-null values in each column. 📈\n\n🔍 Additionally, it checks for any missing values in the train_labels dataframe using the isnull() and sum() methods. 🕵️‍♀️\n\n🔢 The code calculates the number of unique datasets and scenes in the train_labels dataframe using the nunique() method and prints these values to the console. 📉","metadata":{}},{"cell_type":"code","source":"train_labels = pd.read_csv(\"/kaggle/input/image-matching-challenge-2023/train/train_labels.csv\")\n\n# Check the first few rows of the dataset\nprint(\"First few rows of train_labels:\")\nprint(train_labels.head())\n\n# Get dataset information\nprint(\"\\nInformation about train_labels dataset:\")\nprint(train_labels.info())\n\n# Check for missing values in the dataset\nprint(\"\\nMissing values in train_labels dataset:\")\nprint(train_labels.isnull().sum())\n\n# Get unique counts of datasets and scenes\nnum_datasets = train_labels['dataset'].nunique()\nnum_scenes = train_labels['scene'].nunique()\nprint(\"\\nNumber of unique datasets:\", num_datasets)\nprint(\"Number of unique scenes:\", num_scenes)\n","metadata":{"execution":{"iopub.status.busy":"2023-05-03T08:14:39.320026Z","iopub.execute_input":"2023-05-03T08:14:39.320406Z","iopub.status.idle":"2023-05-03T08:14:39.376308Z","shell.execute_reply.started":"2023-05-03T08:14:39.320381Z","shell.execute_reply":"2023-05-03T08:14:39.374762Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"📊 The first block of code creates a count plot using Seaborn library to show the distribution of images per dataset. It uses the countplot() function from the Seaborn library, which takes in the train_labels dataframe and the column name dataset to create the plot. 📈\n\n📝 It also sets the title of the plot using the title() function from Matplotlib library and displays the plot using the show() function from Matplotlib library. 📈\n\n📊 The second block of code creates another count plot using Seaborn library to show the distribution of images per scene. It uses the countplot() function from the Seaborn library, which takes in the train_labels dataframe and the column name scene to create the plot. 📈\n\n👀 It also sets the figure size using the figure() function from Matplotlib library and sets the order of the bars in the plot using the value_counts() and index functions. It sets the title of the plot using the title() function from Matplotlib library, rotates the x-axis labels using the xticks() function from Matplotlib library, and displays the plot using the show() function from Matplotlib library. 📈","metadata":{}},{"cell_type":"code","source":"# Distribution of images per dataset\nsns.countplot(data=train_labels, x='dataset')\nplt.title(\"Images per Dataset\")\nplt.show()\n\n# Distribution of images per scene\nplt.figure(figsize=(12, 6))\nsns.countplot(data=train_labels, x='scene', order=train_labels['scene'].value_counts().index)\nplt.title(\"Images per Scene\")\nplt.xticks(rotation=90)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-05-03T08:16:54.272114Z","iopub.execute_input":"2023-05-03T08:16:54.272555Z","iopub.status.idle":"2023-05-03T08:16:54.735244Z","shell.execute_reply.started":"2023-05-03T08:16:54.272527Z","shell.execute_reply":"2023-05-03T08:16:54.733938Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"🖼️ The code first imports the Image module from the PIL (Python Imaging Library) library and the os module from the standard library. This is to work with images and directories. 📂\n\n🖼️ It then defines a function called display_images that takes in a list of image paths, number of rows and columns, and a figure size as inputs. This function creates a grid of subplots using the subplots() function from the Matplotlib library and iterates over each subplot. It opens the image using the Image.open() function from the PIL library, sets the title of the subplot to the base name of the image using the os.path.basename() function, and displays the image using the imshow() function. Finally, it hides the axes of the subplot using the axis() function from Matplotlib library and displays the plot using the show() function from the Matplotlib library. 📈\n\n🖼️ The code then defines a variable called base_dir that stores the base directory of the image files. It then creates a list of example image paths by concatenating the base directory with the image_path column from the train_labels dataframe for the first 8 images. 📂\n\n🖼️ Finally, it calls the display_images function with the example image paths list as input to display the images in a grid format. 📈","metadata":{}},{"cell_type":"code","source":"from PIL import Image\nimport os\n\n# Function to display a grid of images\ndef display_images(image_paths, nrows=2, ncols=4, figsize=(16, 8)):\n    fig, axes = plt.subplots(nrows, ncols, figsize=figsize)\n    for i, ax in enumerate(axes.flatten()):\n        if i < len(image_paths):\n            img = Image.open(image_paths[i])\n            ax.imshow(img)\n            ax.set_title(os.path.basename(image_paths[i]))\n        ax.axis('off')\n    plt.show()\n\n# Add the base directory to the image paths\nbase_dir = \"/kaggle/input/image-matching-challenge-2023/train/\"\nexample_images = [base_dir + img_path for img_path in train_labels['image_path'][:8].tolist()]\n\n# Display example images\ndisplay_images(example_images)","metadata":{"execution":{"iopub.status.busy":"2023-05-03T08:16:57.504364Z","iopub.execute_input":"2023-05-03T08:16:57.504801Z","iopub.status.idle":"2023-05-03T08:17:00.499548Z","shell.execute_reply.started":"2023-05-03T08:16:57.504769Z","shell.execute_reply":"2023-05-03T08:17:00.497573Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"🔢 The code first imports the NearestNeighbors class from the Scikit-learn library and the Rotation class from the Scipy library. This is to perform nearest neighbor search and work with rotation and translation vectors. 🔍\n\n🔢 It then defines a function called parse_vectors that takes in a pandas dataframe as input and returns two numpy arrays of rotation matrices and translation vectors. This function iterates over each row in the dataframe, converts the rotation matrix and translation vector strings to numpy arrays, and stores them in separate lists. It then returns these lists as numpy arrays. 🔍\n\n🔢 The code then calls the parse_vectors function on the train_labels dataframe to get the rotation matrices and translation vectors. It then calculates the rotation angles using the as_euler() method of the Rotation class from the Scipy library and the translation magnitudes using the norm() method of the numpy library. 📐\n\n📊 The next block of code creates a histogram using Seaborn library to show the distribution of rotation angles in the dataset. It uses the histplot() function from the Seaborn library, which takes in the flattened rotation angles numpy array and the number of bins to create the plot. 📈\n\n📝 It also sets the title, x-label, and y-label of the plot using the title(), xlabel(), and ylabel() functions from Matplotlib library, respectively, and displays the plot using the show() function from Matplotlib library. 📈\n\n📊 The final block of code creates another histogram using Seaborn library to show the distribution of translation magnitudes in the dataset. It uses the histplot() function from the Seaborn library, which takes in the translation magnitudes numpy array and the number of bins to create the plot. 📈\n\n👀 It also sets the title, x-label, and y-label of the plot using the title(), xlabel(), and ylabel() functions from Matplotlib library, respectively, and displays the plot using the show() function from Matplotlib library. 📈","metadata":{}},{"cell_type":"code","source":"from sklearn.neighbors import NearestNeighbors\nfrom scipy.spatial.transform import Rotation\n\n# Function to parse rotation_matrix and translation_vector\ndef parse_vectors(df):\n    rotations = []\n    translations = []\n    for index, row in df.iterrows():\n        rotations.append(np.array(row['rotation_matrix'].split(';'), dtype=float).reshape(3, 3))\n        translations.append(np.array(row['translation_vector'].split(';'), dtype=float))\n    return np.array(rotations), np.array(translations)\n\nrotations, translations = parse_vectors(train_labels)\n\n# Calculate rotation angles and translation magnitudes\nrotation_angles = np.degrees([Rotation.from_matrix(r).as_euler('xyz') for r in rotations])\ntranslation_magnitudes = np.linalg.norm(translations, axis=1)\n\n# Distribution of rotation angles\nplt.figure(figsize=(12, 6))\nsns.histplot(rotation_angles.flatten(), kde=False, bins=50)\nplt.title(\"Distribution of Rotation Angles\")\nplt.xlabel(\"Rotation Angle (Degrees)\")\nplt.ylabel(\"Frequency\")\nplt.show()\n\n# Distribution of translation magnitudes\nplt.figure(figsize=(12, 6))\nsns.histplot(translation_magnitudes, kde=False, bins=50)\nplt.title(\"Distribution of Translation Magnitudes\")\nplt.xlabel(\"Translation Magnitude\")\nplt.ylabel(\"Frequency\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-05-03T08:17:32.611472Z","iopub.execute_input":"2023-05-03T08:17:32.611951Z","iopub.status.idle":"2023-05-03T08:17:33.787085Z","shell.execute_reply.started":"2023-05-03T08:17:32.611919Z","shell.execute_reply":"2023-05-03T08:17:33.786205Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"📊 The code creates a scatter plot using Matplotlib library to show the relationship between rotation angles and translation magnitudes in the dataset. It uses the scatter() function from the Matplotlib library, which takes in the first column of the rotation angles numpy array, the translation magnitudes numpy array, alpha value for transparency, and a label for each rotation axis to create the plot. 📈\n\n📝 It also sets the x-label, y-label, title, and legend of the plot using the xlabel(), ylabel(), title(), and legend() functions from Matplotlib library, respectively, and displays the plot using the show() function from Matplotlib library. 📈\n\nThe scatter plot helps us visualize the relationship between rotation angles and translation magnitudes for different rotation axes.","metadata":{}},{"cell_type":"code","source":"# Scatter plot of rotation angles vs translation magnitudes\nplt.figure(figsize=(12, 8))\nplt.scatter(rotation_angles[:, 0], translation_magnitudes, alpha=0.5, label='Rotation around X-axis')\nplt.scatter(rotation_angles[:, 1], translation_magnitudes, alpha=0.5, label='Rotation around Y-axis')\nplt.scatter(rotation_angles[:, 2], translation_magnitudes, alpha=0.5, label='Rotation around Z-axis')\nplt.xlabel(\"Rotation Angle (Degrees)\")\nplt.ylabel(\"Translation Magnitude\")\nplt.legend()\nplt.title(\"Relationship between Rotation Angles and Translation Magnitudes\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-05-03T03:15:18.935454Z","iopub.execute_input":"2023-05-03T03:15:18.936168Z","iopub.status.idle":"2023-05-03T03:15:19.338230Z","shell.execute_reply.started":"2023-05-03T03:15:18.936129Z","shell.execute_reply":"2023-05-03T03:15:19.337071Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"📊 The code defines a function called plot_camera_poses that takes in a numpy array of translation vectors and a title as inputs. This function creates a 3D scatter plot using the scatter() function from the Matplotlib library and the projection='3d' argument to set the plot to a 3D view. It then sets the x-label, y-label, z-label, and title of the plot using the set_xlabel(), set_ylabel(), set_zlabel(), and title() functions from Matplotlib library, respectively, and displays the plot using the show() function from Matplotlib library. 📈\n\n📊 The code then calls the plot_camera_poses function on the translation vectors numpy array and a title to plot the distribution of camera poses in 3D space. 📈\n\nThe 3D scatter plot helps us visualize the distribution of camera poses in the dataset.","metadata":{}},{"cell_type":"code","source":"# Function to plot 3D camera poses\ndef plot_camera_poses(translations, title):\n    fig = plt.figure(figsize=(12, 8))\n    ax = fig.add_subplot(111, projection='3d')\n    ax.scatter(translations[:, 0], translations[:, 1], translations[:, 2], alpha=0.5)\n    ax.set_xlabel('X')\n    ax.set_ylabel('Y')\n    ax.set_zlabel('Z')\n    plt.title(title)\n    plt.show()\n\n# Plot camera poses in 3D space\nplot_camera_poses(translations, \"Distribution of Camera Poses in 3D Space\")","metadata":{"execution":{"iopub.status.busy":"2023-05-03T08:18:06.571840Z","iopub.execute_input":"2023-05-03T08:18:06.572708Z","iopub.status.idle":"2023-05-03T08:18:06.873519Z","shell.execute_reply.started":"2023-05-03T08:18:06.572668Z","shell.execute_reply":"2023-05-03T08:18:06.870052Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"📊 The code creates a set of three histograms using Seaborn library to show the distribution of rotation angles for each axis in the dataset. It uses the histplot() function from the Seaborn library and creates three subplots using the subplot() function from Matplotlib library. 📈\n\n📝 It sets the title, x-label, and y-label of each subplot using the set_title(), set_xlabel(), and set_ylabel() functions from Matplotlib library, respectively, and displays the plots using the show() function from Matplotlib library. 📈\n\nThe set of histograms helps us visualize the distribution of rotation angles for each axis separately in the dataset.","metadata":{}},{"cell_type":"code","source":"# Histograms of rotation angles for each axis\nplt.figure(figsize=(16, 4))\nax1 = plt.subplot(1, 3, 1)\nsns.histplot(rotation_angles[:, 0], kde=False, bins=50, ax=ax1)\nax1.set_title(\"Distribution of Rotation Angles (X-axis)\")\nax1.set_xlabel(\"Rotation Angle (Degrees)\")\nax1.set_ylabel(\"Frequency\")\n\nax2 = plt.subplot(1, 3, 2)\nsns.histplot(rotation_angles[:, 1], kde=False, bins=50, ax=ax2)\nax2.set_title(\"Distribution of Rotation Angles (Y-axis)\")\nax2.set_xlabel(\"Rotation Angle (Degrees)\")\n\nax3 = plt.subplot(1, 3, 3)\nsns.histplot(rotation_angles[:, 2], kde=False, bins=50, ax=ax3)\nax3.set_title(\"Distribution of Rotation Angles (Z-axis)\")\nax3.set_xlabel(\"Rotation Angle (Degrees)\")\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-05-03T08:18:32.355163Z","iopub.execute_input":"2023-05-03T08:18:32.355579Z","iopub.status.idle":"2023-05-03T08:18:33.161082Z","shell.execute_reply.started":"2023-05-03T08:18:32.355548Z","shell.execute_reply":"2023-05-03T08:18:33.160000Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os\nimport random\nimport subprocess\nimport matplotlib.pyplot as plt\nfrom mpl_toolkits.mplot3d import Axes3D\nfrom PIL import Image","metadata":{"execution":{"iopub.status.busy":"2023-05-03T08:18:45.919681Z","iopub.execute_input":"2023-05-03T08:18:45.920162Z","iopub.status.idle":"2023-05-03T08:18:45.927096Z","shell.execute_reply.started":"2023-05-03T08:18:45.920126Z","shell.execute_reply":"2023-05-03T08:18:45.926189Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"📂 The code sets the base_dir variable to the base directory of the image files. 📂\n\n🔀 The code randomly selects 8 image paths from the train_labels dataframe using the sample() method of the random module and stores them in a variable called example_images. It then concatenates the base directory and the selected image paths using a list comprehension and assigns it back to example_images. 🔀\n\n🖼️ Finally, the code calls the display_images function with the example_images list as input to display the randomly selected example images in a grid format. 📈","metadata":{}},{"cell_type":"code","source":"# Set the base directory\nbase_dir = \"/kaggle/input/image-matching-challenge-2023/train/\"\n\n# Display example images\nexample_images = random.sample(train_labels['image_path'].tolist(), 8)\nexample_images = [base_dir + img_path for img_path in example_images]\ndisplay_images(example_images)","metadata":{"execution":{"iopub.status.busy":"2023-05-03T08:18:47.126878Z","iopub.execute_input":"2023-05-03T08:18:47.127247Z","iopub.status.idle":"2023-05-03T08:18:52.685667Z","shell.execute_reply.started":"2023-05-03T08:18:47.127218Z","shell.execute_reply":"2023-05-03T08:18:52.684485Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"🔀 The code first imports the random module from the standard library. This is to generate random numbers for selecting scenes and images. 🔀\n\n📂 It then sets the base_dir variable to the base directory of the image files. 📂\n\n📋 The code gets the unique scenes from the train_labels dataframe using the unique() method and stores them in a variable called unique_scenes. It then sets the number of scenes to display images from using the num_scenes variable. 📋\n\n🔀 The code randomly selects scenes from the unique_scenes list using the sample() method of the random module and stores them in a variable called selected_scenes. It then creates an empty list called example_images. 🔀\n\n📋 The code then iterates over each selected scene, gets all the image paths for that scene from the train_labels dataframe using boolean indexing and the tolist() method, and selects a random image path from the list using the choice() method of the random module. It concatenates the base directory and the selected image path and appends it to the example_images list. 📋\n\n🖼️ Finally, the code calls the display_images function with the example_images list as input, along with the number of rows, columns, and figure size to display the randomly selected example images from different scenes in a grid format. 📈","metadata":{}},{"cell_type":"code","source":"import random\n\n# Set the base directory\nbase_dir = \"/kaggle/input/image-matching-challenge-2023/train/\"\n\n# Get unique scenes\nunique_scenes = train_labels[\"scene\"].unique()\n\n# Choose a number of scenes to display images from\nnum_scenes = 6\n\n# Randomly select scenes\nselected_scenes = random.sample(list(unique_scenes), num_scenes)\n\n# Get one random image from each selected scene\nexample_images = []\nfor scene in selected_scenes:\n    scene_images = train_labels[train_labels[\"scene\"] == scene][\"image_path\"].tolist()\n    example_image = random.choice(scene_images)\n    example_images.append(base_dir + example_image)\n\n# Display example images from different scenes\ndisplay_images(example_images, nrows=2, ncols=3, figsize=(16, 8))","metadata":{"execution":{"iopub.status.busy":"2023-05-03T08:18:52.687886Z","iopub.execute_input":"2023-05-03T08:18:52.688372Z","iopub.status.idle":"2023-05-03T08:19:02.270679Z","shell.execute_reply.started":"2023-05-03T08:18:52.688296Z","shell.execute_reply":"2023-05-03T08:19:02.268846Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"📂 The code defines a function called print_directory_structure that takes in a starting directory as input. This function uses the os.walk() function to traverse the directory tree and print the directory structure. It sets the level variable to the number of directory levels below the starting directory and the indent variable to four spaces multiplied by the level variable to indent the output. It then prints the name of the directory using the os.path.basename() function and adds a forward slash at the end to denote it's a directory. 📂\n\n📂 The code sets the base_dir variable to the base directory of the image files. It then joins the base_dir and the \"train\" subdirectory using the os.path.join() function and stores it in a variable called train_folder. 📂\n\n📝 The code then calls the print_directory_structure function on the train_folder to print the directory structure of the \"train\" subdirectory. 📝\n\nThe print_directory_structure function helps us visualize the directory structure of the files and folders in a tree-like format.","metadata":{}},{"cell_type":"code","source":"import os\n\ndef print_directory_structure(start_directory):\n    for root, dirs, files in os.walk(start_directory):\n        level = root.replace(start_directory, '').count(os.sep)\n        indent = ' ' * 4 * level\n        print(f\"{indent}{os.path.basename(root)}/\")\n\n# Set the base directory\nbase_dir = \"/kaggle/input/image-matching-challenge-2023/\"\n\n# Print the train folder structure\ntrain_folder = os.path.join(base_dir, \"train\")\nprint_directory_structure(train_folder)","metadata":{"execution":{"iopub.status.busy":"2023-05-03T08:19:23.379273Z","iopub.execute_input":"2023-05-03T08:19:23.379721Z","iopub.status.idle":"2023-05-03T08:19:26.128946Z","shell.execute_reply.started":"2023-05-03T08:19:23.379688Z","shell.execute_reply":"2023-05-03T08:19:26.128065Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"🔀 The code first imports the random module from the standard library. This is to generate random numbers for selecting datasets, scenes, and images. 🔀\n\n📂 It then sets the base_dir variable to the base directory of the image files. It then gets the unique datasets and scenes from the train_labels dataframe using the unique() method and stores them in variables called unique_datasets and unique_scenes, respectively. 📂\n\n📋 The code sets the number of datasets and scenes to display images from using the num_datasets and num_scenes variables. It then creates two empty lists called selected_datasets and selected_scenes. 📋\n\n🔀 The code randomly selects datasets and scenes from the unique_datasets and unique_scenes lists using the sample() method of the random module and stores them in selected_datasets and selected_scenes, respectively. 🔀\n\n📋 The code then iterates over each selected dataset and scene, gets all the image paths for that dataset and scene from the train_labels dataframe using boolean indexing and the tolist() method, and selects a random image path from the list using the choice() method of the random module. If the scene_images list is empty, it skips that iteration. It concatenates the base directory and the selected image path and appends it to the example_images list. 📋\n\n🖼️ Finally, the code calls the display_images function with the example_images list as input, along with the number of rows, columns, and figure size to display the randomly selected example images from different scenes in different datasets in a grid format. The nrows and ncols arguments of the display_images function are set to num_datasets and num_scenes to display the images in multiple rows and columns. 📈\n\nThe code helps us visualize random example images from different scenes in different datasets.","metadata":{}},{"cell_type":"code","source":"import random\n\n# Set the base directory\nbase_dir = \"/kaggle/input/image-matching-challenge-2023/train/\"\n\n# Get unique datasets and scenes\nunique_datasets = train_labels[\"dataset\"].unique()\nunique_scenes = train_labels[\"scene\"].unique()\n\n# Choose a number of datasets and scenes to display images from\nnum_datasets = 3\nnum_scenes = 2\n\n# Randomly select datasets and scenes\nselected_datasets = random.sample(list(unique_datasets), num_datasets)\nselected_scenes = random.sample(list(unique_scenes), num_scenes)\n\n# Get one random image from each selected scene in each dataset\nexample_images = []\nfor dataset in selected_datasets:\n    for scene in selected_scenes:\n        scene_images = train_labels[(train_labels[\"dataset\"] == dataset) & (train_labels[\"scene\"] == scene)][\"image_path\"].tolist()\n        if scene_images:\n            example_image = random.choice(scene_images)\n            example_images.append(base_dir + example_image)\n\n# Display example images from different scenes in different datasets\ndisplay_images(example_images, nrows=num_datasets, ncols=num_scenes, figsize=(16, 8))","metadata":{"execution":{"iopub.status.busy":"2023-05-03T08:20:00.912456Z","iopub.execute_input":"2023-05-03T08:20:00.912926Z","iopub.status.idle":"2023-05-03T08:20:04.547515Z","shell.execute_reply.started":"2023-05-03T08:20:00.912892Z","shell.execute_reply":"2023-05-03T08:20:04.546676Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"📂 The code first imports the necessary modules - os, glob, PIL, and matplotlib.pyplot. The os module is used for file and directory operations, glob is used for pattern matching of file paths, PIL is used to open and manipulate image files, and matplotlib.pyplot is used for displaying the images. 📂\n\n🖼️ The code defines a function called display_sample_images_from_scenes that takes in a train directory as input. It first creates a list called category_folders that contains the full paths of all the subdirectories in the train directory using list comprehension and the os.listdir() and os.path.join() functions. It then iterates over each category_folder in category_folders and creates another list called scene_folders that contains the full paths of all the subdirectories in the category_folder. 🖼️\n\n🖼️ The code then iterates over each scene_folder in scene_folders and gets the full path of the images subdirectory using the os.path.join() function. It then selects the first image file path in the images subdirectory using the glob.glob() function and stores it in a variable called sample_image_path. 🖼️\n\n🖼️ The code opens the sample_image_path using the Image.open() function of the PIL module and stores it in a variable called img. It then displays the image using the imshow() and title() functions of the matplotlib.pyplot module. The title of the image is set to the name of the category and scene folders using the os.path.basename() function. It turns off the axes of the image using the axis() function and finally shows the image using the show() function. 🖼️\n\n📂 The code sets the base_dir variable to the base directory of the image files. It then joins the base_dir and the \"train\" subdirectory using the os.path.join() function and stores it in a variable called train_folder. 📂\n\n🖼️ The code calls the display_sample_images_from_scenes function with the train_folder as input to display sample images from all scenes in the train folder. 🖼️\n\nThis code helps us display sample images from all scenes in the train folder to get a better understanding of the data.","metadata":{}},{"cell_type":"code","source":"import os\nimport glob\nfrom PIL import Image\nimport matplotlib.pyplot as plt\n\ndef display_sample_images_from_scenes(train_directory):\n    category_folders = [os.path.join(train_directory, category) for category in os.listdir(train_directory) if os.path.isdir(os.path.join(train_directory, category))]\n    \n    for category_folder in category_folders:\n        scene_folders = [os.path.join(category_folder, scene) for scene in os.listdir(category_folder) if os.path.isdir(os.path.join(category_folder, scene))]\n        \n        for scene_folder in scene_folders:\n            image_folder = os.path.join(scene_folder, \"images\")\n            sample_image_path = glob.glob(os.path.join(image_folder, \"*\"))[0]\n            \n            img = Image.open(sample_image_path)\n            plt.imshow(img)\n            plt.title(f\"{os.path.basename(category_folder)} - {os.path.basename(scene_folder)}\")\n            plt.axis('off')\n            plt.show()\n\n# Set the base directory\nbase_dir = \"/kaggle/input/image-matching-challenge-2023/\"\n\n# Display sample images from all scenes in the train folder\ntrain_folder = os.path.join(base_dir, \"train\")\ndisplay_sample_images_from_scenes(train_folder)","metadata":{"execution":{"iopub.status.busy":"2023-05-03T08:20:07.342493Z","iopub.execute_input":"2023-05-03T08:20:07.343242Z","iopub.status.idle":"2023-05-03T08:20:25.091538Z","shell.execute_reply.started":"2023-05-03T08:20:07.343203Z","shell.execute_reply":"2023-05-03T08:20:25.090418Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}