{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":84969,"databundleVersionId":10033515,"sourceType":"competition"}],"dockerImageVersionId":30786,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Cryo-ET\n\nCryo-electron tomography (Cryo-ET) is a powerful and advanced imaging method that combines cryogenic techniques with electron microscopy to visualize and reconstruct three-dimensional (3D) structures of biological samples in a near-native, fully hydrated state. Cryo-ET has become a key tool for studying the architecture of cells, organelles, protein complexes, and other macromolecular assemblies at molecular resolution within their natural cellular context, allowing scientists to capture detailed snapshots of biological processes as they occur in living organisms.\n\nThe Cryo-ET process begins with vitrification, in which biological samples—often cells or thin tissue sections—are rapidly frozen by plunging them into liquid ethane at cryogenic temperatures. This process bypasses the formation of damaging ice crystals, instead trapping the sample in an amorphous, glass-like state that preserves molecular structures with minimal alteration. Once the sample is vitrified, it is transferred to an electron microscope, where it remains frozen at temperatures near -180°C throughout the imaging process to prevent damage from the electron beam.\n\nIn the microscope, the sample is tilted incrementally, often by just a few degrees at a time, capturing a series of two-dimensional (2D) images or \"tilt series\" from multiple perspectives. These 2D images are then computationally combined using sophisticated algorithms to generate a 3D tomographic reconstruction. This 3D model reveals the spatial organization and ultrastructure of cellular components, including organelles, protein complexes, and cytoskeletal elements, within their native cellular environment.\n\nOne of the unique strengths of Cryo-ET is its ability to visualize samples in situ, providing contextual information on how individual macromolecules interact and organize within cells. Unlike traditional electron microscopy, which often requires extensive sample preparation and staining that can alter cellular structures, Cryo-ET preserves the sample’s natural state, enabling more accurate structural analysis. This capability is invaluable in fields such as cell biology, microbiology, and structural biology, where understanding the interplay between different cellular components is essential for uncovering the mechanisms behind processes like viral infection, cell signaling, protein synthesis, and energy metabolism.\n\nAdditionally, recent advancements in Cryo-ET, including the integration of direct electron detectors and enhanced image processing techniques, have dramatically improved resolution and data quality. These developments allow scientists to resolve finer structural details, down to the level of individual proteins and macromolecules within the cell. Emerging applications of Cryo-ET include studying complex biological assemblies that were previously challenging to analyze, such as ribosomes, molecular motors, and the architecture of the cytoskeleton, in high resolution and in their native context.\n\nOverall, Cryo-ET provides an unparalleled glimpse into the inner workings of cells and macromolecular structures, making it a transformative tool in modern biology and a vital technique for advancing our understanding of cellular and molecular functions in health and disease.","metadata":{}},{"cell_type":"code","source":"#Bdellovibrio bacteriovorus\nimport matplotlib.pyplot as plt\nimport matplotlib.image as mpimg\nimport requests\nfrom io import BytesIO\n\nurl = 'https://upload.wikimedia.org/wikipedia/commons/3/33/Slice_from_electron_cryotomogram_of_Bdellovibrio_bacteriovorus_cell.jpg'\n\nheaders = {\n    \"User-Agent\": \"Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/87.0.4280.88 Safari/537.36\"\n}\n\ntry:\n    response = requests.get(url, headers=headers)\n    response.raise_for_status() \n\n    img = mpimg.imread(BytesIO(response.content), format='jpg')\n    \n    plt.imshow(img)\n    plt.axis('off') \n    plt.show()\nexcept requests.exceptions.RequestException as e:\n    print(\"Error:\", e)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-09T18:55:46.819189Z","iopub.execute_input":"2024-11-09T18:55:46.820473Z","iopub.status.idle":"2024-11-09T18:55:47.394125Z","shell.execute_reply.started":"2024-11-09T18:55:46.820413Z","shell.execute_reply":"2024-11-09T18:55:47.392796Z"},"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"# Importing modules","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport os\n!pip install zarr -q\nimport zarr\nimport matplotlib.pyplot as plt\nimport plotly.graph_objects as go\nimport imageio.v2 as imageio\nimport time\nimport math\nimport json\nimport base64\nfrom IPython.display import Image, HTML","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2024-11-12T18:05:22.780573Z","iopub.execute_input":"2024-11-12T18:05:22.781070Z","iopub.status.idle":"2024-11-12T18:05:37.045397Z","shell.execute_reply.started":"2024-11-12T18:05:22.781021Z","shell.execute_reply":"2024-11-12T18:05:37.043659Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"# Define Paths\n\n#### First, we need to understand which files are included in this competition. To do this, we’ll modify a built-in function in the notebook that outputs file paths, so that it creates an array of paths instead.","metadata":{}},{"cell_type":"code","source":"paths = []\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        paths.append(os.path.join(dirname, filename))\n\nprint(paths[0])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-12T18:05:37.048669Z","iopub.execute_input":"2024-11-12T18:05:37.049274Z","iopub.status.idle":"2024-11-12T18:05:37.556967Z","shell.execute_reply.started":"2024-11-12T18:05:37.049209Z","shell.execute_reply":"2024-11-12T18:05:37.555591Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"# Data Retrieval\n\n#### Now that we have defined the paths to the files containing the data, we need to retrieve this data. To do so, we’ll define several functions using the `zarr` module that will enable us to accomplish this task, specifically:","metadata":{}},{"cell_type":"markdown","source":"#### `load_zarr_data` Function\n\nThis function loads a Zarr file and returns its root group, providing access to the data within the file.\n\n- **Parameters**:\n  - `zarr_path` (`str`): Path to the Zarr file.\n\n- **Returns**:\n  - `zarr.Group`: The root group of the Zarr file, which serves as the main access point to datasets and subgroups stored in the file.\n\nThe function opens the Zarr file in read-only mode (`mode='r'`) to allow reading without modifying the original file.","metadata":{}},{"cell_type":"code","source":"def load_zarr_data(zarr_path):\n    return zarr.open(zarr_path, mode='r')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-12T18:05:37.558716Z","iopub.execute_input":"2024-11-12T18:05:37.559707Z","iopub.status.idle":"2024-11-12T18:05:37.565205Z","shell.execute_reply.started":"2024-11-12T18:05:37.559648Z","shell.execute_reply":"2024-11-12T18:05:37.564032Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"#### `inspect_zarr_structure` Function\n\nThis function prints the hierarchical structure of a Zarr group, providing an overview of the organization and content of the data.\n\n- **Parameters**:\n  - `zarr_group` (`zarr.Group`): The root group of the Zarr file.\n\nThe function outputs the data structure by calling `zarr_group.tree()`, which displays the tree-like organization of datasets and subgroups within the Zarr file, making it easier to understand the file's contents and navigate the data.","metadata":{}},{"cell_type":"code","source":"def inspect_zarr_structure(zarr_group):\n    print(\"Data structure:\")\n    print(zarr_group.tree())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-12T18:05:37.567670Z","iopub.execute_input":"2024-11-12T18:05:37.568100Z","iopub.status.idle":"2024-11-12T18:05:37.579510Z","shell.execute_reply.started":"2024-11-12T18:05:37.568058Z","shell.execute_reply":"2024-11-12T18:05:37.578051Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"#### `get_data_array` Function\n\nThis function retrieves a specific data array from a given Zarr group, allowing access to individual datasets within the hierarchical structure.\n\n- **Parameters**:\n  - `zarr_group` (`zarr.Group`): The root group of the Zarr file.\n  - `array_name` (`str`): The name of the array to access, with a default value of `'0'`.\n\n- **Returns**:\n  - `zarr.Array`: The specified data array within the Zarr group.\n\n.","metadata":{}},{"cell_type":"code","source":"def get_data_array(zarr_group, array_name='0'):\n    return zarr_group[array_name]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-12T18:05:37.580998Z","iopub.execute_input":"2024-11-12T18:05:37.581413Z","iopub.status.idle":"2024-11-12T18:05:37.596520Z","shell.execute_reply.started":"2024-11-12T18:05:37.581372Z","shell.execute_reply":"2024-11-12T18:05:37.595189Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"#### `display_data_info` Function\n\nThis function provides key information about a specified data array, including its shape, data type, and any attributes, which is useful for understanding the structure and properties of the data.\n\n- **Parameters**:\n  - `data_array` (`zarr.Array`): The data array to inspect.\n\n- **Outputs**:\n  - **Shape**: Prints the shape (dimensions) of the array.\n  - **Data Type**: Prints the data type of elements within the array.\n  - **Attributes**: Prints any associated metadata or attributes stored with the array.\n\n.","metadata":{}},{"cell_type":"code","source":"def display_data_info(data_array):\n    print(\"Data shape:\", data_array.shape)\n    print(\"Data type:\", data_array.dtype)\n    print(\"Array attributes:\", data_array.attrs)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-12T18:05:37.598230Z","iopub.execute_input":"2024-11-12T18:05:37.598641Z","iopub.status.idle":"2024-11-12T18:05:37.611268Z","shell.execute_reply.started":"2024-11-12T18:05:37.598600Z","shell.execute_reply":"2024-11-12T18:05:37.609877Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"#### `display_group_attributes` Function\n\nThis function outputs the attributes of a specified Zarr group, allowing a user to view any metadata or descriptive information attached to the group.\n\n- **Parameters**:\n  - `zarr_group` (`zarr.Group`): The root or specified group of the Zarr file.\n\n- **Outputs**:\n  - **Group Attributes**: Prints the attributes (metadata) of the Zarr group.\n\n.","metadata":{}},{"cell_type":"code","source":"def display_group_attributes(zarr_group):\n    print(\"Group attributes:\", zarr_group.attrs)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-12T18:05:37.613116Z","iopub.execute_input":"2024-11-12T18:05:37.614137Z","iopub.status.idle":"2024-11-12T18:05:37.627684Z","shell.execute_reply.started":"2024-11-12T18:05:37.614093Z","shell.execute_reply":"2024-11-12T18:05:37.626280Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"#### `visualize_slice` Function\n\nThis function visualizes a 2D slice from a 3D data array, providing an easy way to inspect a specific layer of multidimensional data.\n\n- **Parameters**:\n  - `data_array` (`zarr.Array`): The 3D data array containing the data to visualize.\n  - `slice_index` (`int`): The index along the first axis specifying which 2D slice to visualize (default is `0`).\n\n- **Process**:\n  - Retrieves a specific 2D slice of data from the 3D array by selecting the slice at the specified index along the first axis.\n  - Uses `matplotlib`'s `imshow` function to display the slice in grayscale (`cmap='gray'`).\n\n- **Outputs**:\n  - A plot showing the specified slice, with a title indicating the slice index.\n \n.","metadata":{}},{"cell_type":"code","source":"def visualize_slice(data_array, slice_index=0):\n    slice_data = data_array[slice_index, :, :]\n    plt.imshow(slice_data, cmap='gray')\n    plt.title(f'Slice {slice_index} of data')\n    plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-12T18:05:37.871713Z","iopub.execute_input":"2024-11-12T18:05:37.872185Z","iopub.status.idle":"2024-11-12T18:05:37.879294Z","shell.execute_reply.started":"2024-11-12T18:05:37.872138Z","shell.execute_reply":"2024-11-12T18:05:37.878180Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"#### `visualize_all_slices_in_grid` Function\n\nThis function displays all 2D slices of a 3D data array in a single grid layout. This layout allows for easy examination of each slice, all contained within one comprehensive figure.\n\n- **Parameters**:\n  - `data_array` (zarr.Array or numpy.ndarray): A 3D data array where each slice along the first axis will be displayed in a subplot.\n  - `cols` (int, default=5): Specifies the number of columns in the grid layout, impacting the arrangement of slices across rows and columns.\n\n- **Function Workflow**:\n  1. **Grid Calculation**: The function first calculates the number of rows required for the grid, based on the number of slices and specified columns.\n  2. **Slice Plotting**: Each slice is plotted within its own subplot, organized across the calculated grid layout.\n  3. **Axis Management**: The function hides axes for a cleaner look.\n  4. **Empty Subplot Handling**: If extra subplots remain after all slices are plotted, these are removed for visual clarity.\n  5. **Display**: Finally, the function presents the entire grid of slices in a single figure.\n \n.","metadata":{}},{"cell_type":"code","source":"def visualize_all_slices_sequentially(data_array, cols=1):\n    num_slices = data_array.shape[0]\n    rows = (num_slices + cols - 1) // cols  # Round up for rows\n\n    fig, axes = plt.subplots(rows, cols, figsize=(15, rows * 3))\n    axes = axes.flatten()  # Flatten to easily iterate over\n\n    for i in range(num_slices):\n        axes[i].imshow(data_array[i, :, :], cmap='gray')\n        axes[i].set_title(f'Slice {i}')\n        axes[i].axis('off')  # Hide axis ticks\n\n    # Hide any extra subplots\n    for i in range(num_slices, len(axes)):\n        axes[i].axis('off')\n\n    plt.tight_layout()\n    plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-12T18:05:38.931884Z","iopub.execute_input":"2024-11-12T18:05:38.932358Z","iopub.status.idle":"2024-11-12T18:05:38.942490Z","shell.execute_reply.started":"2024-11-12T18:05:38.932315Z","shell.execute_reply":"2024-11-12T18:05:38.940912Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"#### `visualize_all_slices_in_square` Function\n\nThis function is designed to display all 2D slices of a 3D data array in a compact, square-like grid layout for efficient visualization.\n\n- **Parameters**:\n  - `data_array` (zarr.Array or numpy.ndarray): A 3D data array, where each slice along the first axis will be displayed in its own subplot.\n\n- **Function Workflow**:\n  1. **Grid Calculation**: The function calculates the number of rows and columns needed to arrange the slices in a grid that approximates a square, using the total number of slices and aiming for a balanced display.\n  2. **Slice Plotting**: Each slice of the 3D array is shown in an individual subplot, labeled with its index to help identify the slice.\n  3. **Unused Subplot Management**: If there are extra subplots beyond the number of slices, these are hidden to maintain a clean visual structure.\n  4. **Display**: The function uses `plt.show()` to render all subplots in a single grid, providing a comprehensive view of all slices at once.\n\nThis function is beneficial for visualizing multiple layers of 3D data in a structured layout that enhances readability and allows quick comparison across slices.","metadata":{}},{"cell_type":"code","source":"def visualize_all_slices_in_square(data_array, slices=None):\n    if isinstance(slices, slice):\n        selected_slices = data_array[slices]\n        indices = range(*slices.indices(data_array.shape[0]))\n    elif isinstance(slices, list):\n        selected_slices = data_array[slices]\n        indices = slices\n    else:\n        selected_slices = data_array\n        indices = range(data_array.shape[0])\n\n    num_slices = selected_slices.shape[0]\n    cols = math.ceil(math.sqrt(num_slices))  # Calculate columns for a square grid\n    rows = math.ceil(num_slices / cols)      # Calculate necessary rows\n\n    fig, axes = plt.subplots(rows, cols, figsize=(12, 12))\n    fig.suptitle('Selected Slices in Square Layout')\n\n    for i in range(rows * cols):\n        if i < num_slices:\n            slice_data = selected_slices[i, :, :]\n            ax = axes[i // cols, i % cols]\n            ax.imshow(slice_data, cmap='gray')\n            ax.set_title(f'Slice {indices[i]}')\n            ax.axis('off')\n        else:\n            axes[i // cols, i % cols].axis('off')  # Hide unused subplots\n\n    plt.tight_layout()\n    plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-12T18:05:40.514935Z","iopub.execute_input":"2024-11-12T18:05:40.515420Z","iopub.status.idle":"2024-11-12T18:05:40.529006Z","shell.execute_reply.started":"2024-11-12T18:05:40.515377Z","shell.execute_reply":"2024-11-12T18:05:40.527511Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"#### `create_gif_from_slices` Function\n\nThis function generates a GIF animation from all 2D slices in a 3D data array, saving it to a specified file path. This animation provides a sequential visualization of the data across the slices.\n\n- **Parameters**:\n  - `data_array` (`zarr.Array` or `numpy.ndarray`): A 3D data array to animate, where each slice along the first axis is used as a frame in the GIF.\n  - `output_path` (`str`): The file path where the GIF will be saved, defaulting to 'slices_animation.gif'.\n  - `duration` (`float`): Duration of each frame in the GIF, in seconds. This parameter controls the playback speed of the animation.\n\n- **Function Workflow**:\n  1. **Prepare Output Path**: Ensures that the directory for `output_path` exists.\n  2. **Create Frames**: Iterates over each slice in `data_array`, generating a frame for each slice. Each frame is plotted as an image without axes for a clean look, captured as an RGB array, and added to a list of frames.\n  3. **Generate and Save GIF**: Combines the list of frames into a GIF using `imageio.mimsave` and saves it to `output_path`.\n\n- **Returns**:\n  - `output_path` (`str`): The file path to the saved GIF.\n\n.","metadata":{}},{"cell_type":"code","source":"def create_gif_from_slices(data_array, output_path='slices_animation.gif', duration=0.1, loop=0):\n    os.makedirs(os.path.dirname(output_path), exist_ok=True)\n\n    frames = []\n    \n    for i in range(data_array.shape[0]):\n        fig, ax = plt.subplots()\n        ax.imshow(data_array[i, :, :], cmap='gray')\n        ax.set_title(f'Slice {i}')\n        ax.axis('off')\n        \n        fig.canvas.draw()\n        frame_image = np.frombuffer(fig.canvas.tostring_rgb(), dtype=np.uint8)\n        frame_image = frame_image.reshape(fig.canvas.get_width_height()[::-1] + (3,))\n        frames.append(frame_image)\n        \n        plt.close(fig)\n    \n    imageio.mimsave(output_path, frames, duration=duration, loop=loop)\n    \n    print(f\"GIF saved at {output_path}\")\n    return output_path","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-12T18:05:43.325795Z","iopub.execute_input":"2024-11-12T18:05:43.326876Z","iopub.status.idle":"2024-11-12T18:05:43.336572Z","shell.execute_reply.started":"2024-11-12T18:05:43.326824Z","shell.execute_reply":"2024-11-12T18:05:43.335369Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"code","source":"zarr_path = '/kaggle/input/czii-cryo-et-object-identification/train/static/ExperimentRuns/TS_6_4/VoxelSpacing10.000/denoised.zarr'","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-12T18:05:45.076102Z","iopub.execute_input":"2024-11-12T18:05:45.077195Z","iopub.status.idle":"2024-11-12T18:05:45.082527Z","shell.execute_reply.started":"2024-11-12T18:05:45.077145Z","shell.execute_reply":"2024-11-12T18:05:45.081076Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"zarr_group = load_zarr_data(zarr_path)\n\ninspect_zarr_structure(zarr_group)\n\ndata_array = get_data_array(zarr_group)\n\ndisplay_data_info(data_array)\n\ndisplay_group_attributes(zarr_group)\n\nvisualize_slice(data_array)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-12T18:05:45.997526Z","iopub.execute_input":"2024-11-12T18:05:45.998013Z","iopub.status.idle":"2024-11-12T18:05:48.672801Z","shell.execute_reply.started":"2024-11-12T18:05:45.997933Z","shell.execute_reply":"2024-11-12T18:05:48.671621Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#visualize_all_slices_sequentially(data_array, cols=5)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-09T18:08:51.941241Z","iopub.execute_input":"2024-11-09T18:08:51.942436Z","iopub.status.idle":"2024-11-09T18:08:51.947772Z","shell.execute_reply.started":"2024-11-09T18:08:51.942375Z","shell.execute_reply":"2024-11-09T18:08:51.946380Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"visualize_all_slices_in_square(data_array, slice(0, 25))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-09T18:17:05.368605Z","iopub.execute_input":"2024-11-09T18:17:05.369131Z","iopub.status.idle":"2024-11-09T18:17:08.858933Z","shell.execute_reply.started":"2024-11-09T18:17:05.369080Z","shell.execute_reply":"2024-11-09T18:17:08.857675Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"gif_path = '/kaggle/working/slices_animation.gif'\ncreate_gif_from_slices(data_array, output_path=gif_path, duration=0.1)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-09T18:31:23.954645Z","iopub.execute_input":"2024-11-09T18:31:23.955636Z","iopub.status.idle":"2024-11-09T18:33:15.385413Z","shell.execute_reply.started":"2024-11-09T18:31:23.955546Z","shell.execute_reply":"2024-11-09T18:33:15.384207Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"Image(filename=gif_path)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-09T18:33:15.387842Z","iopub.execute_input":"2024-11-09T18:33:15.388341Z","iopub.status.idle":"2024-11-09T18:33:15.851736Z","shell.execute_reply.started":"2024-11-09T18:33:15.388287Z","shell.execute_reply":"2024-11-09T18:33:15.849443Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"####  `json_to_pandas` Function\nThis function converts a JSON data array into a pandas.DataFrame format, organizing the data for convenient analysis and manipulation. It extracts key information from each JSON object and creates a table where each row represents a single data point with fields such as location_x, location_y, location_z, transformation, and instance_id.\n\n- **Parameters**:\n\n    - json_data (list): A list of JSON objects containing data to be processed.\n\nEach object should contain an array called points, where each element represents a data point with information about location, transformation, and instance_id.","metadata":{}},{"cell_type":"code","source":"def json_to_pandas(json_data):\n    records = []\n    for item in json_data:\n        for point in item.get('points', []):\n            record = {\n                'location_x': point.get('location', {}).get('x', None),\n                'location_y': point.get('location', {}).get('y', None),\n                'location_z': point.get('location', {}).get('z', None),\n                'transformation': point.get('transformation', None),\n                'instance_id': point.get('instance_id', None)\n            }\n            records.append(record)\n\n    df = pd.DataFrame(records)\n    return df","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-12T18:05:53.336667Z","iopub.execute_input":"2024-11-12T18:05:53.337280Z","iopub.status.idle":"2024-11-12T18:05:53.346420Z","shell.execute_reply.started":"2024-11-12T18:05:53.337234Z","shell.execute_reply":"2024-11-12T18:05:53.345039Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"json_data = []\n\nfor i in range(590, 632):\n    try:\n        with open(paths[i], 'r', encoding='utf-8') as file:\n            data = json.load(file)\n            json_data.append(data)\n    except (FileNotFoundError, json.JSONDecodeError) as e:\n        print(f\"Error {paths[i]}: {e}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-12T18:05:54.608961Z","iopub.execute_input":"2024-11-12T18:05:54.609723Z","iopub.status.idle":"2024-11-12T18:05:54.665456Z","shell.execute_reply.started":"2024-11-12T18:05:54.609676Z","shell.execute_reply":"2024-11-12T18:05:54.664049Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"objects = json_to_pandas(json_data)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-12T18:05:55.294214Z","iopub.execute_input":"2024-11-12T18:05:55.294684Z","iopub.status.idle":"2024-11-12T18:05:55.307417Z","shell.execute_reply.started":"2024-11-12T18:05:55.294640Z","shell.execute_reply":"2024-11-12T18:05:55.306028Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"objects.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-12T18:05:55.962593Z","iopub.execute_input":"2024-11-12T18:05:55.963156Z","iopub.status.idle":"2024-11-12T18:05:55.979382Z","shell.execute_reply.started":"2024-11-12T18:05:55.963107Z","shell.execute_reply":"2024-11-12T18:05:55.978160Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"x, y, z = objects['location_x'], objects['location_y'], objects['location_z']\n\nfig = go.Figure(data=[go.Scatter3d(\n    x=x, y=y, z=z,\n    mode='markers',\n    marker=dict(\n        size=5,\n        color=z,\n        colorscale='Viridis',\n        opacity=0.8\n    ),\n)])\n\nfig.update_layout(\n    scene=dict(\n        xaxis_title='X Coordinate',\n        yaxis_title='Y Coordinate',\n        zaxis_title='Z Coordinate'\n    ),\n    title=\"3D Map of Objects\",\n    width=700,\n    margin=dict(r=10, l=10, b=10, t=40)\n)\n\nfig.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-12T18:05:57.126873Z","iopub.execute_input":"2024-11-12T18:05:57.127372Z","iopub.status.idle":"2024-11-12T18:05:57.154878Z","shell.execute_reply.started":"2024-11-12T18:05:57.127325Z","shell.execute_reply":"2024-11-12T18:05:57.153456Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"---\n\n# WORK IN PROGRESS\n\nNext steps:\n- projection object markers to slice\n- object detection in images\n- model building\n\n.","metadata":{}}]}