{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":91249,"databundleVersionId":11294684,"sourceType":"competition"}],"dockerImageVersionId":30918,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## BYU - Locating Bacterial Flagellar Motors 2025","metadata":{}},{"cell_type":"markdown","source":"## 1. Setup","metadata":{}},{"cell_type":"code","source":"from glob import glob\nfrom pathlib import Path\nimport numpy as np\nimport pandas as pd\nimport cv2\nimport matplotlib.pyplot as plt","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2025-03-14T09:44:34.176682Z","iopub.execute_input":"2025-03-14T09:44:34.177065Z","iopub.status.idle":"2025-03-14T09:44:34.181792Z","shell.execute_reply.started":"2025-03-14T09:44:34.177038Z","shell.execute_reply":"2025-03-14T09:44:34.180587Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"competition_dataset_directory = Path('/kaggle/input/byu-locating-bacterial-flagellar-motors-2025')\n\npd.set_option('display.max_rows', 100)\npd.set_option('display.max_columns', 100)\npd.set_option('display.width', 1000)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-14T10:07:05.519968Z","iopub.execute_input":"2025-03-14T10:07:05.520425Z","iopub.status.idle":"2025-03-14T10:07:05.525561Z","shell.execute_reply.started":"2025-03-14T10:07:05.520395Z","shell.execute_reply":"2025-03-14T10:07:05.524352Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"training_tomograms = glob(str(competition_dataset_directory / 'train' / '*'))\ntraining_images = glob(str(competition_dataset_directory / 'train' / '*' / '*.jpg'))\ndf_train = pd.read_csv(competition_dataset_directory / 'train_labels.csv')\n\nprint(f'Training Tomogram Count {len(training_tomograms)}')\nprint(f'Training Images Count {len(training_images)}')\nprint(f'Training Set Shape: {df_train.shape}')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-14T09:43:12.848619Z","iopub.execute_input":"2025-03-14T09:43:12.848938Z","iopub.status.idle":"2025-03-14T09:43:23.605466Z","shell.execute_reply.started":"2025-03-14T09:43:12.848913Z","shell.execute_reply":"2025-03-14T09:43:23.604441Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 2. Introduction\n\n### Cryogenic electron tomography (cryo-ET)\n\nCryogenic Electron Tomography (cryo-ET) is an advanced imaging technique used to create high-resolution (10–40 Å) 3D models of tiny structures, such as biological molecules and cells. It works by tilting a frozen sample under a transmission electron microscope (Cryo-TEM) to capture multiple 2D images from different angles. These images are then combined to form a detailed 3D reconstruction, similar to how a CT scan works for the human body. Unlike other electron microscopy methods, cryo-ET keeps samples at extremely low temperatures (< -150°C), preserving their natural state in a thin layer of vitreous ice without dehydration or chemical damage.\n\n![cryo_et](https://i.ibb.co/0yk5JwtX/Screenshot-from-2025-03-07-20-05-59.png)\n\n### Flagellar Motors\n\nA flagellar motor is a tiny, natural molecular machine that helps bacteria move. It works like a biological rotary engine, spinning a long, whip-like structure called a flagellum to push the bacterium forward. The motor is embedded in the bacterial cell wall and powered by ions (like protons). It rotates the flagellum at high speeds, allowing the bacterium to swim toward food or escape harmful environments. This movement is crucial for processes like chemotaxis (moving toward or away from chemicals) and even bacterial infections.\n\n![flagellar_motor](https://i.ibb.co/q398w2T7/Screenshot-from-2025-03-07-19-58-39.png)\n\nThe problem is to identify the presence and location of flagellar motors in 3D tomograms of bacteria, reconstructed from stacks of 2D slices. The training dataset consists of **648** tomograms, totaling **269,194** images, with a structured label set of **737** entries. These slices represent different cross-sections of the bacteria, and when stacked together, they form a 3D tomogram.","metadata":{}},{"cell_type":"code","source":"display(df_train)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-14T09:43:23.607159Z","iopub.execute_input":"2025-03-14T09:43:23.607554Z","iopub.status.idle":"2025-03-14T09:43:23.647985Z","shell.execute_reply.started":"2025-03-14T09:43:23.607519Z","shell.execute_reply":"2025-03-14T09:43:23.646672Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 3. Labels\n\nThe labels in this competition represent the **center coordinates** of flagellar motors in a 3D tomogram, given as **(z, y, x)** positions. Since the evaluation metric is based on **Euclidean distance**, a motor is considered correctly detected if the predicted location falls within a **1000 angstrom (Å)** radius of the true center. This effectively means that each motor's label defines a spherical region in which predictions are counted as true positives. However, flagellar motors are not perfect spheres, they are more cylindrical, and their orientation within the tomogram is unknown, adding another layer of complexity to detection. Tomograms that do not contain a flagellar motor are labeled with (-1, -1, -1) to indicate the absence of a motor.","metadata":{}},{"cell_type":"code","source":"def visualize_counts(df, value, title):\n        \n    fig, ax = plt.subplots(figsize=(24, 16))\n    ax.barh(\n        y=np.arange(df.shape[0]),\n        width=df['count'].values,\n        align='center'\n    )\n    ax.set_yticks(np.arange(df.shape[0]))\n    ax.set_yticklabels([\n        f'{val} - Count: {count} ({percentage:.2f}%)' for val, count, percentage in zip(\n            df[value].values,\n            df['count'].values,\n            df['percentage'].values,\n        )\n    ])\n    ax.set_xlabel('')\n    ax.tick_params(axis='x', labelsize=17.5, pad=10)\n    ax.tick_params(axis='y', labelsize=17.5, pad=10)\n    ax.set_title(title, size=20, pad=15)\n    plt.gca().invert_yaxis()\n    plt.show()\n","metadata":{"trusted":true,"_kg_hide-input":true,"execution":{"iopub.status.busy":"2025-03-14T09:43:23.648916Z","iopub.execute_input":"2025-03-14T09:43:23.649328Z","iopub.status.idle":"2025-03-14T09:43:23.655945Z","shell.execute_reply.started":"2025-03-14T09:43:23.649300Z","shell.execute_reply":"2025-03-14T09:43:23.655077Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"In the training set, the majority of tomograms contain zero or one flagellar motor. Specifically, **48.3% (313 tomograms)** have no motors, while **44.1% (286 tomograms)** contain a single motor. A smaller fraction of tomograms have multiple motors, with **30 tomograms (4.6%)** containing two motors, and even fewer cases with three or more. The most extreme cases include one tomogram with **10 motors (0.15%)** and a few with four to six motors. This distribution suggests that while single-motor tomograms are common, handling cases with multiple motors requires extra attention. The test set consists exclusively of tomograms with **zero or one** motor.","metadata":{}},{"cell_type":"code","source":"df_train_n_motors = df_train.groupby('tomo_id')['Number of motors'].first().reset_index(drop=True).value_counts().reset_index()\ndf_train_n_motors['percentage'] = df_train_n_motors['count'] / df_train_n_motors['count'].sum(axis=0) * 100\n\nvisualize_counts(\n    df=df_train_n_motors,\n    value='Number of motors',\n    title='Number of motors Value Counts'\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-14T09:43:23.656884Z","iopub.execute_input":"2025-03-14T09:43:23.657177Z","iopub.status.idle":"2025-03-14T09:43:24.180465Z","shell.execute_reply.started":"2025-03-14T09:43:23.657154Z","shell.execute_reply":"2025-03-14T09:43:24.179393Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Some tomograms are listed as containing two motors, but only one label is provided for them in the training data. This suggests possible inconsistencies in the labels, where an expected motor location might be missing.","metadata":{}},{"cell_type":"code","source":"df_train['tomogram_row_count'] = df_train.groupby('tomo_id')['tomo_id'].transform('count')\ndf_missing_labels = df_train.loc[(df_train['tomogram_row_count'] != df_train['Number of motors']) & (df_train['Number of motors'] != 0)]\ndisplay(df_missing_labels)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-14T09:43:24.181539Z","iopub.execute_input":"2025-03-14T09:43:24.181872Z","iopub.status.idle":"2025-03-14T09:43:24.202197Z","shell.execute_reply.started":"2025-03-14T09:43:24.181835Z","shell.execute_reply":"2025-03-14T09:43:24.201051Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The distribution of flagellar motor coordinates in the training set shows variability across all three axes. The z-axis (Motor axis 0) has a mean of 167 with values ranging from 0 to 466, while the y-axis (Motor axis 1) ranges from 59 to 904 with a mean of 482, and the x-axis (Motor axis 2) ranges from 30 to 902 with a mean of 492. The standard deviation is smallest for the z-axis (78.1), indicating that motors are more tightly clustered along this dimension, whereas the y and x axes have broader distributions (198.6 and 214.8, respectively).","metadata":{}},{"cell_type":"code","source":"def visualize_histogram(df, columns, title):\n\n    fig, ax = plt.subplots(figsize=(24, 6))\n    for column in columns:\n        ax.hist(df[column], 32, alpha=0.5, label=column)\n    ax.tick_params(axis='x', labelsize=15)\n    ax.tick_params(axis='y', labelsize=15)\n    ax.set_title(title, size=20, pad=15)\n    ax.legend(prop={'size': 14})\n    plt.show()\n\n\ndf_train_with_motors = df_train.loc[df_train['Motor axis 0'] != -1].reset_index(drop=True)\n\nvisualize_histogram(\n    df_train_with_motors,\n    columns=['Motor axis 0', 'Motor axis 1', 'Motor axis 2'],\n    title='Label Coordinate Histograms'\n)\n\ndisplay(df_train_with_motors[['Motor axis 0', 'Motor axis 1', 'Motor axis 2']].describe())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-14T09:43:24.203400Z","iopub.execute_input":"2025-03-14T09:43:24.203791Z","iopub.status.idle":"2025-03-14T09:43:24.641476Z","shell.execute_reply.started":"2025-03-14T09:43:24.203750Z","shell.execute_reply":"2025-03-14T09:43:24.640391Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The tomogram shapes in the training set exhibit low cardinality, meaning there are only a few distinct shape variations. The z-axis (depth) varies more than the y (height) and x (width) axes, with a wide range from 300 to 800. In contrast, the y-axis and x-axis dimensions are more consistent, with means of 950 and 954, respectively, and standard deviations of only 64.9 and 97.2. The most common shapes cluster around (300, 928, 928) and (500, 960, 956), with a few outliers reaching up to (800, 1912, 1847). This suggests that many tomograms share the same or very similar dimensions.","metadata":{"_kg_hide-input":false}},{"cell_type":"code","source":"visualize_histogram(\n    df_train,\n    columns=['Array shape (axis 0)', 'Array shape (axis 1)', 'Array shape (axis 2)'],\n    title='Tomogram Shape Histograms'\n)\n\ndisplay(df_train[['Array shape (axis 0)', 'Array shape (axis 1)', 'Array shape (axis 2)']].describe())","metadata":{"trusted":true,"_kg_hide-input":true,"execution":{"iopub.status.busy":"2025-03-14T09:43:24.644747Z","iopub.execute_input":"2025-03-14T09:43:24.645076Z","iopub.status.idle":"2025-03-14T09:43:25.155531Z","shell.execute_reply.started":"2025-03-14T09:43:24.645014Z","shell.execute_reply":"2025-03-14T09:43:25.154441Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The normalized label coordinate distributions provide insight into where flagellar motors tend to be located within the tomograms. On average, motors are positioned at ~46.7% of the z-depth, ~50.9% of the y-height, and ~52.5% of the x-width, indicating a slight central tendency, especially in the y and x axes. However, there is notable spread, with standard deviations of ~16.5% (z), ~21.0% (y), and ~22.8% (x). The minimum values suggest that some motors are located very close to the tomogram edges, while the maximum values reach over 90% of each axis, confirming that motors can appear throughout the tomogram volume.","metadata":{}},{"cell_type":"code","source":"df_train_with_motors['Normalized Motor axis 0'] = df_train_with_motors['Motor axis 0'] / df_train_with_motors['Array shape (axis 0)']\ndf_train_with_motors['Normalized Motor axis 1'] = df_train_with_motors['Motor axis 1'] / df_train_with_motors['Array shape (axis 1)']\ndf_train_with_motors['Normalized Motor axis 2'] = df_train_with_motors['Motor axis 2'] / df_train_with_motors['Array shape (axis 2)']","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-14T09:43:25.157563Z","iopub.execute_input":"2025-03-14T09:43:25.157983Z","iopub.status.idle":"2025-03-14T09:43:25.165216Z","shell.execute_reply.started":"2025-03-14T09:43:25.157942Z","shell.execute_reply":"2025-03-14T09:43:25.163889Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"visualize_histogram(\n    df_train_with_motors,\n    columns=['Normalized Motor axis 0', 'Normalized Motor axis 1', 'Normalized Motor axis 2'],\n    title='Normalized Label Coordinate Histograms'\n)\n\ndisplay(df_train_with_motors[['Normalized Motor axis 0', 'Normalized Motor axis 1', 'Normalized Motor axis 2']].describe())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-14T09:43:25.166478Z","iopub.execute_input":"2025-03-14T09:43:25.166929Z","iopub.status.idle":"2025-03-14T09:43:25.612347Z","shell.execute_reply.started":"2025-03-14T09:43:25.166889Z","shell.execute_reply":"2025-03-14T09:43:25.611327Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The heatmaps of label coordinates provide a clearer view of how flagellar motors are distributed across different tomographic planes. The YX plane (top-down view) reveals that motors are fairly scattered across the width and height of the tomograms, suggesting that their placement along these axes is relatively unconstrained. However, in the ZY (side view) and ZX (front view) planes, the distribution shows a more noticeable concentration near the center. This suggests that motors tend to be positioned around the midsections of the tomograms along the depth (Z) axis rather than being uniformly spread. The slight central bias in these planes could be due to biological factors influencing motor placement or dataset biases related to how tomograms were captured and annotated.","metadata":{}},{"cell_type":"code","source":"def visualize_heatmap(heatmap, title):\n\n    fig, ax = plt.subplots(figsize=(8, 8))\n    ax.imshow(heatmap, cmap=plt.cm.hot)\n    ax.set_xlabel('')\n    ax.set_ylabel('')\n    ax.tick_params(axis='x', labelsize=15, pad=10)\n    ax.tick_params(axis='y', labelsize=15, pad=10)\n    ax.set_title(title, size=15, pad=12.5, loc='center', wrap=True)\n    plt.show()\n\n\nyx_coordinate_heatmap = np.zeros((1000, 1000))\n\nfor _, row in df_train_with_motors.iterrows():\n    yx_coordinate_heatmap[\n        int(row['Normalized Motor axis 1'] * 1000) - 3:int(row['Normalized Motor axis 1'] * 1000) + 3,\n        int(row['Normalized Motor axis 2'] * 1000) - 3:int(row['Normalized Motor axis 2'] * 1000) + 3,\n        \n    ] += 1\n\nvisualize_heatmap(heatmap=yx_coordinate_heatmap, title='Label Coordinate Heatmap on YX Plane')\n\nzy_coordinate_heatmap = np.zeros((1000, 1000))\n\nfor _, row in df_train_with_motors.iterrows():\n    zy_coordinate_heatmap[\n        int(row['Normalized Motor axis 0'] * 1000) - 3:int(row['Normalized Motor axis 0'] * 1000) + 3,\n        int(row['Normalized Motor axis 1'] * 1000) - 3:int(row['Normalized Motor axis 1'] * 1000) + 3,\n        \n    ] += 1\n\nvisualize_heatmap(heatmap=zy_coordinate_heatmap, title='Label Coordinate Heatmap on ZY Plane')\n\nzx_coordinate_heatmap = np.zeros((1000, 1000))\n\nfor _, row in df_train_with_motors.iterrows():\n    zx_coordinate_heatmap[\n        int(row['Normalized Motor axis 0'] * 1000) - 3:int(row['Normalized Motor axis 0'] * 1000) + 3,\n        int(row['Normalized Motor axis 2'] * 1000) - 3:int(row['Normalized Motor axis 2'] * 1000) + 3,\n        \n    ] += 1\n\nvisualize_heatmap(heatmap=zx_coordinate_heatmap, title='Label Coordinate Heatmap on ZX Plane')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-14T09:43:25.613316Z","iopub.execute_input":"2025-03-14T09:43:25.613652Z","iopub.status.idle":"2025-03-14T09:43:26.808313Z","shell.execute_reply.started":"2025-03-14T09:43:25.613626Z","shell.execute_reply":"2025-03-14T09:43:26.807153Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 4. Real-World Units\n\nThe units used in this dataset are in angstroms (Å), a common measurement in structural biology and microscopy. One angstrom (1Å) is equal to 0.1 nanometers (nm), making it a convenient unit for describing molecular and subcellular structures. Since the voxel spacing column provides the resolution of the tomograms in angstroms per voxel, it allows for the conversion of pixel-based measurements into real-world physical dimensions. This ensures that distances, areas, and volumes can be analyzed in meaningful metric units rather than arbitrary pixel values, enabling comparisons across tomograms with different resolutions. Additionally, the competition metric considers a 1000 Å sphere around each labeled motor coordinate when determining true positives.","metadata":{}},{"cell_type":"markdown","source":"Voxel spacing and angstroms per voxel essentially represent the same concept. They both define the real-world physical size of each voxel in the tomograms. They tell us how many angstroms correspond to a single voxel along each axis. The voxel spacing values in the dataset are grouped around certain common numbers rather than being spread out evenly. The average voxel spacing is 15.07 Å, with most values falling between 13.1 Å and 16.1 Å. The most frequent values; 13.1 Å, 15.6 Å, and 19.7 Å appear often, showing that the tomograms were likely captured using a few specific settings. The smallest spacing recorded is 6.5 Å, while the largest is 19.7 Å. Since voxel spacing determines how large objects appear in real-world units, these clusters are important when comparing flagellar motor sizes across different tomograms.","metadata":{}},{"cell_type":"code","source":"df_train_voxel_spacings = df_train.groupby('tomo_id')['Voxel spacing'].first().reset_index(drop=True).value_counts().reset_index()\ndf_train_voxel_spacings['percentage'] = df_train_voxel_spacings['count'] / df_train_voxel_spacings['count'].sum(axis=0) * 100\n\nvisualize_counts(\n    df=df_train_voxel_spacings,\n    value='Voxel spacing',\n    title='Voxel spacing Value Counts'\n)\n\nvisualize_histogram(\n    df_train.groupby('tomo_id')['Voxel spacing'].first().reset_index(),\n    columns=['Voxel spacing'],\n    title='Voxel Spacing Histogram'\n)\n\ndisplay(df_train.groupby('tomo_id')[['Voxel spacing']].first().describe())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-14T09:43:26.809251Z","iopub.execute_input":"2025-03-14T09:43:26.809530Z","iopub.status.idle":"2025-03-14T09:43:27.494479Z","shell.execute_reply.started":"2025-03-14T09:43:26.809507Z","shell.execute_reply":"2025-03-14T09:43:27.493438Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The total number of voxels in a tomogram represents the full 3D grid of discrete volume elements used to capture structural details. Each voxel corresponds to a small cubic region of space, and the total count depends on the tomogram’s dimensions along the X, Y, and Z axes.\n\nFrom the dataset statistics, the average tomogram contains approximately 382 million voxels, with a standard deviation of 197 million, indicating significant variation in tomogram sizes. The smallest tomogram has 258 million voxels, while the largest reaches nearly 1.77 billion voxels, highlighting the wide range of spatial resolutions and field-of-view among samples. Notably, 50% of tomograms have fewer than 267 million voxels, showing that many follow a common size pattern, while the upper quartile extends to 442 million voxels. This variation implies that models must be capable of handling tomograms of different sizes efficiently, as voxel counts influence both computational requirements and the density of information available for flagellar motor detection.","metadata":{}},{"cell_type":"code","source":"df_train['total_voxels'] = df_train['Array shape (axis 0)'] * df_train['Array shape (axis 1)'] * df_train['Array shape (axis 2)']","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-14T09:43:27.495439Z","iopub.execute_input":"2025-03-14T09:43:27.495750Z","iopub.status.idle":"2025-03-14T09:43:27.501583Z","shell.execute_reply.started":"2025-03-14T09:43:27.495727Z","shell.execute_reply":"2025-03-14T09:43:27.500305Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"visualize_histogram(\n    df_train.groupby('tomo_id')['total_voxels'].first().reset_index(),\n    columns=['total_voxels'],\n    title='Total Voxels Histogram'\n)\n\ndisplay(df_train.groupby('tomo_id')[['total_voxels']].first().describe())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-14T09:43:27.502654Z","iopub.execute_input":"2025-03-14T09:43:27.502946Z","iopub.status.idle":"2025-03-14T09:43:27.840711Z","shell.execute_reply.started":"2025-03-14T09:43:27.502914Z","shell.execute_reply":"2025-03-14T09:43:27.839591Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"It makes sense to examine required voxel proximity because it directly impacts how precise a model's predictions need to be in voxel space. Since a prediction is considered correct if it falls within a 1000 Å sphere, the number of voxels within that range depends on the voxel spacing of each tomogram. A lower voxel spacing means more voxels fit within 1000 Å, increasing the voxel proximity value, while a higher voxel spacing reduces it.\n\nThe histogram shows that most tomograms require predictions to be within 50 to 76 voxels to be correct, with a median of 64 voxels. The mean is 67.8, suggesting a slight skew toward higher values. The minimum is 50.76, meaning the least dense tomograms require roughly 51 voxels for correctness, while the maximum of 153.85 indicates some tomograms have much finer resolution, requiring a significantly higher number of voxels within the 1000 Å range.\n\nThis variation in required voxel proximity highlights a key challenge for models: handling different voxel resolutions across tomograms. Models must be able to adjust predictions accordingly, ensuring they remain within the correct voxel range across different tomogram scales.","metadata":{}},{"cell_type":"code","source":"df_train['required_voxel_proximity'] = 1000 / df_train['Voxel spacing']","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-14T09:43:27.841869Z","iopub.execute_input":"2025-03-14T09:43:27.842336Z","iopub.status.idle":"2025-03-14T09:43:27.847846Z","shell.execute_reply.started":"2025-03-14T09:43:27.842302Z","shell.execute_reply":"2025-03-14T09:43:27.846781Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"visualize_histogram(\n    df_train.groupby('tomo_id')['required_voxel_proximity'].first().reset_index(),\n    columns=['required_voxel_proximity'],\n    title='Required Voxel Proximity Histogram'\n)\n\ndisplay(df_train.groupby('tomo_id')[['required_voxel_proximity']].first().describe())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-14T09:43:27.848799Z","iopub.execute_input":"2025-03-14T09:43:27.849103Z","iopub.status.idle":"2025-03-14T09:43:28.155399Z","shell.execute_reply.started":"2025-03-14T09:43:27.849079Z","shell.execute_reply":"2025-03-14T09:43:28.154432Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The valid voxel count represents the number of voxels within each tomogram that fall within the 1,000-angstrom proximity threshold around the motor’s center. This measurement helps estimate the potential search space for a correct prediction, as a prediction is considered valid if it falls within this region.\n\nOn average, a tomogram contains 1.41 million valid voxels, though this varies significantly, with a standard deviation of 1.06 million. The smallest valid voxel count is around 548,000, while the largest reaches 15.25 million, reflecting differences in voxel spacing and tomogram dimensions. Half of the tomograms have fewer than 1.1 million valid voxels, while the upper quartile extends to 1.86 million, suggesting that some tomograms require handling significantly larger search areas. The presence of a few tomograms with exceptionally high valid voxel counts underscores the need for models that can adapt to varying tomogram sizes while maintaining precision in detection.","metadata":{}},{"cell_type":"code","source":"df_train['valid_voxels'] = ((4 / 3) * np.pi * (df_train['required_voxel_proximity'] ** 3)).astype(int)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-14T09:43:28.156270Z","iopub.execute_input":"2025-03-14T09:43:28.156549Z","iopub.status.idle":"2025-03-14T09:43:28.162998Z","shell.execute_reply.started":"2025-03-14T09:43:28.156513Z","shell.execute_reply":"2025-03-14T09:43:28.161825Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"visualize_histogram(\n    df_train.groupby('tomo_id')['valid_voxels'].first().reset_index(),\n    columns=['valid_voxels'],\n    title='Valid Voxels Histogram'\n)\n\ndisplay(df_train.groupby('tomo_id')[['valid_voxels']].first().describe())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-14T09:43:28.164204Z","iopub.execute_input":"2025-03-14T09:43:28.164473Z","iopub.status.idle":"2025-03-14T09:43:28.484764Z","shell.execute_reply.started":"2025-03-14T09:43:28.164451Z","shell.execute_reply":"2025-03-14T09:43:28.483679Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The valid voxel percentage represents the fraction of voxels within a tomogram that fall inside the 1,000-angstrom proximity threshold relative to the total voxel count. This percentage provides insight into how much of the 3D volume is relevant for classification.\n\nOn average, only 0.39% of voxels in a tomogram are considered valid for classification, with values ranging from 0.08% to 0.87%. The standard deviation of 0.20% indicates moderate variability across tomograms. Half of the tomograms have a valid voxel percentage above 0.41%, suggesting that even in the most favorable cases, only a small fraction of the tomogram is useful for identifying motors. The low percentages highlight the challenge of localization. Models must efficiently focus on these small yet crucial regions.","metadata":{}},{"cell_type":"code","source":"df_train['valid_voxel_percentage'] = df_train['valid_voxels'] / df_train['total_voxels'] * 100","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-14T09:43:28.485870Z","iopub.execute_input":"2025-03-14T09:43:28.486233Z","iopub.status.idle":"2025-03-14T09:43:28.492438Z","shell.execute_reply.started":"2025-03-14T09:43:28.486206Z","shell.execute_reply":"2025-03-14T09:43:28.491216Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"visualize_histogram(\n    df_train.groupby('tomo_id')['valid_voxel_percentage'].first().reset_index(),\n    columns=['valid_voxel_percentage'],\n    title='Valid Voxel Percentage Histogram'\n)\n\ndisplay(df_train.groupby('tomo_id')[['valid_voxel_percentage']].first().describe())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-14T09:43:28.493571Z","iopub.execute_input":"2025-03-14T09:43:28.493921Z","iopub.status.idle":"2025-03-14T09:43:28.836730Z","shell.execute_reply.started":"2025-03-14T09:43:28.493896Z","shell.execute_reply":"2025-03-14T09:43:28.835697Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Examining the correlations between voxel spacing, total voxel count, and valid voxel percentage provides further insight. There is a strong negative correlation (-0.68) between voxel spacing and valid voxel percentage, meaning that as voxel spacing increases, the proportion of useful voxels decreases. This makes sense because larger voxel spacing reduces the number of voxels that fall within the 1,000-angstrom proximity threshold. Additionally, total voxel count and valid voxel percentage are negatively correlated (-0.46), indicating that larger tomograms tend to have a lower proportion of valid voxels. Meanwhile, voxel spacing and total voxel count have a weaker negative correlation (-0.20), suggesting that larger tomograms do not necessarily have significantly different voxel spacings. These relationships emphasize the importance of voxel resolution in determining how much of a tomogram is actually useful for motor localization.","metadata":{}},{"cell_type":"code","source":"df_train.groupby('tomo_id')[['Voxel spacing', 'total_voxels', 'valid_voxel_percentage']].first().corr()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-14T09:43:28.837755Z","iopub.execute_input":"2025-03-14T09:43:28.838142Z","iopub.status.idle":"2025-03-14T09:43:28.852578Z","shell.execute_reply.started":"2025-03-14T09:43:28.838119Z","shell.execute_reply":"2025-03-14T09:43:28.851467Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 5. Tomograms\n\nTomograms used in cryo-electron tomography (cryo-ET) are typically composed of 8-bit grayscale images, where each voxel intensity represents electron density. These images are captured as a series of 2D slices that, when stacked together, form a 3D reconstruction of a biological sample at nanometer resolution. Since the images are grayscale, each pixel has an intensity ranging from 0 to 255, with lower values representing less-dense regions and higher values indicating denser structures.","metadata":{}},{"cell_type":"code","source":"def visualize_slice_with_annotations(df, tomo_id, slice_idx):\n\n    image_paths = sorted(glob(str(competition_dataset_directory / 'train' / tomo_id / '*')))\n    image = cv2.imread(image_paths[slice_idx], cv2.IMREAD_GRAYSCALE)\n    image_rgb = cv2.cvtColor(image, cv2.COLOR_GRAY2RGB)\n\n    tomo_mask = df['tomo_id'] == tomo_id\n    motor_center_coordinates = df.loc[tomo_mask, ['Motor axis 0', 'Motor axis 1', 'Motor axis 2']].values\n    voxel_spacing = df.loc[tomo_mask, 'Voxel spacing'].values[0]\n\n    fig, ax = plt.subplots(figsize=(8, 8))\n    \n    if len(motor_center_coordinates) == 1 and motor_center_coordinates.sum() == -3:\n        ax.imshow(image_rgb)\n    else:\n        overlay = image_rgb.copy()\n        alpha = 0.25\n        motor_radius = int(round(1000 / voxel_spacing))\n    \n        for z, y, x in motor_center_coordinates:\n            \n            depth_diff = abs(z - slice_idx)\n            if depth_diff <= motor_radius:\n                visible_radius = int(round((motor_radius ** 2 - depth_diff ** 2) ** 0.5))\n                cv2.circle(image_rgb, (int(x), int(y)), visible_radius, (255, 0, 0), -1)\n                cv2.circle(image_rgb, (int(x), int(y)), visible_radius, (0, 255, 0), 2)\n                annotation_text = f'({int(x)}, {int(y)}, {int(z)}) r: {visible_radius:.2f}'\n                cv2.putText(\n                    image_rgb,\n                    annotation_text,\n                    (int(x) + 25, int(y) - 10),\n                    cv2.FONT_HERSHEY_SIMPLEX,\n                    0.65,\n                    (0, 0, 255),\n                    2,\n                    cv2.LINE_AA\n                )\n\n        cv2.addWeighted(overlay, alpha, image_rgb, 1 - alpha, 0, image_rgb)\n        ax.imshow(image_rgb)\n        \n    ax.axis('off')\n    ax.set_title(f'{tomo_id} - Slice: {slice_idx} - Voxel Spacing: {voxel_spacing}\\nImage Shape: {image.shape}', size=15, pad=12.5, loc='center', wrap=True)\n    plt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-14T09:46:38.768709Z","iopub.execute_input":"2025-03-14T09:46:38.769101Z","iopub.status.idle":"2025-03-14T09:46:38.779554Z","shell.execute_reply.started":"2025-03-14T09:46:38.769071Z","shell.execute_reply":"2025-03-14T09:46:38.778465Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The tomogram **tomo_003acc** has the lowest voxel spacing (**6.5**) among all the samples, meaning that each voxel represents a smaller physical distance compared to other tomograms. As a result, objects within this tomogram appear significantly larger when visualized, even though their true physical size remains the same. This effect occurs because the same number of pixels now covers a much smaller real-world volume, effectively increasing the apparent size of structures such as the motors. The high level of detail in similar tomograms may provide more precise localization of objects, but it also requires careful interpretation to avoid misjudging relative sizes when comparing across different tomograms.","metadata":{}},{"cell_type":"code","source":"visualize_slice_with_annotations(df_train, 'tomo_003acc', slice_idx=200)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-14T09:49:30.976214Z","iopub.execute_input":"2025-03-14T09:49:30.976562Z","iopub.status.idle":"2025-03-14T09:49:31.757501Z","shell.execute_reply.started":"2025-03-14T09:49:30.976538Z","shell.execute_reply":"2025-03-14T09:49:31.756192Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The tomogram **tomo_00e047** has a voxel spacing of **15.6**, which is close to the average among all samples. This means the objects within the tomogram appear at a typical scale, neither excessively large nor too small compared to the dataset as a whole. With this voxel spacing, the visualization provides a balanced level of detail, maintaining sufficient resolution to distinguish structures while avoiding excessive magnification effects seen in lower-spacing tomograms. As a result, it serves as a representative example for analyzing object appearances in the dataset under standard imaging conditions.","metadata":{}},{"cell_type":"code","source":"visualize_slice_with_annotations(df_train, 'tomo_00e047', slice_idx=200)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-14T10:00:18.959169Z","iopub.execute_input":"2025-03-14T10:00:18.959500Z","iopub.status.idle":"2025-03-14T10:00:19.489261Z","shell.execute_reply.started":"2025-03-14T10:00:18.959477Z","shell.execute_reply":"2025-03-14T10:00:19.487996Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The tomogram **tomo_00e463** has a voxel spacing of **19.7**, which is the highest value in the dataset. This larger voxel spacing results in objects appearing smaller compared to those in tomograms with lower voxel spacing. Despite this reduced size, it is notable for containing **6** motors, which is one of the highest counts in the dataset. The combination of a higher voxel spacing and multiple motors presents a unique challenge for visualization and analysis, as the motors may appear more condensed and the finer details might be harder to distinguish without proper scaling.","metadata":{}},{"cell_type":"code","source":"visualize_slice_with_annotations(df_train, 'tomo_00e463', slice_idx=250)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-14T10:02:56.290802Z","iopub.execute_input":"2025-03-14T10:02:56.291183Z","iopub.status.idle":"2025-03-14T10:02:56.832714Z","shell.execute_reply.started":"2025-03-14T10:02:56.291155Z","shell.execute_reply":"2025-03-14T10:02:56.831217Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null}]}