{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceId":91249,"databundleVersionId":11294684,"isSourceIdPinned":false,"sourceType":"competition"},{"sourceId":14776167,"sourceType":"datasetVersion","datasetId":9445236}],"dockerImageVersionId":30918,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## BYU - Locating Bacterial Flagellar Motors 2025","metadata":{}},{"cell_type":"code","source":"from glob import glob\nfrom pathlib import Path\nimport numpy as np\nimport pandas as pd\nimport cv2\nimport matplotlib.pyplot as plt","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2026-02-09T10:55:54.011338Z","iopub.execute_input":"2026-02-09T10:55:54.011707Z","iopub.status.idle":"2026-02-09T10:55:54.790010Z","shell.execute_reply.started":"2026-02-09T10:55:54.011675Z","shell.execute_reply":"2026-02-09T10:55:54.789079Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\n\nemdb_dir = Path(\n    \"/kaggle/input/dataset-for-flagellar-motors-motor-and-non-motor/tomograms\"\n)\n\nmotor_imgs = glob(str(emdb_dir / \"Motor\" / \"*.png\"))\nnonmotor_imgs = glob(str(emdb_dir / \"Non motor\" / \"*.png\"))\n\nprint(\"Motor images:\", len(motor_imgs))\nprint(\"Non-motor images:\", len(nonmotor_imgs))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-09T10:58:45.932546Z","iopub.execute_input":"2026-02-09T10:58:45.932851Z","iopub.status.idle":"2026-02-09T10:58:45.967834Z","shell.execute_reply.started":"2026-02-09T10:58:45.932830Z","shell.execute_reply":"2026-02-09T10:58:45.967186Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"competition_dataset_directory = Path('/kaggle/input/byu-locating-bacterial-flagellar-motors-2025')\n\npd.set_option('display.max_rows', 100)\npd.set_option('display.max_columns', 100)\npd.set_option('display.width', 1000)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-09T11:04:09.604284Z","iopub.execute_input":"2026-02-09T11:04:09.604685Z","iopub.status.idle":"2026-02-09T11:04:09.609127Z","shell.execute_reply.started":"2026-02-09T11:04:09.604658Z","shell.execute_reply":"2026-02-09T11:04:09.608241Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"training_tomograms = glob(str(competition_dataset_directory / 'train' / '*'))\ntraining_images = glob(str(competition_dataset_directory / 'train' / '*' / '*.jpg'))\ndf_train = pd.read_csv(competition_dataset_directory / 'train_labels.csv')\n\nprint(f'Counting of Training Tomograms {len(training_tomograms)}')\nprint(f'Counting of Training Images {len(training_images)}')\nprint(f'Training Set Shape: {df_train.shape}')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-09T11:04:09.797370Z","iopub.execute_input":"2026-02-09T11:04:09.797688Z","iopub.status.idle":"2026-02-09T11:04:11.064247Z","shell.execute_reply.started":"2026-02-09T11:04:09.797666Z","shell.execute_reply":"2026-02-09T11:04:11.063413Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 2. Introduction\n\n### Cryogenic electron tomography (cryo-ET)\n\nCryogenic Electron Tomography (cryo-ET) is an advanced imaging technique used to create high-resolution (10–40 Å) 3D models of tiny structures, such as biological molecules and cells. It works by tilting a frozen sample under a transmission electron microscope (Cryo-TEM) to capture multiple 2D images from different angles. These images are then combined to form a detailed 3D reconstruction, similar to how a CT scan works for the human body. Unlike other electron microscopy methods, cryo-ET keeps samples at extremely low temperatures (< -150°C), preserving their natural state in a thin layer of vitreous ice without dehydration or chemical damage.\n\n![cryo_et](https://i.ibb.co/0yk5JwtX/Screenshot-from-2025-03-07-20-05-59.png)\n\n### Flagellar Motors\n\nA flagellar motor is a tiny, natural molecular machine that helps bacteria move. It works like a biological rotary engine, spinning a long, whip-like structure called a flagellum to push the bacterium forward. The motor is embedded in the bacterial cell wall and powered by ions (like protons). It rotates the flagellum at high speeds, allowing the bacterium to swim toward food or escape harmful environments. This movement is crucial for processes like chemotaxis (moving toward or away from chemicals) and even bacterial infections.\n\n![flagellar_motor](https://i.ibb.co/q398w2T7/Screenshot-from-2025-03-07-19-58-39.png)\n\nThe problem is to identify the presence and location of flagellar motors in 3D tomograms of bacteria, reconstructed from stacks of 2D slices. The training dataset consists of **648** tomograms, totaling **269,194** images, with a structured label set of **737** entries. These slices represent different cross-sections of the bacteria, and when stacked together, they form a 3D tomogram.","metadata":{}},{"cell_type":"code","source":"display(df_train)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-02T19:04:37.267375Z","iopub.execute_input":"2026-02-02T19:04:37.267733Z","iopub.status.idle":"2026-02-02T19:04:37.311641Z","shell.execute_reply.started":"2026-02-02T19:04:37.267693Z","shell.execute_reply":"2026-02-02T19:04:37.310790Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 3. Labels\n\nIn this competition, the labels correspond to the central coordinates of flagellar motors within a 3D tomogram, expressed as (z, y, x) positions. The evaluation is based on Euclidean distance, meaning a motor is considered correctly detected if the predicted coordinates lie within a 1000 angstrom (Å) radius of the true center. In essence, each motor’s label defines a spherical region, and predictions inside this region are treated as true positives. It is important to note that flagellar motors are not perfect spheres—they are more cylindrical in shape, and their orientation in the tomogram is unknown, which adds complexity to the detection task. Tomograms without any flagellar motor are labeled as (-1, -1, -1) to indicate the absence of a motor.","metadata":{}},{"cell_type":"code","source":"def visualize_counts(df, value, title):\n        \n    fig, ax = plt.subplots(figsize=(26, 18))\n    ax.barh(\n        y=np.arange(df.shape[0]),\n        width=df['count'].values,\n        align='center'\n    )\n    ax.set_yticks(np.arange(df.shape[0]))\n    ax.set_yticklabels([\n        f'{val} - Count: {count} ({percentage:.2f}%)' for val, count, percentage in zip(\n            df[value].values,\n            df['count'].values,\n            df['percentage'].values,\n        )\n    ])\n    ax.set_xlabel('')\n    ax.tick_params(axis='x', labelsize=17.5, pad=10)\n    ax.tick_params(axis='y', labelsize=17.5, pad=10)\n    ax.set_title(title, size=20, pad=15)\n    plt.gca().invert_yaxis()\n    plt.show()\n","metadata":{"trusted":true,"_kg_hide-input":true,"execution":{"iopub.status.busy":"2026-02-02T19:04:37.312551Z","iopub.execute_input":"2026-02-02T19:04:37.312863Z","iopub.status.idle":"2026-02-02T19:04:37.319079Z","shell.execute_reply.started":"2026-02-02T19:04:37.312812Z","shell.execute_reply":"2026-02-02T19:04:37.318411Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"In the training set, most tomograms contain either zero or one flagellar motor. Specifically, 48.3% (313 tomograms) have no motors, and 44.1% (286 tomograms) contain a single motor. A smaller portion of tomograms include multiple motors: 30 tomograms (4.6%) have two motors, and even fewer contain three or more. The most extreme instances include one tomogram with 10 motors (0.15%), along with a few cases having four to six motors. This distribution indicates that while tomograms with a single motor are prevalent, special attention is needed for detecting multiple motors. The test set, in contrast, contains only tomograms with zero or one motor.","metadata":{}},{"cell_type":"code","source":"df_train_n_motors = df_train.groupby('tomo_id')['Number of motors'].first().reset_index(drop=True).value_counts().reset_index()\ndf_train_n_motors['percentage'] = df_train_n_motors['count'] / df_train_n_motors['count'].sum(axis=0) * 100\n\nvisualize_counts(\n    df=df_train_n_motors,\n    value='Number of motors',\n    title='Value count motors number'\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-02T19:04:37.319820Z","iopub.execute_input":"2026-02-02T19:04:37.320149Z","iopub.status.idle":"2026-02-02T19:04:37.824246Z","shell.execute_reply.started":"2026-02-02T19:04:37.320120Z","shell.execute_reply":"2026-02-02T19:04:37.823417Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Some tomograms are indicated as having two motors, yet the training data provides only a single label. This points to potential inconsistencies in the labeling, where one of the expected motor positions may be missing:","metadata":{}},{"cell_type":"code","source":"df_train['tomogram_row_count'] = df_train.groupby('tomo_id')['tomo_id'].transform('count')\ndf_missing_labels = df_train.loc[(df_train['tomogram_row_count'] != df_train['Number of motors']) & (df_train['Number of motors'] != 0)]\ndisplay(df_missing_labels)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-02T19:04:37.825328Z","iopub.execute_input":"2026-02-02T19:04:37.825596Z","iopub.status.idle":"2026-02-02T19:04:37.842532Z","shell.execute_reply.started":"2026-02-02T19:04:37.825576Z","shell.execute_reply":"2026-02-02T19:04:37.841711Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The coordinates of flagellar motors in the training set exhibit variability along all three axes. Along the z-axis (Motor axis 0), values range from 0 to 466 with a mean of 167, while the y-axis (Motor axis 1) spans 59 to 904 with an average of 482, and the x-axis (Motor axis 2) ranges from 30 to 902 with a mean of 492. The z-axis shows the smallest standard deviation (78.1), indicating that motors are more tightly clustered along this axis, whereas the y and x axes display wider spreads, with standard deviations of 198.6 and 214.8, respectively.","metadata":{}},{"cell_type":"code","source":"def visualize_histogram(df, columns, title):\n\n    fig, ax = plt.subplots(figsize=(28, 7))\n    for column in columns:\n        ax.hist(df[column], 32, alpha=0.5, label=column)\n    ax.tick_params(axis='x', labelsize=14)\n    ax.tick_params(axis='y', labelsize=14)\n    ax.set_title(title, size=18, pad=15)\n    ax.legend(prop={'size': 15})\n    plt.show()\n\n\ndf_train_with_motors = df_train.loc[df_train['Motor axis 0'] != -1].reset_index(drop=True)\n\nvisualize_histogram(\n    df_train_with_motors,\n    columns=['Motor axis 0', 'Motor axis 1', 'Motor axis 2'],\n    title='Label Coordinate Histograms'\n)\n\ndisplay(df_train_with_motors[['Motor axis 0', 'Motor axis 1', 'Motor axis 2']].describe())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-02T19:04:37.845245Z","iopub.execute_input":"2026-02-02T19:04:37.845477Z","iopub.status.idle":"2026-02-02T19:04:38.254767Z","shell.execute_reply.started":"2026-02-02T19:04:37.845452Z","shell.execute_reply":"2026-02-02T19:04:38.254027Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The tomogram shapes in the training set show low diversity, with only a few distinct variations. The z-axis (depth) exhibits the greatest variability, ranging from 300 to 800, while the y-axis (height) and x-axis (width) are relatively consistent, with means of 950 and 954 and standard deviations of 64.9 and 97.2, respectively. The most common tomogram shapes cluster around (300, 928, 928) and (500, 960, 956), though a few outliers reach as high as (800, 1912, 1847). This indicates that many tomograms share identical or very similar dimensions.","metadata":{"_kg_hide-input":false}},{"cell_type":"code","source":"visualize_histogram(\n    df_train,\n    columns=['Array shape (axis 0)', 'Array shape (axis 1)', 'Array shape (axis 2)'],\n    title='Shape of Tomogram Histograms'\n)\n\ndisplay(df_train[['Array shape (axis 0)', 'Array shape (axis 1)', 'Array shape (axis 2)']].describe())","metadata":{"trusted":true,"_kg_hide-input":true,"execution":{"iopub.status.busy":"2026-02-02T19:04:38.256271Z","iopub.execute_input":"2026-02-02T19:04:38.256527Z","iopub.status.idle":"2026-02-02T19:04:38.715822Z","shell.execute_reply.started":"2026-02-02T19:04:38.256506Z","shell.execute_reply":"2026-02-02T19:04:38.714743Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The normalized distributions of label coordinates reveal the typical locations of flagellar motors within the tomograms. On average, motors are found at approximately 46.7% of the z-depth, 50.9% of the y-height, and 52.5% of the x-width, showing a slight central tendency, particularly along the y and x axes. Nevertheless, there is considerable variability, with standard deviations of about 16.5% (z), 21.0% (y), and 22.8% (x). The minimum values indicate that some motors are positioned very close to the edges of the tomograms, while the maximum values exceed 90% along each axis, confirming that motors can appear throughout the entire tomogram volume.","metadata":{}},{"cell_type":"code","source":"df_train_with_motors['Normalized Motor axis 0'] = df_train_with_motors['Motor axis 0'] / df_train_with_motors['Array shape (axis 0)']\ndf_train_with_motors['Normalized Motor axis 1'] = df_train_with_motors['Motor axis 1'] / df_train_with_motors['Array shape (axis 1)']\ndf_train_with_motors['Normalized Motor axis 2'] = df_train_with_motors['Motor axis 2'] / df_train_with_motors['Array shape (axis 2)']","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-02T19:04:38.716782Z","iopub.execute_input":"2026-02-02T19:04:38.717369Z","iopub.status.idle":"2026-02-02T19:04:38.723008Z","shell.execute_reply.started":"2026-02-02T19:04:38.717337Z","shell.execute_reply":"2026-02-02T19:04:38.722161Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"visualize_histogram(\n    df_train_with_motors,\n    columns=['Normalized Motor axis 0', 'Normalized Motor axis 1', 'Normalized Motor axis 2'],\n    title='Normalized Label Coordinate Histograms'\n)\n\ndisplay(df_train_with_motors[['Normalized Motor axis 0', 'Normalized Motor axis 1', 'Normalized Motor axis 2']].describe())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-02T19:04:38.723435Z","iopub.execute_input":"2026-02-02T19:04:38.723686Z","iopub.status.idle":"2026-02-02T19:04:39.118767Z","shell.execute_reply.started":"2026-02-02T19:04:38.723658Z","shell.execute_reply":"2026-02-02T19:04:39.118127Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The heatmaps of label coordinates offer a clearer picture of how flagellar motors are distributed across different tomographic planes. In the YX plane (top-down view), motors appear fairly scattered across the width and height, indicating that their placement along these axes is relatively unconstrained. In contrast, the ZY (side view) and ZX (front view) planes show a stronger concentration near the center, suggesting that motors are more likely to be located around the midsections of the tomograms along the depth (Z) axis. This slight central tendency may reflect biological factors affecting motor positioning or potential biases in how the tomograms were captured and annotated.","metadata":{}},{"cell_type":"code","source":"def visualize_heatmap(heatmap, title):\n\n    fig, ax = plt.subplots(figsize=(9, 9))\n    ax.imshow(heatmap, cmap=plt.cm.hot)\n    ax.set_xlabel('')\n    ax.set_ylabel('')\n    ax.tick_params(axis='x', labelsize=16, pad=10)\n    ax.tick_params(axis='y', labelsize=16, pad=10)\n    ax.set_title(title, size=16, pad=13.5, loc='center', wrap=True)\n    plt.show()\n\n\nyx_coordinate_heatmap = np.zeros((1000, 1000))\n\nfor _, row in df_train_with_motors.iterrows():\n    yx_coordinate_heatmap[\n        int(row['Normalized Motor axis 1'] * 1000) - 3:int(row['Normalized Motor axis 1'] * 1000) + 3,\n        int(row['Normalized Motor axis 2'] * 1000) - 3:int(row['Normalized Motor axis 2'] * 1000) + 3,\n        \n    ] += 1\n\nvisualize_heatmap(heatmap=yx_coordinate_heatmap, title='Label Coordinate Heatmap on YX Plane')\n\nzy_coordinate_heatmap = np.zeros((1000, 1000))\n\nfor _, row in df_train_with_motors.iterrows():\n    zy_coordinate_heatmap[\n        int(row['Normalized Motor axis 0'] * 1000) - 3:int(row['Normalized Motor axis 0'] * 1000) + 3,\n        int(row['Normalized Motor axis 1'] * 1000) - 3:int(row['Normalized Motor axis 1'] * 1000) + 3,\n        \n    ] += 1\n\nvisualize_heatmap(heatmap=zy_coordinate_heatmap, title='Label Coordinate Heatmap on ZY Plane')\n\nzx_coordinate_heatmap = np.zeros((1000, 1000))\n\nfor _, row in df_train_with_motors.iterrows():\n    zx_coordinate_heatmap[\n        int(row['Normalized Motor axis 0'] * 1000) - 3:int(row['Normalized Motor axis 0'] * 1000) + 3,\n        int(row['Normalized Motor axis 2'] * 1000) - 3:int(row['Normalized Motor axis 2'] * 1000) + 3,\n        \n    ] += 1\n\nvisualize_heatmap(heatmap=zx_coordinate_heatmap, title='Label Coordinate Heatmap on ZX Plane')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-02T19:04:39.119520Z","iopub.execute_input":"2026-02-02T19:04:39.119772Z","iopub.status.idle":"2026-02-02T19:04:40.216218Z","shell.execute_reply.started":"2026-02-02T19:04:39.119750Z","shell.execute_reply":"2026-02-02T19:04:40.215289Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 4. Real World Units\nThe dataset uses angstroms (Å) as the unit of measurement, which is common in structural biology and microscopy. One angstrom (1 Å) equals 0.1 nanometers (nm), making it suitable for describing molecular and subcellular structures. The voxel spacing column specifies the tomogram resolution in angstroms per voxel, allowing pixel-based measurements to be converted into real-world physical dimensions. This enables meaningful analysis of distances, areas, and volumes, and allows comparisons across tomograms with varying resolutions. Furthermore, the competition metric defines a 1000 Å radius around each labeled motor coordinate to determine true positive detections.","metadata":{}},{"cell_type":"markdown","source":"Voxel spacing and angstroms per voxel represent the same concept, defining the real-world physical size of each voxel in the tomograms. They indicate how many angstroms correspond to a single voxel along each axis. In the dataset, voxel spacing values are concentrated around a few common numbers rather than being evenly distributed. The average voxel spacing is 15.07 Å, with most values ranging between 13.1 Å and 16.1 Å. The most frequent values—13.1 Å, 15.6 Å, and 19.7 Å—suggest that the tomograms were likely acquired using a few specific settings. The minimum and maximum recorded spacings are 6.5 Å and 19.7 Å, respectively. Because voxel spacing determines how large objects appear in real-world units, these clusters are important when comparing flagellar motor sizes across different tomograms.","metadata":{}},{"cell_type":"code","source":"df_train_voxel_spacings = df_train.groupby('tomo_id')['Voxel spacing'].first().reset_index(drop=True).value_counts().reset_index()\ndf_train_voxel_spacings['percentage'] = df_train_voxel_spacings['count'] / df_train_voxel_spacings['count'].sum(axis=0) * 100\n\nvisualize_counts(\n    df=df_train_voxel_spacings,\n    value='Voxel spacing',\n    title='Voxel spacing Value Counts'\n)\n\nvisualize_histogram(\n    df_train.groupby('tomo_id')['Voxel spacing'].first().reset_index(),\n    columns=['Voxel spacing'],\n    title='Voxel Spacing Histogram'\n)\n\ndisplay(df_train.groupby('tomo_id')[['Voxel spacing']].first().describe())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-02T19:04:40.217387Z","iopub.execute_input":"2026-02-02T19:04:40.217741Z","iopub.status.idle":"2026-02-02T19:04:40.822908Z","shell.execute_reply.started":"2026-02-02T19:04:40.217709Z","shell.execute_reply":"2026-02-02T19:04:40.822251Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The total number of voxels in a tomogram represents the complete 3D grid of discrete volume elements used to capture structural details. Each voxel corresponds to a small cubic region of space, and the overall count depends on the tomogram’s dimensions along the X, Y, and Z axes.\n\nAccording to the dataset statistics, the average tomogram contains roughly 382 million voxels, with a standard deviation of 197 million, reflecting considerable variation in tomogram sizes. The smallest tomogram has 258 million voxels, while the largest approaches 1.77 billion voxels, illustrating the wide range of spatial resolutions and fields of view among samples. Approximately 50% of tomograms have fewer than 267 million voxels, indicating a common size pattern, while the upper quartile extends to 442 million voxels. This variability means that models need to efficiently handle tomograms of different sizes, as voxel counts directly affect both computational requirements and the amount of information available for detecting flagellar motors.","metadata":{}},{"cell_type":"code","source":"df_train['total_voxels'] = df_train['Array shape (axis 0)'] * df_train['Array shape (axis 1)'] * df_train['Array shape (axis 2)']","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-02T19:04:40.823676Z","iopub.execute_input":"2026-02-02T19:04:40.823975Z","iopub.status.idle":"2026-02-02T19:04:40.828467Z","shell.execute_reply.started":"2026-02-02T19:04:40.823939Z","shell.execute_reply":"2026-02-02T19:04:40.827678Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"visualize_histogram(\n    df_train.groupby('tomo_id')['total_voxels'].first().reset_index(),\n    columns=['total_voxels'],\n    title='Total Voxels Histogram'\n)\n\ndisplay(df_train.groupby('tomo_id')[['total_voxels']].first().describe())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-02T19:04:40.829153Z","iopub.execute_input":"2026-02-02T19:04:40.829329Z","iopub.status.idle":"2026-02-02T19:04:41.132450Z","shell.execute_reply.started":"2026-02-02T19:04:40.829313Z","shell.execute_reply":"2026-02-02T19:04:41.131754Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Examining the required voxel proximity is important because it determines how precise a model’s predictions must be in voxel space. A prediction is considered correct if it falls within a 1000 Å radius, but the number of voxels within this distance depends on the voxel spacing of each tomogram. Lower voxel spacing results in more voxels within 1000 Å, increasing the voxel proximity value, whereas higher spacing decreases it.\n\nThe histogram shows that most tomograms require predictions to be within 50 to 76 voxels to be considered correct, with a median of 64 voxels. The mean is 67.8, indicating a slight skew toward higher values. The minimum value is 50.76, meaning the least dense tomograms require roughly 51 voxels, while the maximum of 153.85 corresponds to tomograms with finer resolution, needing many more voxels within the 1000 Å range.\n\nThis variation in required voxel proximity underscores a key challenge: models must handle different voxel resolutions across tomograms, adjusting predictions to remain within the correct voxel range for each scale.","metadata":{}},{"cell_type":"code","source":"df_train['required_voxel_proximity'] = 1000 / df_train['Voxel spacing']","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-02T19:04:41.133253Z","iopub.execute_input":"2026-02-02T19:04:41.133539Z","iopub.status.idle":"2026-02-02T19:04:41.137798Z","shell.execute_reply.started":"2026-02-02T19:04:41.133516Z","shell.execute_reply":"2026-02-02T19:04:41.137168Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"visualize_histogram(\n    df_train.groupby('tomo_id')['required_voxel_proximity'].first().reset_index(),\n    columns=['required_voxel_proximity'],\n    title='Required Voxel Proximity Histogram'\n)\n\ndisplay(df_train.groupby('tomo_id')[['required_voxel_proximity']].first().describe())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-02T19:04:41.138655Z","iopub.execute_input":"2026-02-02T19:04:41.138985Z","iopub.status.idle":"2026-02-02T19:04:41.407494Z","shell.execute_reply.started":"2026-02-02T19:04:41.138955Z","shell.execute_reply":"2026-02-02T19:04:41.406825Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The valid voxel count indicates the number of voxels in each tomogram that fall within the 1,000-angstrom proximity threshold around a motor’s center. This metric provides an estimate of the potential search space for a correct prediction, as any prediction within this region is considered valid.\n\nOn average, a tomogram contains 1.41 million valid voxels, though there is considerable variation, with a standard deviation of 1.06 million. The smallest valid voxel count is approximately 548,000, while the largest reaches 15.25 million, reflecting differences in voxel spacing and tomogram dimensions. Half of the tomograms have fewer than 1.1 million valid voxels, and the upper quartile extends to 1.86 million, indicating that some tomograms require handling much larger search areas. The presence of a few tomograms with exceptionally high valid voxel counts highlights the need for models capable of adapting to varying tomogram sizes while maintaining detection precision.","metadata":{}},{"cell_type":"code","source":"df_train['valid_voxels'] = ((4 / 3) * np.pi * (df_train['required_voxel_proximity'] ** 3)).astype(int)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-02T19:04:41.408382Z","iopub.execute_input":"2026-02-02T19:04:41.408728Z","iopub.status.idle":"2026-02-02T19:04:41.413546Z","shell.execute_reply.started":"2026-02-02T19:04:41.408686Z","shell.execute_reply":"2026-02-02T19:04:41.412931Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"visualize_histogram(\n    df_train.groupby('tomo_id')['valid_voxels'].first().reset_index(),\n    columns=['valid_voxels'],\n    title='Valid Voxels Histogram'\n)\n\ndisplay(df_train.groupby('tomo_id')[['valid_voxels']].first().describe())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-02T19:04:41.414542Z","iopub.execute_input":"2026-02-02T19:04:41.414896Z","iopub.status.idle":"2026-02-02T19:04:41.702372Z","shell.execute_reply.started":"2026-02-02T19:04:41.414842Z","shell.execute_reply":"2026-02-02T19:04:41.701632Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The valid voxel count represents the number of voxels within a tomogram that lie inside the 1,000-angstrom proximity threshold around a motor’s center. This serves as an estimate of the search space for a correct prediction, since any prediction within this region is considered valid.\n\nOn average, tomograms contain 1.41 million valid voxels, with considerable variation (standard deviation of 1.06 million). The smallest tomogram has around 548,000 valid voxels, while the largest reaches 15.25 million, reflecting differences in voxel spacing and tomogram dimensions. Half of the tomograms have fewer than 1.1 million valid voxels, with the upper quartile extending to 1.86 million, showing that some tomograms require much larger search areas. The presence of a few tomograms with exceptionally high valid voxel counts underscores the need for models that can adapt to varying tomogram sizes while maintaining detection accuracy.","metadata":{}},{"cell_type":"code","source":"df_train['valid_voxel_percentage'] = df_train['valid_voxels'] / df_train['total_voxels'] * 100","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-02T19:04:41.703152Z","iopub.execute_input":"2026-02-02T19:04:41.703456Z","iopub.status.idle":"2026-02-02T19:04:41.708199Z","shell.execute_reply.started":"2026-02-02T19:04:41.703433Z","shell.execute_reply":"2026-02-02T19:04:41.707501Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"visualize_histogram(\n    df_train.groupby('tomo_id')['valid_voxel_percentage'].first().reset_index(),\n    columns=['valid_voxel_percentage'],\n    title='Valid Voxel Percentage Histogram'\n)\n\ndisplay(df_train.groupby('tomo_id')[['valid_voxel_percentage']].first().describe())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-02T19:04:41.708958Z","iopub.execute_input":"2026-02-02T19:04:41.709148Z","iopub.status.idle":"2026-02-02T19:04:42.018381Z","shell.execute_reply.started":"2026-02-02T19:04:41.709131Z","shell.execute_reply":"2026-02-02T19:04:42.017438Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Analyzing the correlations between voxel spacing, total voxel count, and the percentage of valid voxels provides additional insight. There is a strong negative correlation (-0.68) between voxel spacing and valid voxel percentage, indicating that as voxel spacing increases, the proportion of useful voxels decreases. This is expected, since larger voxel spacing reduces the number of voxels falling within the 1,000-angstrom proximity threshold. Total voxel count and valid voxel percentage also show a negative correlation (-0.46), suggesting that larger tomograms generally have a smaller proportion of valid voxels. In contrast, the correlation between voxel spacing and total voxel count is weaker (-0.20), implying that larger tomograms do not necessarily have markedly different voxel spacings. These relationships highlight the critical role of voxel resolution in determining how much of a tomogram is usable for flagellar motor localization.","metadata":{}},{"cell_type":"code","source":"df_train.groupby('tomo_id')[['Voxel spacing', 'total_voxels', 'valid_voxel_percentage']].first().corr()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-02T19:04:42.021581Z","iopub.execute_input":"2026-02-02T19:04:42.021819Z","iopub.status.idle":"2026-02-02T19:04:42.033962Z","shell.execute_reply.started":"2026-02-02T19:04:42.021799Z","shell.execute_reply":"2026-02-02T19:04:42.033124Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 5. Tomograms\n\nTomograms used in cryo-electron tomography (cryo-ET) are typically composed of 8-bit grayscale images, where each voxel intensity represents electron density. These images are captured as a series of 2D slices that, when stacked together, form a 3D reconstruction of a biological sample at nanometer resolution. Since the images are grayscale, each pixel has an intensity ranging from 0 to 255, with lower values representing less-dense regions and higher values indicating denser structures.","metadata":{}},{"cell_type":"code","source":"def visualize_slice_with_annotations(df, tomo_id, slice_idx):\n\n    image_paths = sorted(glob(str(competition_dataset_directory / 'train' / tomo_id / '*')))\n    image = cv2.imread(image_paths[slice_idx], cv2.IMREAD_GRAYSCALE)\n    image_rgb = cv2.cvtColor(image, cv2.COLOR_GRAY2RGB)\n\n    tomo_mask = df['tomo_id'] == tomo_id\n    motor_center_coordinates = df.loc[tomo_mask, ['Motor axis 0', 'Motor axis 1', 'Motor axis 2']].values\n    voxel_spacing = df.loc[tomo_mask, 'Voxel spacing'].values[0]\n\n    fig, ax = plt.subplots(figsize=(8, 8))\n    \n    if len(motor_center_coordinates) == 1 and motor_center_coordinates.sum() == -3:\n        ax.imshow(image_rgb)\n    else:\n        overlay = image_rgb.copy()\n        alpha = 0.25\n        motor_radius = int(round(1000 / voxel_spacing))\n    \n        for z, y, x in motor_center_coordinates:\n            \n            depth_diff = abs(z - slice_idx)\n            if depth_diff <= motor_radius:\n                visible_radius = int(round((motor_radius ** 2 - depth_diff ** 2) ** 0.5))\n                cv2.circle(image_rgb, (int(x), int(y)), visible_radius, (255, 0, 0), -1)\n                cv2.circle(image_rgb, (int(x), int(y)), visible_radius, (0, 255, 0), 2)\n                annotation_text = f'({int(x)}, {int(y)}, {int(z)}) r: {visible_radius:.2f}'\n                cv2.putText(\n                    image_rgb,\n                    annotation_text,\n                    (int(x) + 25, int(y) - 10),\n                    cv2.FONT_HERSHEY_SIMPLEX,\n                    0.65,\n                    (0, 0, 255),\n                    2,\n                    cv2.LINE_AA\n                )\n\n        cv2.addWeighted(overlay, alpha, image_rgb, 1 - alpha, 0, image_rgb)\n        ax.imshow(image_rgb)\n        \n    ax.axis('off')\n    ax.set_title(f'{tomo_id} - Slice: {slice_idx} - Voxel Spacing: {voxel_spacing}\\nImage Shape: {image.shape}', size=15, pad=12.5, loc='center', wrap=True)\n    plt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-02T19:04:42.035129Z","iopub.execute_input":"2026-02-02T19:04:42.035411Z","iopub.status.idle":"2026-02-02T19:04:42.043696Z","shell.execute_reply.started":"2026-02-02T19:04:42.035391Z","shell.execute_reply":"2026-02-02T19:04:42.042879Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The tomogram **tomo_003acc** has the lowest voxel spacing (**6.5**) among all the samples, meaning that each voxel represents a smaller physical distance compared to other tomograms. As a result, objects within this tomogram appear significantly larger when visualized, even though their true physical size remains the same. This effect occurs because the same number of pixels now covers a much smaller real-world volume, effectively increasing the apparent size of structures such as the motors. The high level of detail in similar tomograms may provide more precise localization of objects, but it also requires careful interpretation to avoid misjudging relative sizes when comparing across different tomograms.","metadata":{}},{"cell_type":"code","source":"visualize_slice_with_annotations(df_train, 'tomo_003acc', slice_idx=200)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-02T19:04:42.044609Z","iopub.execute_input":"2026-02-02T19:04:42.044929Z","iopub.status.idle":"2026-02-02T19:04:42.855475Z","shell.execute_reply.started":"2026-02-02T19:04:42.044883Z","shell.execute_reply":"2026-02-02T19:04:42.854604Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The tomogram tomo_00e047 has a voxel spacing of 15.6 Å, which is close to the dataset’s average. This indicates that the objects within the tomogram appear at a typical scale—neither unusually large nor too small relative to other samples. At this voxel spacing, the visualization offers a balanced level of detail, providing sufficient resolution to distinguish structures without the exaggerated magnification seen in tomograms with lower spacing. Consequently, this tomogram serves as a representative example for examining object appearances under standard imaging conditions.","metadata":{}},{"cell_type":"code","source":"visualize_slice_with_annotations(df_train, 'tomo_00e047', slice_idx=200)","metadata":{"trusted":true,"_kg_hide-input":true,"execution":{"iopub.status.busy":"2026-02-02T19:04:42.856621Z","iopub.execute_input":"2026-02-02T19:04:42.856980Z","iopub.status.idle":"2026-02-02T19:04:43.331391Z","shell.execute_reply.started":"2026-02-02T19:04:42.856949Z","shell.execute_reply":"2026-02-02T19:04:43.330398Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The tomogram tomo_00e463 has a voxel spacing of 19.7 Å, the highest in the dataset. This larger spacing causes objects to appear smaller compared to tomograms with lower voxel spacing. Despite this, it contains 6 motors, one of the highest counts in the dataset. The combination of high voxel spacing and multiple motors poses a unique challenge for visualization and analysis, as the motors appear more condensed and finer structural details may be harder to distinguish without proper scaling.","metadata":{}},{"cell_type":"code","source":"visualize_slice_with_annotations(df_train, 'tomo_00e463', slice_idx=250)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-02-02T19:04:43.332474Z","iopub.execute_input":"2026-02-02T19:04:43.332740Z","iopub.status.idle":"2026-02-02T19:04:43.788542Z","shell.execute_reply.started":"2026-02-02T19:04:43.332718Z","shell.execute_reply":"2026-02-02T19:04:43.787554Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null}]}