{"metadata":{"kernelspec":{"name":"python3","display_name":"Python 3","language":"python"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":91249,"databundleVersionId":11294684,"sourceType":"competition"},{"sourceId":10974966,"sourceType":"datasetVersion","datasetId":6829219}],"dockerImageVersionId":30918,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# <p style=\"font-family:JetBrains Mono; font-weight:normal; letter-spacing: 2px; color:#1AAB70; font-size:140%; text-align:left;padding: 0px; border-bottom: 3px solid #1AAB70\">Libraries</p>","metadata":{}},{"cell_type":"code","source":"from IPython.display import display, HTML\nfrom IPython.display import Image as Im\n\nimport os\n\nfrom glob import glob\nfrom tqdm.auto import tqdm\ntqdm.pandas()\n\n\nimport numpy as np\nimport pandas as pd\nfrom pandas.io.formats.style import Styler\n\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nfrom PIL import Image\n\n\nimport warnings \nwarnings.filterwarnings('ignore')\n\nfrom colorama import Style, Fore\n\nRC = {\n    \"axes.facecolor\": \"#F8F8F8\",\n    \"figure.facecolor\": \"#F8F8F8\",\n    \"axes.edgecolor\": \"#000000\",\n    \"grid.color\": \"#EBEBE7\" + \"30\",\n    \"font.family\": \"serif\",\n    \"axes.labelcolor\": \"#000000\",\n    \"xtick.color\": \"#000000\",\n    \"ytick.color\": \"#000000\",\n    \"grid.alpha\": 0.4\n}\n\n\nPALETTE = ['#302c36', '#037d97', '#91013E', '#C09741',\n           '#EC5B6D', '#90A6B1', '#6ca957', '#D8E3E2']\n\n\nsns.set(rc=RC)\n\nclass ColorStyle:\n    def __init__(self):\n        self.red = Style.BRIGHT + Fore.RED\n        self.blk = Style.BRIGHT + Fore.BLACK\n        self.gld = Style.BRIGHT + Fore.YELLOW\n        self.grn = Style.BRIGHT + Fore.GREEN\n        self.cya = Style.BRIGHT + Fore.CYAN\n        self.mgt = Style.BRIGHT + Fore.MAGENTA\n        self.blu = Style.BRIGHT + Fore.BLUE\n        self.cya_v2 = self._init_new_color(30, 255, 255)  # 153, 229, 255\n        self.res = Style.RESET_ALL\n\n    @staticmethod\n    def _init_new_color(r, g, b, background=False):\n        return f'\\033[{\"48\" if background else \"38\"};2;{r};{g};{b}m'\n\n    def color_from_rgb(self, r, g, b):\n        return self._init_new_color(r, g, b)\n\n\ncS = ColorStyle()","metadata":{"_kg_hide-input":true,"trusted":true,"execution":{"iopub.status.busy":"2025-03-10T01:52:03.946033Z","iopub.execute_input":"2025-03-10T01:52:03.946304Z","iopub.status.idle":"2025-03-10T01:52:06.652850Z","shell.execute_reply.started":"2025-03-10T01:52:03.946284Z","shell.execute_reply":"2025-03-10T01:52:06.651850Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<p align=\"right\">\n  <img src=\"https://api.monosnap.com/file/download?id=MkDfAf8fTGK1CZ4HdmwYTqvyWFYDhE\"/>\n</p>","metadata":{"execution":{"iopub.status.busy":"2025-03-10T01:36:56.133654Z","iopub.execute_input":"2025-03-10T01:36:56.133955Z","iopub.status.idle":"2025-03-10T01:36:56.138740Z","shell.execute_reply.started":"2025-03-10T01:36:56.133927Z","shell.execute_reply":"2025-03-10T01:36:56.137440Z"}}},{"cell_type":"markdown","source":"\n# <p style=\"font-family:JetBrains Mono; font-weight:normal; letter-spacing: 1px; color:#1AAB70; font-size:140%; text-align:left;padding: 0px; border-bottom: 3px solid #1AAB70\">Intro</p>\n\nThe bacterial flagellum is a complex, multiprotein nanomachine that allows bacteria to swim in liquid environments. Flagella from different species contain structurally conserved components. The cytoplasmic C-ring, which integrates chemotaxis signals, abuts the inner membrane MS-ring and interacts with stator proteins that anchor in the peptidoglycan and transform chemical energy into flagellar rotation (1)\n\n<p align=\"right\">\n  <img src=\"https://api.monosnap.com/file/download?id=lWpF3ZH6CnMWiHB66uw9Xk35opIb89\"/>\n</p>\n\n\n1. Minamino T, Imada K. 2015. The bacterial flagellar motor and its structural diversity. Trends Microbiol 23:267–274.\n2. Zhu SSchniederberend M, Zhitnitsky D, Jain R, Galán JE, Kazmierczak BI, Liu J2019.In Situ Structures of Polar and Lateral Flagella Revealed by Cryo-Electron Tomography. J Bacteriol201:10.1128/jb.00117-19.https://doi.org/10.1128/jb.00117-19","metadata":{}},{"cell_type":"markdown","source":"# <p style=\"font-family:JetBrains Mono; font-weight:normal; letter-spacing: 1px; color:#1AAB70; font-size:140%; text-align:left;padding: 0px; border-bottom: 3px solid #1AAB70\">Dataset Description</p>\n\nIn this competition, you are tasked with finding flagellar motor centers in 3D tomograms. A tomogram is a 3D volumetric representation of an object. Each tomogram is provided as a set of 2D image slices (JPEG) stored in a unique directory. Your goal is to predict points where a flagellar motor is located when one is present.\n\n## <p style=\"font-family:JetBrains Mono; font-weight:normal; letter-spacing: 1px; color:#1AAB70; font-size:140%; text-align:left;padding: 0px; border-bottom: 3px solid #1AAB70\">Files and Directories</p>\n\n## <p style=\"font-family:JetBrains Mono; font-weight:normal; letter-spacing: 1px; color:#1AAB70; font-size:120%; text-align:left;padding: 0px;\">train/</p>\nDirectory of subdirectories, each containing a stack of tomogram slices for training. Each tomogram subdirectory comprises JPEGs, where each JPEG is a 2D slice of a tomogram.\n\n## <p style=\"font-family:JetBrains Mono; font-weight:normal; letter-spacing: 1px; color:#1AAB70; font-size:120%; text-align:left;padding: 0px;\">train_labels.csv</p>\nTraining data labels. Each row represents a unique motor location, not a unique tomogram.\n\n- **`row_id`** - Index of the row  \n- **`tomo_id`** - Unique identifier of the tomogram (some tomograms have multiple motors)  \n- **`Motor axis 0`** - Z-coordinate of the motor (which slice it is located on)  \n- **`Motor axis 1`** - Y-coordinate of the motor  \n- **`Motor axis 2`** - X-coordinate of the motor  \n- **`Array shape axis 0`** - Z-axis length (number of slices in the tomogram)  \n- **`Array shape axis 1`** - Y-axis length (width of each slice)  \n- **`Array shape axis 2`** - X-axis length (height of each slice)  \n- **`Voxel spacing`** - Scaling of the tomogram (angstroms per voxel)  \n- **`Number of motors`** - Number of motors in the tomogram \n\n(each row represents a motor, so tomograms with multiple motors have several rows)\n\n## <p style=\"font-family:JetBrains Mono; font-weight:normal; letter-spacing: 1px; color:#1AAB70; font-size:120%; text-align:left;padding: 0px;\">test/</p>\nDirectory with 3 directories of dummy test tomograms; the rerun test dataset contains approximately 900 tomograms. **The test data only contain tomograms with one or zero motors**.\n\n## <p style=\"font-family:JetBrains Mono; font-weight:normal; letter-spacing: 1px; color:#1AAB70; font-size:120%; text-align:left;padding: 0px;\">sample_submission.csv</p>\nSample submission file in the correct format.  \n(If you predict that no motor exists, set `Motor axis 0`, `Motor axis 1`, and `Motor axis 2` to `-1`).","metadata":{}},{"cell_type":"markdown","source":"# <p style=\"font-family:JetBrains Mono; font-weight:normal; letter-spacing: 1px; color:#1AAB70; font-size:140%; text-align:left;padding: 0px; border-bottom: 3px solid #1AAB70\">Data Schema</p>","metadata":{}},{"cell_type":"code","source":"!apt-get install tree -qq","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-10T02:32:51.953819Z","iopub.execute_input":"2025-03-10T02:32:51.954251Z","iopub.status.idle":"2025-03-10T02:32:54.752548Z","shell.execute_reply.started":"2025-03-10T02:32:51.954224Z","shell.execute_reply":"2025-03-10T02:32:54.751258Z"},"_kg_hide-input":true,"_kg_hide-output":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"!tree -L 1 /kaggle/input/byu-locating-bacterial-flagellar-motors-2025\n!tree -L 2 /kaggle/input/byu-locating-bacterial-flagellar-motors-2025 | head -n 10","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-10T02:32:56.664462Z","iopub.execute_input":"2025-03-10T02:32:56.664817Z","iopub.status.idle":"2025-03-10T02:32:57.370865Z","shell.execute_reply.started":"2025-03-10T02:32:56.664790Z","shell.execute_reply":"2025-03-10T02:32:57.369887Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"There are **648** unique subdirectories in the `train` folder, each corresponding to a specific tomogram. Each train tomogram subdirectory contains `2D` slices, which can be stacked to reconstruct a `3D` tomogram.  \n\nIn total, there are **269,194** `.jpg` slices across all train tomogram subdirectories. Some `3D` tomograms can be storage-intensive, reaching up to **1.7GB** in size. When unpacked, the total volume of all train examples amounts to approximately **236GB** (`uint8` format).  \n\n---\n\n#### **Explanation and Schema**  \n\nThe dataset consists of standard **train** and **test** folders. The **test** folder contains three directories of dummy test tomograms, while the actual rerun test dataset includes approximately **900** tomograms. **Test data only contains tomograms with either one or zero motors.**  \n\nThe **train** folder contains **648** tomograms, each sliced into `.jpg` files. The `train_labels.csv` file provides the `z, y, x` coordinates of the target of interest. If no target is present in a tomogram, the coordinates are set to `-1, -1, -1`. Similarly, during submission, `-1, -1, -1` should be used when predicting `0` targets.  \n\n\n#### **Below** is the descriptive data schema and a visualization of a tomogram: \n","metadata":{}},{"cell_type":"markdown","source":"<p align=\"right\">\n  <img src=\"https://api.monosnap.com/file/download?id=lSegY3jtLWIdysgn4yiYFQvJvdbBW1\"/>\n</p>","metadata":{"execution":{"iopub.status.busy":"2025-03-10T01:47:01.927681Z","iopub.execute_input":"2025-03-10T01:47:01.927953Z","iopub.status.idle":"2025-03-10T01:47:01.933054Z","shell.execute_reply.started":"2025-03-10T01:47:01.927932Z","shell.execute_reply":"2025-03-10T01:47:01.932002Z"}}},{"cell_type":"markdown","source":"This notebook provides a code snippet of how effectively display dataframes in a tidy format using Pandas Styler class.\n\nWe are going to leverage CSS styling language to manipulate many parameters including colors, fonts, borders, background, format and make our tables interactive.\n\n\n**Reference**:\n* [Pandas Table Visualization](https://pandas.pydata.org/pandas-docs/stable/user_guide/style.html).\n\n* At this point, you will require a certain level of understanding in web development.\n\n---\nPrimarily, you will have to modify the CSS of the `td`, `tr`, and `th` tags.\n\n**You can refer** to the following materials to learn **HTML/CSS**:\n* [w3schools HTML Tutorial](https://www.w3schools.com/html/default.asp)\n* [w3schools CSS Reference](https://www.w3schools.com/cssref/index.php)","metadata":{}},{"cell_type":"code","source":"def magnify(is_test: bool = False):\n    base_color = '#1AAB70'\n    if is_test:\n        highlight_target_row = []\n    else:\n        highlight_target_row = [dict(selector='tr:last-child',\n                                     props=[('background-color', f'{base_color}'+'20')])]\n\n    return [dict(selector=\"th\",\n                 props=[(\"font-size\", \"11pt\"),\n                        ('background-color', f'{base_color}'),\n                        ('color', 'white'),\n                        ('font-weight', 'bold'),\n                        ('border-bottom', '0.1px solid white'),\n                        ('border-left', '0.1px solid white'),\n                        ('text-align', 'right')]),\n\n            dict(selector='th.blank.level0', \n                props=[('font-weight', 'bold'),\n                       ('border-left', '1.7px solid white'),\n                       ('background-color', 'white')]),\n\n            dict(selector=\"td\",\n                 props=[('padding', \"0.5em 1em\"),\n                        ('text-align', 'right')]),\n\n            dict(selector=\"th:hover\",\n                 props=[(\"font-size\", \"14pt\")]),\n\n            dict(selector=\"tr:hover td:hover\",\n                 props=[('max-width', '250px'),\n                        ('font-size', '14pt'),\n                        ('color', f'{base_color}'),\n                        ('font-weight', 'bold'),\n                        ('background-color', 'white'),\n                        ('border', f'1px dashed {base_color}')]),\n\n             dict(selector=\"caption\",\n                  props=[(('caption-side', 'bottom'))])] + highlight_target_row\n\ndef stylize_min_max_count(pivot_table):\n    \"\"\"Waps the min_max_count pivot_table into the Styler.\n\n        Args:\n            df: |min_train| max_train |min_test |max_test |top10_counts_train |top_10_counts_train|\n\n        Returns:\n            s: the dataframe wrapped into Styler.\n    \"\"\"\n    s = pivot_table\n    # A formatting dictionary for controlling each column precision (.000 <-). \n    di_frmt = {(i if i.startswith('m') else i):\n              ('{:.3f}' if i.startswith('m') else '{:}') for i in s.columns}\n\n    s = s.style.set_table_styles(magnify(True))\\\n        .format(di_frmt)\\\n        .set_caption(f\"The train and test datasets min, max, top10 values side by side (hover to magnify).\")\n    return s\n  \n    \ndef stylize_describe(df: pd.DataFrame, dataset_name: str = 'train', is_test: bool = False) -> Styler:\n    \"\"\"Applies .descibe() method to the df and wraps it into the Styler.\n    \n        Args:\n            df: any dataframe (train/test/origin)\n            dataset_name: default 'train'\n            is_test: the bool parameter passed into magnify() function\n                     in order to control the highlighting of the last row.\n                     \n        Returns:\n            s: the dataframe wrapped into Styler.\n    \"\"\"\n    s = df.describe().T\n    # A formatting dictionary for controlling each column precision (.000 <-). \n    di_frmt = {(i if i == 'count' else i):\n              ('{:.0f}' if i == 'count' else '{:.3f}') for i in s.columns}\n    \n    s = s.style.set_table_styles(magnify(is_test))\\\n        .format(di_frmt)\\\n        .set_caption(f\"The {dataset_name} dataset descriptive statistics (hover to magnify).\")\n    return s\n\ndef stylize_simple(df: pd.DataFrame, caption: str) -> Styler:\n    \"\"\"Waps the min_max_count pivot_table into the Styler.\n\n        Args:\n            df: any dataframe (train/test/origin)\n\n        Returns:\n            s: the dataframe wrapped into Styler.\n    \"\"\"\n    s = df\n    s = s.style.set_table_styles(magnify(True)).set_caption(f\"{caption}\")\n    return s","metadata":{"_kg_hide-input":true,"trusted":true,"execution":{"iopub.status.busy":"2025-03-10T01:52:06.951235Z","iopub.execute_input":"2025-03-10T01:52:06.951689Z","iopub.status.idle":"2025-03-10T01:52:06.964394Z","shell.execute_reply.started":"2025-03-10T01:52:06.951651Z","shell.execute_reply":"2025-03-10T01:52:06.963475Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"WORKDIR = '/kaggle/input/byu-locating-bacterial-flagellar-motors-2025/'\nTRAIN_DIR = WORKDIR + 'train'\n\ntrain = pd.read_csv(WORKDIR + 'train_labels.csv')\nsub = pd.read_csv(WORKDIR + 'sample_submission.csv')\n\ndisplay(stylize_simple(train.head(), '* train_labels.csv '))\ndisplay(stylize_simple(sub, '* sample_submission.csv '))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-10T01:52:06.965395Z","iopub.execute_input":"2025-03-10T01:52:06.965592Z","iopub.status.idle":"2025-03-10T01:52:07.017429Z","shell.execute_reply.started":"2025-03-10T01:52:06.965575Z","shell.execute_reply":"2025-03-10T01:52:07.016412Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Note:**\n- We are predicting coordinates of a motor (z, y, x wrt 0, 1, 2)","metadata":{}},{"cell_type":"code","source":"stylize_describe(train)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-10T01:52:07.018176Z","iopub.execute_input":"2025-03-10T01:52:07.018396Z","iopub.status.idle":"2025-03-10T01:52:07.056824Z","shell.execute_reply.started":"2025-03-10T01:52:07.018373Z","shell.execute_reply":"2025-03-10T01:52:07.055232Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(f'{cS.red}Number of Unique tomo_ids in train.csv:{cS.res} {len(train.tomo_id.unique())}')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-10T01:52:07.058276Z","iopub.execute_input":"2025-03-10T01:52:07.058871Z","iopub.status.idle":"2025-03-10T01:52:07.068313Z","shell.execute_reply.started":"2025-03-10T01:52:07.058829Z","shell.execute_reply":"2025-03-10T01:52:07.067464Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Group by `tomo_id` and perform aggregation\nagg_train = train.groupby('tomo_id').agg(\n    mean_voxel_spacing=('Voxel spacing', 'mean'),\n    n_unique_array_shape_0=('Array shape (axis 0)', 'nunique'),\n    n_unique_array_shape_1=('Array shape (axis 1)', 'nunique'),\n    n_unique_array_shape_2=('Array shape (axis 2)', 'nunique'),\n    mean_number_of_motors=('Number of motors', 'mean')\n).reset_index()\n\n# Check for consistency: Array shapes should be consistent within each `tomo_id`\nagg_train['array_shapes_consistent'] = (\n    (agg_train['n_unique_array_shape_0'] == 1) &\n    (agg_train['n_unique_array_shape_1'] == 1) &\n    (agg_train['n_unique_array_shape_2'] == 1)\n)\n\n# Validate motor axis consistency: Motor axis should be within the array coordinates\nmotor_axis_validity = (\n    (train['Motor axis 0'] >= 0) &\n    (train['Motor axis 0'] <= train['Array shape (axis 0)']) &\n    (train['Motor axis 1'] >= 0) &\n    (train['Motor axis 1'] <= train['Array shape (axis 1)']) &\n    (train['Motor axis 2'] >= 0) &\n    (train['Motor axis 2'] <= train['Array shape (axis 2)'])\n)\n\ntrain['motor_axis_valid'] = motor_axis_validity\n\n# Group by `tomo_id` again to validate number of motors matches number of rows per `tomo_id`\nmotors_per_tomo = train.groupby('tomo_id').size().reset_index(name='num_rows')\nmotors_per_tomo = motors_per_tomo.merge(agg_train[['tomo_id', 'mean_number_of_motors']], on='tomo_id')\n\n# Validate number of motors for each `tomo_id`\nmotors_per_tomo['num_motors_match_rows'] = motors_per_tomo['num_rows'] == motors_per_tomo['mean_number_of_motors']\n\n# # Displaying the results\ndisplay(stylize_simple(agg_train.head(3), '* sanity check for the shapes consistency across the rows.'))\ndisplay(stylize_simple(motors_per_tomo.head(3), '* sanity check for the Number of motors consistency across the rows.'))\n\n# Print validation results\nprint(f\"{cS.red}* Aggregated Train Data{cS.res}\")\nprint(f\"Consistency of Array Shapes all: {agg_train['array_shapes_consistent'].all()}\")\nprint(f\"Motor Axis Validity all: {motor_axis_validity.all()}\")\nprint(f\"Number of Motors Match Rows all: {motors_per_tomo['num_motors_match_rows'].all()}\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-10T01:52:07.069195Z","iopub.execute_input":"2025-03-10T01:52:07.069492Z","iopub.status.idle":"2025-03-10T01:52:07.135697Z","shell.execute_reply.started":"2025-03-10T01:52:07.069464Z","shell.execute_reply":"2025-03-10T01:52:07.134837Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"stylize_simple(train[train['motor_axis_valid'].eq(False)].head(5), '* train with Number of motors equal to 0, Motor Axes are -1, -1, -1')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-10T01:52:07.138221Z","iopub.execute_input":"2025-03-10T01:52:07.138523Z","iopub.status.idle":"2025-03-10T01:52:07.151877Z","shell.execute_reply.started":"2025-03-10T01:52:07.138498Z","shell.execute_reply":"2025-03-10T01:52:07.150291Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"stylize_simple(motors_per_tomo[motors_per_tomo['num_motors_match_rows'].eq(False)].head(5), '* number of train rows not matched to number of motors, total 290')","metadata":{"scrolled":true,"trusted":true,"execution":{"iopub.status.busy":"2025-03-10T01:52:07.153608Z","iopub.execute_input":"2025-03-10T01:52:07.153833Z","iopub.status.idle":"2025-03-10T01:52:07.173836Z","shell.execute_reply.started":"2025-03-10T01:52:07.153816Z","shell.execute_reply":"2025-03-10T01:52:07.172801Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"stylize_simple(train[train.tomo_id.eq('tomo_003acc')], '* looking at the specific example where number of motors does not match number of rows')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-10T01:52:07.174753Z","iopub.execute_input":"2025-03-10T01:52:07.175047Z","iopub.status.idle":"2025-03-10T01:52:07.200241Z","shell.execute_reply.started":"2025-03-10T01:52:07.175024Z","shell.execute_reply":"2025-03-10T01:52:07.199401Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"set(train[train['motor_axis_valid'].eq(False)].tomo_id) ^ set(motors_per_tomo[motors_per_tomo['num_motors_match_rows'].eq(False)].tomo_id)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-10T01:52:07.200953Z","iopub.execute_input":"2025-03-10T01:52:07.201216Z","iopub.status.idle":"2025-03-10T01:52:07.222644Z","shell.execute_reply.started":"2025-03-10T01:52:07.201200Z","shell.execute_reply":"2025-03-10T01:52:07.221260Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"stylize_simple(train[train.tomo_id.isin(['tomo_2b3cdf', 'tomo_62eea8', 'tomo_c84b8e', 'tomo_e6f7f7'])], '* non zero examples with single row but two motors')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-10T01:52:07.223648Z","iopub.execute_input":"2025-03-10T01:52:07.223946Z","iopub.status.idle":"2025-03-10T01:52:07.248659Z","shell.execute_reply.started":"2025-03-10T01:52:07.223913Z","shell.execute_reply":"2025-03-10T01:52:07.247397Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Note:**\n- We are predicting coordinates of a motor (z, y, x wrt 0, 1, 2)\n- We got more records than actual train parent dirs, which obviously suggests that for some ids we got multiple coordinates (motors).\n- Some tomos might missing motors labels completely and marked as -1, -1, -1 w.r.t motor_axis_{0,1,2}\n- We got numbers of motors equals to 2 for `'tomo_2b3cdf', 'tomo_62eea8', 'tomo_c84b8e', 'tomo_e6f7f7'` though there is only one row for each of them in the `train.csv`","metadata":{}},{"cell_type":"markdown","source":"## <p style=\"font-family:'IBM Plex Sans'; font-weight:bold; letter-spacing: 1px; color:#1AAB70; font-size:120%; text-align:left;padding: 0px;\">Key Observations:</p>  \n\n- **Motor Coordinates**: The **Motor axis 0, 1, and 2** have a minimum value of **-1**, likely indicating missing or absent motors.  \n- **Array Shape**: The **z-axis (axis 0)** varies between **300 and 800**, while the **x and y dimensions** are quite large, with max values exceeding **1800 pixels**. This suggests that the tomogram slices have a high spatial resolution.  \n- **Voxel Spacing**: The voxel spacing varies significantly (from **6.5 to 19.7**), indicating different levels of resolution across tomograms.  \n- **Number of Motors**: Most tomograms contain **one or zero motors**, but some outliers have up to **10 motors** in a single tomogram.  \n---\n\n## <p style=\"font-family:'IBM Plex Sans'; font-weight:bold; letter-spacing: 1px; color:#1AAB70; font-size:120%; text-align:left;padding: 0px;\">Takeaways & Future Considerations:</p>  \n\n- **Image Size**: The tomograms have **large width and height**, which may impact storage and computation. We might consider patching, downsampling or resizing them if necessary.  \n- **Voxel Resampling**: Due to **variable voxel spacing**, resampling to a consistent resolution might be required for better model performance.  \n- **Handling Missing Data**: The presence of `-1` values in motor coordinates suggests we need a strategy to handle missing or absent motors effectively.","metadata":{}},{"cell_type":"markdown","source":"# <p style=\"font-family:JetBrains Mono; font-weight:normal; letter-spacing: 1px; color:#1AAB70; font-size:140%; text-align:left;padding: 0px; border-bottom: 3px solid #1AAB70\">Dataset Sanity Check & Target Counts</p>","metadata":{}},{"cell_type":"code","source":"%%time\n# Use glob to find all files inside the subdirectories\nfiles = glob(f'{TRAIN_DIR}/**/*')\n\nprint(f\"{cS.red}* len train files {len(files)}{cS.res}\")\n\n\ndef split_path(path):\n    from pathlib import Path\n    p = Path(path)\n    \n    # Extract parent directory, base name, and stem\n    tomo_id = p.parent.name\n    base_name = p.name\n    stem = p.stem\n    ext = p.suffix\n    \n    # Extract the slice number from the stem by removing 'slice_' and converting to integer\n    slice_num = int(stem.replace('slice_', '')) if 'slice_' in stem else None\n    return tomo_id, base_name, slice_num, ext\n\n# Inits train_paths_df\ntrain_paths_df = pd.DataFrame(files, columns=['abs_path_jpeg'])\n\n# Apply the split_path function to the 'abs_path_jpeg' column\ntrain_paths_df[\n    ['tomo_id', 'base_name', 'slice_num', 'ext']\n] = train_paths_df['abs_path_jpeg'].progress_apply(lambda x: pd.Series(split_path(x)))\n\nstylize_simple(train_paths_df.head(5), '* traversing TRAIN_DIR and collecting the file paths parsed in train_paths_df')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-10T01:52:07.249746Z","iopub.execute_input":"2025-03-10T01:52:07.250063Z","iopub.status.idle":"2025-03-10T01:52:47.397685Z","shell.execute_reply.started":"2025-03-10T01:52:07.250037Z","shell.execute_reply":"2025-03-10T01:52:47.396777Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Display the value counts for 'ext' column\next_value_counts = train_paths_df.ext.value_counts().to_dict()\nprint(f\"{cS.red}* Value Counts for ext: {cS.res}{ext_value_counts}\")  \n\n# Display the number of unique parent directories\nlen_uniq_tomo = len(train_paths_df.tomo_id.unique())\nprint(f\"{cS.red}* Number of Unique tomo_ids: {cS.res}{len_uniq_tomo}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-10T01:52:47.398417Z","iopub.execute_input":"2025-03-10T01:52:47.398647Z","iopub.status.idle":"2025-03-10T01:52:47.438315Z","shell.execute_reply.started":"2025-03-10T01:52:47.398625Z","shell.execute_reply":"2025-03-10T01:52:47.437170Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"num_slices_per_tomo = train_paths_df.groupby(by=['tomo_id'])['slice_num'].count().reset_index() \nnum_motors_per_tomo = train.groupby(by=['tomo_id'])['Number of motors'].first().reset_index()\nmerged_slices_motors = num_slices_per_tomo.merge(num_motors_per_tomo, how='left', on='tomo_id')\n\n# Create a count plot for 'Number of motors'\nplt.figure(figsize=(10, 6))\nax = sns.countplot(data=num_motors_per_tomo, x='Number of motors', palette=PALETTE)\n\n# Despine top and right\nsns.despine(top=True, right=True)\n\n# Add annotations on top of the bars\nfor p in ax.patches:\n    height = p.get_height()\n    total = len(merged_slices_motors)\n    percentage = (height / total) * 100\n    ax.annotate(f'{height}\\n({percentage:.1f}%)', \n                (p.get_x() + p.get_width() / 2., height + 5), \n                ha='center', va='center', \n                xytext=(0, 10), \n                textcoords='offset points')\n\nplt.title('Count of Number of Motors', color='black')\nplt.xlabel('Number of Motors')\nplt.ylabel('Count')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-10T01:52:47.439440Z","iopub.execute_input":"2025-03-10T01:52:47.439721Z","iopub.status.idle":"2025-03-10T01:52:47.766193Z","shell.execute_reply.started":"2025-03-10T01:52:47.439698Z","shell.execute_reply":"2025-03-10T01:52:47.764755Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"fig = plt.figure(figsize=(15, 8))  \nax = sns.histplot(data=merged_slices_motors, x='slice_num', hue='Number of motors', palette=PALETTE, multiple='dodge', shrink=0.8)\n\n# Add annotations on top of the bars\nfor p in ax.patches:\n    height = p.get_height()\n    if height > 0:\n        ax.annotate(\n            f'{int(height)}',                          # Text to display (height of the bar)\n            (p.get_x() + p.get_width() / 2., height),  # Position of the annotation\n            ha='center',    # Horizontal alignment\n            va='bottom',    # Vertical alignment\n            fontsize=8,     # Font size\n            color='black',  # Text color\n            fontweight='bold'\n        )\n\n# Set x-axis ticks with a step size of 50\nax.set_xticks(range(250, int(merged_slices_motors['slice_num'].max()) + 50, 50))\nax.set_ylim(0, 320)\nsns.despine(top=True, right=True)\nplt.title('Number of motors vs number of slices per tomo_id', color='black')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-10T01:52:47.767075Z","iopub.execute_input":"2025-03-10T01:52:47.767365Z","iopub.status.idle":"2025-03-10T01:52:48.228234Z","shell.execute_reply.started":"2025-03-10T01:52:47.767322Z","shell.execute_reply":"2025-03-10T01:52:48.227117Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## <p style=\"font-family:'IBM Plex Sans'; font-weight:bold; letter-spacing: 1px; color:#1AAB70; font-size:120%; text-align:left;padding: 0px;\">Key Observations:</p>  \n\n- **File Extensions**: The dataset only consists of **.jpg** files, with a count of **269,194**.\n- **Slice Number Distribution**: The `slice_num` distribution ranges from **0 to 799**, with the median slice being at **207**. \n- **Tomo_ids**: There are **648 unique Tomo_ids**, which indicate the number of distinct tomograms available in the dataset.\n- **Slice number**: greater than `750` got `0` Number of motors, is it a leak?! Also it is intersting to see distinct clusters of tomograms with a specific number of slices, for example `300` slices - `1st cluster`, `400` - `second`, `500`, `3rd` and `750` - `4th`.\n---\n\n## <p style=\"font-family:'IBM Plex Sans'; font-weight:bold; letter-spacing: 1px; color:#1AAB70; font-size:120%; text-align:left;padding: 0px;\">Takeaways & Future Considerations:</p>  \n\n- **Large Image Size**: The tomograms are large, both in width and height, which could be a challenge in terms of memory and computation. Consideration should be given to techniques such as image patching, downsampling, or resizing the images to manage these constraints.\n- **Voxel Resampling**: Due to the **variable voxel spacing**, it might be necessary to resample the tomograms to a consistent resolution for improved model generalization and performance.\n- **Handling Missing Data**: The presence of `-1` values in the motor coordinates suggests the need for a strategy to deal with absent motors.\n","metadata":{}},{"cell_type":"markdown","source":"# <p style=\"font-family:JetBrains Mono; font-weight:normal; letter-spacing: 1px; color:#1AAB70; font-size:140%; text-align:left;padding: 0px; border-bottom: 3px solid #1AAB70\">Dataset Stats</p>","metadata":{}},{"cell_type":"code","source":"%%time\nfrom multiprocessing import Pool, cpu_count\n\ndef process_group(tomo_id, group):\n    volume = []\n    for _, row in group.iterrows():\n        img = Image.open(row['abs_path_jpeg'])\n        img_array = np.array(img)\n        volume.append(img_array)\n    \n    volume = np.stack(volume)\n    z, y, x = volume.shape\n    \n    # Calculate volume size in megabytes\n    volume_size_mb = (volume.nbytes) / (1024 * 1024)\n    \n    return {\n        'tomo_id': tomo_id,\n        'min': np.min(volume),\n        'max': np.max(volume),\n        'std': np.std(volume),\n        'mean': np.mean(volume),\n        'z_shape': z,\n        'y_shape': y,\n        'x_shape': x,\n        'volume_size_mb': volume_size_mb  \n    }\n\n# Sort the dataframe by parent_dir and slice_num\ntrain_paths_df = train_paths_df.sort_values(by=['tomo_id', 'slice_num'])\n\n# Group by parent_dir\ngrouped = list(train_paths_df.groupby('tomo_id'))[:2]\n\n# Use multiprocessing to process groups in parallel\nif os.path.exists('/kaggle/input/2025-byu-locating-bacterial-motors-public-repo/volume_stats.csv'):\n    results_df = pd.read_csv('/kaggle/input/2025-byu-locating-bacterial-motors-public-repo/volume_stats.csv')\nelse:\n    with Pool(cpu_count()) as pool:\n        results = list(tqdm(pool.starmap(process_group, [(tomo_id, group) for tomo_id, group in grouped]), desc=\"Processing groups\"))\n\n    # Combine results into a dataframe and save\n    results_df = pd.DataFrame(results)\n    results_df.to_csv('volume_stats.csv', index=False)\n    \nstylize_simple(results_df.head(5), '* results_df combines tomograms as volumes and provides basic volume stats')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-10T01:52:48.229300Z","iopub.execute_input":"2025-03-10T01:52:48.229588Z","iopub.status.idle":"2025-03-10T01:52:48.376072Z","shell.execute_reply.started":"2025-03-10T01:52:48.229565Z","shell.execute_reply":"2025-03-10T01:52:48.375196Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Calculate aspect ratio (y_shape / x_shape)\nresults_df['aspect_ratio'] = results_df['y_shape'] / results_df['x_shape']\nmerged_df = results_df.merge(num_motors_per_tomo, on='tomo_id', how='left')\n\n# Create a 3x2 grid of subplots\nfig, axes = plt.subplots(3, 2, figsize=(15, 18))\n\n# Plot mean distribution\nsns.histplot(data=merged_df, x='mean', hue='Number of motors', palette=PALETTE, alpha=0.7, multiple='stack', kde=True, ax=axes[0, 0])\naxes[0, 0].set_title('Mean Distribution for Each Tomogram', color='black')\naxes[0, 0].set_xlabel('Mean Intensity')\naxes[0, 0].set_ylabel('Frequency')\n\n# Plot std distribution\nsns.histplot(data=merged_df, x='std', hue='Number of motors', palette=PALETTE, alpha=0.7, multiple='stack', kde=True, ax=axes[0, 1])\naxes[0, 1].set_title('Standard Deviation Distribution for Each Tomogram', color='black')\naxes[0, 1].set_xlabel('Standard Deviation')\naxes[0, 1].set_ylabel('Frequency')\n\n# Plot y_shape distribution\nsns.histplot(data=merged_df, x='y_shape', color=PALETTE[1], kde=True, ax=axes[1, 0])\naxes[1, 0].set_title('Y Shape Distribution for Each Tomogram', color='black')\naxes[1, 0].set_xlabel('Y Shape')\naxes[1, 0].set_ylabel('Frequency')\n\n# Plot x_shape distribution\nsns.histplot(data=merged_df, x='x_shape', color=PALETTE[1], kde=True, ax=axes[1, 1])\naxes[1, 1].set_title('X Shape Distribution for Each Tomogram', color='black')\naxes[1, 1].set_xlabel('X Shape')\naxes[1, 1].set_ylabel('Frequency')\n\n# Scatterplot of x_shape vs y_shape\nsns.scatterplot(data=merged_df, x='x_shape', y='y_shape', hue='Number of motors', palette=PALETTE, alpha=0.7, ax=axes[2, 0])\naxes[2, 0].set_title('X Shape vs Y Shape', color='black')\naxes[2, 0].set_xlabel('X Shape')\naxes[2, 0].set_ylabel('Y Shape')\n\n# Histplot of aspect ratio\nsns.histplot(data=merged_df, x='aspect_ratio', hue='Number of motors', palette=PALETTE, alpha=0.7, multiple='stack', kde=True, ax=axes[2, 1])\naxes[2, 1].set_title('Aspect Ratio Distribution', color='black')\naxes[2, 1].set_xlabel('Aspect Ratio (y_shape / x_shape)')\naxes[2, 1].set_ylabel('Frequency')\n\n# Adjust layout\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-10T01:52:48.377116Z","iopub.execute_input":"2025-03-10T01:52:48.377374Z","iopub.status.idle":"2025-03-10T01:52:51.851921Z","shell.execute_reply.started":"2025-03-10T01:52:48.377331Z","shell.execute_reply":"2025-03-10T01:52:51.850884Z"},"_kg_hide-input":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## <p style=\"font-family:'IBM Plex Sans'; font-weight:bold; letter-spacing: 1px; color:#1AAB70; font-size:120%; text-align:left;padding: 0px;\">Ref:</p>  \n\n- Credit to [@welshonionman](https://www.kaggle.com/welshonionman) for the original insights (found the same issue, but posted ealier).  \n- See the related [discussion](https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/566699) for more details.  \n","metadata":{}},{"cell_type":"code","source":"fig = plt.figure(figsize=(15, 6))\nsns.histplot(data=merged_df, x='volume_size_mb', hue='Number of motors', palette=PALETTE)\nsns.despine(top=True, right=True)\nplt.title('Histogram of tomogram sizes in MB', color='black')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-10T01:52:51.852677Z","iopub.execute_input":"2025-03-10T01:52:51.852873Z","iopub.status.idle":"2025-03-10T01:52:52.440658Z","shell.execute_reply.started":"2025-03-10T01:52:51.852856Z","shell.execute_reply":"2025-03-10T01:52:52.439674Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"stylize_simple(merged_df[merged_df['aspect_ratio'].lt(0.95)], 'possible apect ratio (AR) leak - AR < 0.95 = 0 Number of motors')","metadata":{"scrolled":true,"trusted":true,"execution":{"iopub.status.busy":"2025-03-10T01:52:52.441452Z","iopub.execute_input":"2025-03-10T01:52:52.441675Z","iopub.status.idle":"2025-03-10T01:52:52.453530Z","shell.execute_reply.started":"2025-03-10T01:52:52.441649Z","shell.execute_reply":"2025-03-10T01:52:52.452458Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## <p style=\"font-family:'IBM Plex Sans'; font-weight:bold; letter-spacing: 1px; color:#1AAB70; font-size:120%; text-align:left;padding: 0px;\">Key Observations:</p>  \n\n- **Image Aspect Ratio**: Some images with a particular aspect ratio contain **0** motors—curious if this could be a pseudo-leak.  \n- **Image Aspect Ratio**: For most records, the aspect ratio is close to **1**, meaning the `X` and `Y` dimensions are nearly equal, forming almost square `2D` slices.  \n- **Image Sizes**: Some `3D` volumes exceed **1.6GB** in storage space.  \n- **Mean Intensity & Aspect Ratio**: There appears to be a cluster of images that consistently contain **0** motors (`mean intensity` < `0.75` or `aspect_ratio` < `0.95` or `600` < `volume_size_mb` < `800` ).  \n\n---\n\n## <p style=\"font-family:'IBM Plex Sans'; font-weight:bold; letter-spacing: 1px; color:#1AAB70; font-size:120%; text-align:left;padding: 0px;\">Takeaways & Future Considerations:</p>  \n\n- **Post-processing & Probing**: We could explore whether applying `mean intensity` or `aspect ratio` thresholds allows us to identify tomograms with **0** motors and whether this improves results.  \n\n","metadata":{}},{"cell_type":"markdown","source":"# <p style=\"font-family:JetBrains Mono; font-weight:normal; letter-spacing: 1px; color:#1AAB70; font-size:140%; text-align:left;padding: 0px; border-bottom: 3px solid #1AAB70\">Visualization</p>  \n\n\nTODO","metadata":{}},{"cell_type":"markdown","source":"# <p style=\"font-family:JetBrains Mono; font-weight:normal; letter-spacing: 1px; color:#1AAB70; font-size:140%; text-align:left;padding: 0px; border-bottom: 3px solid #1AAB70\">VTK Visualization of the Volume</p>  \n\n\n## <p style=\"font-family:'IBM Plex Sans'; font-weight:bold; letter-spacing: 1px; color:#1AAB70; font-size:120%; text-align:left;padding: 0px;\">Ref:</p>  \n\n- Credit to [@johnpayne0](https://www.kaggle.com/johnpayne0) for the original insights.  \n- See the related [discussion](https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/566633) for more details.  \n- just `pip install vtk`, uncomment and run code below.\n\n---\nP.S. I tried to wrap VTK so it would work interactively in the Notebook, but alas! At some point, I considered it not worth the effort. Let me know if you figure that out.","metadata":{}},{"cell_type":"code","source":"!cp \"/kaggle/input/2025-byu-locating-bacterial-motors-public-repo/Screencastfrom03-09-2025025314PM-ezgif.com-video-to-gif-converter.gif\" .","metadata":{"trusted":true,"_kg_hide-input":true,"execution":{"iopub.status.busy":"2025-03-10T02:11:43.251966Z","iopub.execute_input":"2025-03-10T02:11:43.252250Z","iopub.status.idle":"2025-03-10T02:11:43.439528Z","shell.execute_reply.started":"2025-03-10T02:11:43.252232Z","shell.execute_reply":"2025-03-10T02:11:43.438434Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<p align=\"right\">\n  <center><img src=\"Screencastfrom03-09-2025025314PM-ezgif.com-video-to-gif-converter.gif\"/></center>\n</p>","metadata":{"execution":{"iopub.status.busy":"2025-03-10T02:11:44.162945Z","iopub.execute_input":"2025-03-10T02:11:44.163267Z","iopub.status.idle":"2025-03-10T02:11:44.169620Z","shell.execute_reply.started":"2025-03-10T02:11:44.163245Z","shell.execute_reply":"2025-03-10T02:11:44.168101Z"},"_kg_hide-input":false,"_kg_hide-output":false}},{"cell_type":"code","source":"# import imageio.v2 as imageio\n# import vtkmodules.all as vtk\n# from vtkmodules.util.numpy_support import numpy_to_vtk\n\n# folder_path = WORKDIR + \"/data/train/tomo_3e7407\"\n\n# slice_files = sorted([file for file in os.listdir(folder_path) if file.endswith('.jpg')])\n# slices = [imageio.imread(os.path.join(folder_path, file)) for file in slice_files]\n\n# volume = np.stack(slices, axis=0)\n# volume = (1.0 - (volume - np.min(volume)) / (np.max(volume) - np.min(volume))) * 255\n# volume = volume.astype(np.uint8)\n\n# vtk_data = numpy_to_vtk(volume.ravel(), deep=True, array_type=vtk.VTK_UNSIGNED_CHAR)\n# image_data = vtk.vtkImageData()\n# image_data.SetDimensions(volume.shape[::-1])\n# image_data.GetPointData().SetScalars(vtk_data)\n\n# mapper = vtk.vtkSmartVolumeMapper()\n# mapper.SetInputData(image_data)\n\n# volume_property = vtk.vtkVolumeProperty()\n# composite_function = vtk.vtkPiecewiseFunction()\n# composite_function.AddPoint(0, 0.0)\n# composite_function.AddPoint(255, 0.01)\n# volume_property.SetScalarOpacity(composite_function)\n\n# color = vtk.vtkColorTransferFunction()\n# color.AddRGBPoint(0, 0.0, 0.0, 0.0)\n# color.AddRGBPoint(255, 1.0, 1.0, 1.0)\n# volume_property.SetColor(color)\n\n# volume_actor = vtk.vtkVolume()\n# volume_actor.SetMapper(mapper)\n# volume_actor.SetProperty(volume_property)\n\n# renderer = vtk.vtkRenderer()\n# renderer.AddVolume(volume_actor)\n# renderer.SetBackground(0, 0, 0)\n\n# render_window = vtk.vtkRenderWindow()\n# render_window.AddRenderer(renderer)\n# render_window.SetSize(1280, 720)\n\n# interactor = vtk.vtkRenderWindowInteractor()\n# interactor.SetRenderWindow(render_window)\n\n# interactor_style = vtk.vtkInteractorStyleTrackballCamera()\n# interactor.SetInteractorStyle(interactor_style)\n\n# #render_window.Render()\n\n\n# interactor.Start()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-10T01:52:53.058384Z","iopub.execute_input":"2025-03-10T01:52:53.058642Z","iopub.status.idle":"2025-03-10T01:52:53.063501Z","shell.execute_reply.started":"2025-03-10T01:52:53.058623Z","shell.execute_reply":"2025-03-10T01:52:53.061515Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <p style=\"font-family:JetBrains Mono; font-weight:normal; letter-spacing: 1px; color:#1AAB70; font-size:140%; text-align:left;padding: 0px; border-bottom: 3px solid #1AAB70\">Outro and future work</p>  \n\nThe work is still in progress. I hope to continue working on it and add more to EDA, training and inference part. Good luck in the competition!","metadata":{}},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}