{
  "id": 583128,
  "title": "20th place solution -- Keypoint-Based Dual-Graph Predictor",
  "url": "/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/583128",
  "author_name": "Tom",
  "post_date": "2025-06-05T00:05:03.794000",
  "votes": 38,
  "comment_count": 8,
  "views": 0,
  "content": "<h1>Acknowledgement</h1>\n<p>Although I'm actually solo (with <a href=\"https://www.kaggle.com/daphne4sg\" target=\"_blank\">@daphne4sg</a> ) in this competition, I want to thank all my teammates <a href=\"https://www.kaggle.com/goodcoder\" target=\"_blank\">@goodcoder</a>, <a href=\"https://www.kaggle.com/bigochampion\" target=\"_blank\">@bigochampion</a> and <a href=\"https://www.kaggle.com/cybersimar08\" target=\"_blank\">@cybersimar08</a> who help me in another LLM competition and share good things about personal life. I learn a lot from you guys. I also want to thank <a href=\"https://www.kaggle.com/andrewjdarley\" target=\"_blank\">@andrewjdarley</a> for hosting such amazing competition. This project gives a lot of helps for my job interview. I have high willingness to do further development because your goal is very interesting.  </p>\n<h1>Short Story About This Solution</h1>\n<p>This solution originates from the best project of my career, which I completed during my master’s degree a year ago. I collaborated with a director at Micron Technology through an industrial partnership. The project was extremely challenging—we were required to build a high-performance anomaly detector using only 88 HBM scanned images due to data privacy constraints.</p>\n<p>Additionally, the labels were a mix of AOI-judged results and human expert annotations. Our objective was to model the preferences of the human expert. Some parts of my BYU solution were adapted from this HBM project, such as keypoint generation and label inference from graph structures.</p>\n<p>However, the most innovative aspect of my approach was leveraging few-shot prediction with a vision-language model (VLM) instead of traditional object detectors like YOLO. Essentially, I used prompts such as “this HBM is anomalous” or “HBM is OK” to represent the presence or absence of a target, allowing the model to make an initial classification guess.</p>\n<p>My solution calibrates the text-image embedding space using graph-based methods to enable the VLM to perform effectively for our specific task. As a result, I was able to train the model with only about 20 images, and it successfully detected nearly all anomalies identified by human experts in the HBM dataset.</p>\n<h1>Solution</h1>\n<h2>Initial Idea</h2>\n<p>The initial idea can refer to the <a href=\"https://www.kaggle.com/code/tom99763/first-idea-to-extract-keypoints-byu/notebook\" target=\"_blank\">notebook</a> where I made it public earlier. I want to produce a lot of points then refining them by building their relationship with labels. My strategy is sampling around the bacteria since motor existed in the head of it. If I aggregate all the keypoints  and summarize them, the correct semantic might can be captured and get away from noisy labels. However, the biggest issue is how to represent the feature of each point, so I only use patch feature extracted from a simple CNN then building GNN to predict the location of motor. It turns out a really bad model. </p>\n<h2>Further Improvement -- YOLO as Feature Extractor and Keypoint Generator</h2>\n<p>It's funny that I almost leaked my solution in <a href=\"https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/573491\" target=\"_blank\">this post</a> two month ago. I had short discussion with <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> about this development. After this discussion, I start to implement it and the idea becomes the pipeline showing below.  So basically I use the C3K2 feature map of YOLO11 because the feature in each spatial position can represent refined object. Setting small confidence threshold (conf=0.05) can filter enough object locations determined from YOLO11. This design boosts the score, improving basic approach a lot.</p>\n<h2>Node Feature -- C3K2 + RandomWalkPE</h2>\n<p>Feature makes very big difference for motor prediction. I've tried almost every types of features like HoG, Daisy, EfficientNet, ResNet and different YOLO versions, all of them do not work well except YOLO11, but the score is still not good enough. After several days, I found <code>AddRandomWalkPE</code> on the <a href=\"https://pytorch-geometric.readthedocs.io/en/2.5.1/generated/torch_geometric.transforms.AddRandomWalkPE.html\" target=\"_blank\">PYG document</a>. I suddenly realize this is the missing ingredient in 3d task. The score instantly boosts after I add it. So the node feature is C3K2 + <code>RanadomWalkPE</code> with <code>walk_length=8</code>. I think the main reason it improves the score is because it can represent the position of each point, implicitly modeling the position of motor.</p>\n<h2>Create More 3D points</h2>\n<p>YOLO successfully generates a lot of points by setting low confidence threshold, but it is still sparse in 3d space and has high possibility missing object existence. So my strategy for solving this issue is using the predicted point as center then uniformly sampling other points within a radius. I set <code>radius</code> and <code>num_samples</code> to control the range and density, respectively. I found using this strategy for training GNN also boosts score. </p>\n<h2>Label Development -- Similarity within Radius</h2>\n<p>Since annotating a single point as motor can not completely represent the actual semantic (that's why host gives so much radius tolerance in the metric), I label a point as positive by the following conditions:</p>\n<ul>\n<li>Features similarity between current location and ground truth location larger than <code>thr_sim</code>.</li>\n<li>The current location is within the 3d radius <code>thr</code> of ground truth location.</li>\n</ul>\n<pre><code> ():\n    label = train_labels[train_labels.tomo_id == tomo_id]\n    (, label[[, , ]].values[])\n    n_motors = label[].item()\n     n_motors == :\n         np.zeros(points.shape[], )\n    d, h, w = label[[, , ]].values[]\n    loc = label[[, , ]].values[]\n    tz, ty, tx = loc.astype()\n    fd, fh, fw = feat.shape[], feat.shape[], feat.shape[]\n    z = tz.astype()\n    y = ((ty/h) * fh).astype()\n    x = ((tx/w) * fw).astype()\n    \n    extract_feat = extract_feat / np.linalg.norm(extract_feat, axis=-, keepdims=)\n    target_feat = feat[z, :, y, x][, :]\n    target_feat = target_feat / np.linalg.norm(target_feat, axis=-, keepdims=)\n    sim = extract_feat @ target_feat.T\n    bool1 = sim[:, ]&gt;thr_sim\n\n    \n    dist = np.linalg.norm(points - loc[, :], axis=-) \n    bool2 = dist&lt;thr\n\n    \n    label = np.array(bool1 &amp; bool2, dtype=)\n     label\n</code></pre>\n<h2>Choice of Graphs</h2>\n<p>I experimented with various graph structures and ultimately found that combining a k-NN graph with a <a href=\"https://en.wikipedia.org/wiki/Delaunay_triangulation\" target=\"_blank\">Delaunay graph</a> yields the best results. The k-NN graph provides high density, while the Delaunay graph establishes connections between distant clusters. By leveraging both properties, the combined approach results in a more generalized model. The experiment results below clearly explain why someone only take 16 submission attemps and go to sleep for a while.</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>Public Score</th>\n<th>Private Score</th>\n<th>CV</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Knn Graph</td>\n<td><strong>0.835</strong></td>\n<td><strong>0.835</strong></td>\n<td>0.968</td>\n</tr>\n<tr>\n<td>Radius Graph</td>\n<td>0.826</td>\n<td>0.830</td>\n<td>0.966</td>\n</tr>\n<tr>\n<td>Delaunay Graph</td>\n<td>0.801</td>\n<td>0.806</td>\n<td>0.962</td>\n</tr>\n<tr>\n<td>Knn + Radius</td>\n<td>0.841</td>\n<td>0.835</td>\n<td>0.966</td>\n</tr>\n<tr>\n<td>Knn + Delaunay</td>\n<td><strong>0.856</strong></td>\n<td><strong>0.843</strong></td>\n<td>0.966</td>\n</tr>\n</tbody>\n</table>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4310004%2F6627841c58b70434ff5c698f41fffe0a%2Fgraphs.png?generation=1749125582409642&amp;alt=media\" alt=\"\"></p>\n<h2>Pipeline</h2>\n<p>First training a YOLO on the competition dataset. I don't use external dataset because I think the annotator is not the same person so it may gives differenet judgement for the motor existence. Then training the two GraphSages for each fold with <code>Binary Cross Entropy</code> with <code>positive_weight=8</code>. The ensemble function is maximum in my selected submission, but HDBSCAN is better in some scenarios. The training data for GNN is collected from the points filtered by YOLO which contains both positive and negative data. No extra dataset being used here.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4310004%2F82efe55faaa132866f20dbda41f235e7%2F123.png?generation=1749081901431650&amp;alt=media\" alt=\"\"></p>\n<h2>YOLO Size</h2>\n<p>Unfortunately I'm trapped from my experiments and cannot make the right choice in the final minutes of deadline. YOLO11L predicts a lot of false positives which is good in this competition. But I didin't select it. This means I still have to improve my mindset and experience as a data scientist.</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>Public Score</th>\n<th>Private Score</th>\n<th>CV</th>\n<th>Sub Time</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>YOLO11L</td>\n<td>0.856</td>\n<td><strong>0.843</strong></td>\n<td>0.968</td>\n<td>7HR</td>\n</tr>\n<tr>\n<td>YOLO11X</td>\n<td><strong>0.865</strong></td>\n<td>0.831</td>\n<td><strong>0.983</strong></td>\n<td>11HR</td>\n</tr>\n</tbody>\n</table>\n<h2>What didn't work</h2>\n<ul>\n<li>All the segmentation models </li>\n<li>Pure keypoints generation </li>\n<li>Larger model</li>\n</ul>\n<h1>Code Release</h1>\n<p><a href=\"https://www.kaggle.com/code/tom99763/21th-place-solution?scriptVersionId=240706129\" target=\"_blank\">yolo11l inference notebook</a> (0.843 pb, best private score) <br>\n<a href=\"https://www.kaggle.com/code/tom99763/21th-place-solution-yolo11x?scriptVersionId=243423147\" target=\"_blank\">yolo11x interence notebook</a> (0.831 pb, final submission) <br>\n<a href=\"https://github.com/tom99763/21th-place-solution-BYU\" target=\"_blank\">Github training code</a></p>\n<h1>Experiment Record</h1>\n<p>My pipeline is quite complicated, verifying those methods take me a lot of engineering effort. Here is all the experiment records in the past three months: <a href=\"https://docs.google.com/spreadsheets/d/1KCo8G_LF3T7FMXkVHWeKxo4vz24zoY7i697syACEaUY/edit?usp=sharing\" target=\"_blank\">Google Sheet</a></p>\n<p>If you have any question please comment below or use <a href=\"tom99763@gmail.com\" target=\"_blank\">email</a> &amp; <a href=\"www.linkedin.com/in/lin-chieh-huang-4b2231227\" target=\"_blank\">linkin</a> to contact me! <br>\nYou're welcome to share my solution on your blog—it's a great honor for me! I hope others will find it helpful and refer to it in future competitions.</p>",
  "messages": [
    {
      "id": 3217362,
      "postDate": "2025-06-05T00:05:03.793Z",
      "content": "<h1>Acknowledgement</h1>\n<p>Although I'm actually solo (with <a href=\"https://www.kaggle.com/daphne4sg\" target=\"_blank\">@daphne4sg</a> ) in this competition, I want to thank all my teammates <a href=\"https://www.kaggle.com/goodcoder\" target=\"_blank\">@goodcoder</a>, <a href=\"https://www.kaggle.com/bigochampion\" target=\"_blank\">@bigochampion</a> and <a href=\"https://www.kaggle.com/cybersimar08\" target=\"_blank\">@cybersimar08</a> who help me in another LLM competition and share good things about personal life. I learn a lot from you guys. I also want to thank <a href=\"https://www.kaggle.com/andrewjdarley\" target=\"_blank\">@andrewjdarley</a> for hosting such amazing competition. This project gives a lot of helps for my job interview. I have high willingness to do further development because your goal is very interesting.  </p>\n<h1>Short Story About This Solution</h1>\n<p>This solution originates from the best project of my career, which I completed during my master’s degree a year ago. I collaborated with a director at Micron Technology through an industrial partnership. The project was extremely challenging—we were required to build a high-performance anomaly detector using only 88 HBM scanned images due to data privacy constraints.</p>\n<p>Additionally, the labels were a mix of AOI-judged results and human expert annotations. Our objective was to model the preferences of the human expert. Some parts of my BYU solution were adapted from this HBM project, such as keypoint generation and label inference from graph structures.</p>\n<p>However, the most innovative aspect of my approach was leveraging few-shot prediction with a vision-language model (VLM) instead of traditional object detectors like YOLO. Essentially, I used prompts such as “this HBM is anomalous” or “HBM is OK” to represent the presence or absence of a target, allowing the model to make an initial classification guess.</p>\n<p>My solution calibrates the text-image embedding space using graph-based methods to enable the VLM to perform effectively for our specific task. As a result, I was able to train the model with only about 20 images, and it successfully detected nearly all anomalies identified by human experts in the HBM dataset.</p>\n<h1>Solution</h1>\n<h2>Initial Idea</h2>\n<p>The initial idea can refer to the <a href=\"https://www.kaggle.com/code/tom99763/first-idea-to-extract-keypoints-byu/notebook\" target=\"_blank\">notebook</a> where I made it public earlier. I want to produce a lot of points then refining them by building their relationship with labels. My strategy is sampling around the bacteria since motor existed in the head of it. If I aggregate all the keypoints  and summarize them, the correct semantic might can be captured and get away from noisy labels. However, the biggest issue is how to represent the feature of each point, so I only use patch feature extracted from a simple CNN then building GNN to predict the location of motor. It turns out a really bad model. </p>\n<h2>Further Improvement -- YOLO as Feature Extractor and Keypoint Generator</h2>\n<p>It's funny that I almost leaked my solution in <a href=\"https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/573491\" target=\"_blank\">this post</a> two month ago. I had short discussion with <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> about this development. After this discussion, I start to implement it and the idea becomes the pipeline showing below.  So basically I use the C3K2 feature map of YOLO11 because the feature in each spatial position can represent refined object. Setting small confidence threshold (conf=0.05) can filter enough object locations determined from YOLO11. This design boosts the score, improving basic approach a lot.</p>\n<h2>Node Feature -- C3K2 + RandomWalkPE</h2>\n<p>Feature makes very big difference for motor prediction. I've tried almost every types of features like HoG, Daisy, EfficientNet, ResNet and different YOLO versions, all of them do not work well except YOLO11, but the score is still not good enough. After several days, I found <code>AddRandomWalkPE</code> on the <a href=\"https://pytorch-geometric.readthedocs.io/en/2.5.1/generated/torch_geometric.transforms.AddRandomWalkPE.html\" target=\"_blank\">PYG document</a>. I suddenly realize this is the missing ingredient in 3d task. The score instantly boosts after I add it. So the node feature is C3K2 + <code>RanadomWalkPE</code> with <code>walk_length=8</code>. I think the main reason it improves the score is because it can represent the position of each point, implicitly modeling the position of motor.</p>\n<h2>Create More 3D points</h2>\n<p>YOLO successfully generates a lot of points by setting low confidence threshold, but it is still sparse in 3d space and has high possibility missing object existence. So my strategy for solving this issue is using the predicted point as center then uniformly sampling other points within a radius. I set <code>radius</code> and <code>num_samples</code> to control the range and density, respectively. I found using this strategy for training GNN also boosts score. </p>\n<h2>Label Development -- Similarity within Radius</h2>\n<p>Since annotating a single point as motor can not completely represent the actual semantic (that's why host gives so much radius tolerance in the metric), I label a point as positive by the following conditions:</p>\n<ul>\n<li>Features similarity between current location and ground truth location larger than <code>thr_sim</code>.</li>\n<li>The current location is within the 3d radius <code>thr</code> of ground truth location.</li>\n</ul>\n<pre><code> ():\n    label = train_labels[train_labels.tomo_id == tomo_id]\n    (, label[[, , ]].values[])\n    n_motors = label[].item()\n     n_motors == :\n         np.zeros(points.shape[], )\n    d, h, w = label[[, , ]].values[]\n    loc = label[[, , ]].values[]\n    tz, ty, tx = loc.astype()\n    fd, fh, fw = feat.shape[], feat.shape[], feat.shape[]\n    z = tz.astype()\n    y = ((ty/h) * fh).astype()\n    x = ((tx/w) * fw).astype()\n    \n    extract_feat = extract_feat / np.linalg.norm(extract_feat, axis=-, keepdims=)\n    target_feat = feat[z, :, y, x][, :]\n    target_feat = target_feat / np.linalg.norm(target_feat, axis=-, keepdims=)\n    sim = extract_feat @ target_feat.T\n    bool1 = sim[:, ]&gt;thr_sim\n\n    \n    dist = np.linalg.norm(points - loc[, :], axis=-) \n    bool2 = dist&lt;thr\n\n    \n    label = np.array(bool1 &amp; bool2, dtype=)\n     label\n</code></pre>\n<h2>Choice of Graphs</h2>\n<p>I experimented with various graph structures and ultimately found that combining a k-NN graph with a <a href=\"https://en.wikipedia.org/wiki/Delaunay_triangulation\" target=\"_blank\">Delaunay graph</a> yields the best results. The k-NN graph provides high density, while the Delaunay graph establishes connections between distant clusters. By leveraging both properties, the combined approach results in a more generalized model. The experiment results below clearly explain why someone only take 16 submission attemps and go to sleep for a while.</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>Public Score</th>\n<th>Private Score</th>\n<th>CV</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Knn Graph</td>\n<td><strong>0.835</strong></td>\n<td><strong>0.835</strong></td>\n<td>0.968</td>\n</tr>\n<tr>\n<td>Radius Graph</td>\n<td>0.826</td>\n<td>0.830</td>\n<td>0.966</td>\n</tr>\n<tr>\n<td>Delaunay Graph</td>\n<td>0.801</td>\n<td>0.806</td>\n<td>0.962</td>\n</tr>\n<tr>\n<td>Knn + Radius</td>\n<td>0.841</td>\n<td>0.835</td>\n<td>0.966</td>\n</tr>\n<tr>\n<td>Knn + Delaunay</td>\n<td><strong>0.856</strong></td>\n<td><strong>0.843</strong></td>\n<td>0.966</td>\n</tr>\n</tbody>\n</table>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4310004%2F6627841c58b70434ff5c698f41fffe0a%2Fgraphs.png?generation=1749125582409642&amp;alt=media\" alt=\"\"></p>\n<h2>Pipeline</h2>\n<p>First training a YOLO on the competition dataset. I don't use external dataset because I think the annotator is not the same person so it may gives differenet judgement for the motor existence. Then training the two GraphSages for each fold with <code>Binary Cross Entropy</code> with <code>positive_weight=8</code>. The ensemble function is maximum in my selected submission, but HDBSCAN is better in some scenarios. The training data for GNN is collected from the points filtered by YOLO which contains both positive and negative data. No extra dataset being used here.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4310004%2F82efe55faaa132866f20dbda41f235e7%2F123.png?generation=1749081901431650&amp;alt=media\" alt=\"\"></p>\n<h2>YOLO Size</h2>\n<p>Unfortunately I'm trapped from my experiments and cannot make the right choice in the final minutes of deadline. YOLO11L predicts a lot of false positives which is good in this competition. But I didin't select it. This means I still have to improve my mindset and experience as a data scientist.</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>Public Score</th>\n<th>Private Score</th>\n<th>CV</th>\n<th>Sub Time</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>YOLO11L</td>\n<td>0.856</td>\n<td><strong>0.843</strong></td>\n<td>0.968</td>\n<td>7HR</td>\n</tr>\n<tr>\n<td>YOLO11X</td>\n<td><strong>0.865</strong></td>\n<td>0.831</td>\n<td><strong>0.983</strong></td>\n<td>11HR</td>\n</tr>\n</tbody>\n</table>\n<h2>What didn't work</h2>\n<ul>\n<li>All the segmentation models </li>\n<li>Pure keypoints generation </li>\n<li>Larger model</li>\n</ul>\n<h1>Code Release</h1>\n<p><a href=\"https://www.kaggle.com/code/tom99763/21th-place-solution?scriptVersionId=240706129\" target=\"_blank\">yolo11l inference notebook</a> (0.843 pb, best private score) <br>\n<a href=\"https://www.kaggle.com/code/tom99763/21th-place-solution-yolo11x?scriptVersionId=243423147\" target=\"_blank\">yolo11x interence notebook</a> (0.831 pb, final submission) <br>\n<a href=\"https://github.com/tom99763/21th-place-solution-BYU\" target=\"_blank\">Github training code</a></p>\n<h1>Experiment Record</h1>\n<p>My pipeline is quite complicated, verifying those methods take me a lot of engineering effort. Here is all the experiment records in the past three months: <a href=\"https://docs.google.com/spreadsheets/d/1KCo8G_LF3T7FMXkVHWeKxo4vz24zoY7i697syACEaUY/edit?usp=sharing\" target=\"_blank\">Google Sheet</a></p>\n<p>If you have any question please comment below or use <a href=\"tom99763@gmail.com\" target=\"_blank\">email</a> &amp; <a href=\"www.linkedin.com/in/lin-chieh-huang-4b2231227\" target=\"_blank\">linkin</a> to contact me! <br>\nYou're welcome to share my solution on your blog—it's a great honor for me! I hope others will find it helpful and refer to it in future competitions.</p>",
      "rawMarkdown": "#Acknowledgement\nAlthough I'm actually solo (with @daphne4sg ) in this competition, I want to thank all my teammates @goodcoder, @bigochampion and @cybersimar08 who help me in another LLM competition and share good things about personal life. I learn a lot from you guys. I also want to thank @andrewjdarley for hosting such amazing competition. This project gives a lot of helps for my job interview. I have high willingness to do further development because your goal is very interesting.  \n\n#Short Story About This Solution\nThis solution originates from the best project of my career, which I completed during my master’s degree a year ago. I collaborated with a director at Micron Technology through an industrial partnership. The project was extremely challenging—we were required to build a high-performance anomaly detector using only 88 HBM scanned images due to data privacy constraints.\n\nAdditionally, the labels were a mix of AOI-judged results and human expert annotations. Our objective was to model the preferences of the human expert. Some parts of my BYU solution were adapted from this HBM project, such as keypoint generation and label inference from graph structures.\n\nHowever, the most innovative aspect of my approach was leveraging few-shot prediction with a vision-language model (VLM) instead of traditional object detectors like YOLO. Essentially, I used prompts such as “this HBM is anomalous” or “HBM is OK” to represent the presence or absence of a target, allowing the model to make an initial classification guess.\n\nMy solution calibrates the text-image embedding space using graph-based methods to enable the VLM to perform effectively for our specific task. As a result, I was able to train the model with only about 20 images, and it successfully detected nearly all anomalies identified by human experts in the HBM dataset.\n\n# Solution \n\n## Initial Idea\nThe initial idea can refer to the [notebook](https://www.kaggle.com/code/tom99763/first-idea-to-extract-keypoints-byu/notebook) where I made it public earlier. I want to produce a lot of points then refining them by building their relationship with labels. My strategy is sampling around the bacteria since motor existed in the head of it. If I aggregate all the keypoints  and summarize them, the correct semantic might can be captured and get away from noisy labels. However, the biggest issue is how to represent the feature of each point, so I only use patch feature extracted from a simple CNN then building GNN to predict the location of motor. It turns out a really bad model. \n\n##Further Improvement -- YOLO as Feature Extractor and Keypoint Generator \nIt's funny that I almost leaked my solution in [this post](https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/573491) two month ago. I had short discussion with @hengck23 about this development. After this discussion, I start to implement it and the idea becomes the pipeline showing below.  So basically I use the C3K2 feature map of YOLO11 because the feature in each spatial position can represent refined object. Setting small confidence threshold (conf=0.05) can filter enough object locations determined from YOLO11. This design boosts the score, improving basic approach a lot.\n\n## Node Feature -- C3K2 + RandomWalkPE\nFeature makes very big difference for motor prediction. I've tried almost every types of features like HoG, Daisy, EfficientNet, ResNet and different YOLO versions, all of them do not work well except YOLO11, but the score is still not good enough. After several days, I found `AddRandomWalkPE` on the [PYG document](https://pytorch-geometric.readthedocs.io/en/2.5.1/generated/torch_geometric.transforms.AddRandomWalkPE.html). I suddenly realize this is the missing ingredient in 3d task. The score instantly boosts after I add it. So the node feature is C3K2 + `RanadomWalkPE` with `walk_length=8`. I think the main reason it improves the score is because it can represent the position of each point, implicitly modeling the position of motor.\n\n\n## Create More 3D points \nYOLO successfully generates a lot of points by setting low confidence threshold, but it is still sparse in 3d space and has high possibility missing object existence. So my strategy for solving this issue is using the predicted point as center then uniformly sampling other points within a radius. I set `radius` and `num_samples` to control the range and density, respectively. I found using this strategy for training GNN also boosts score. \n\n## Label Development -- Similarity within Radius\nSince annotating a single point as motor can not completely represent the actual semantic (that's why host gives so much radius tolerance in the metric), I label a point as positive by the following conditions:\n* Features similarity between current location and ground truth location larger than `thr_sim`.\n* The current location is within the 3d radius `thr` of ground truth location.\n\n```python\ndef feat_labeling(points, feat, extract_feat, tomo_id, train_labels, thr = 10, thr_sim=0.5):\n    label = train_labels[train_labels.tomo_id == tomo_id]\n    print('location:', label[['Motor axis 0', 'Motor axis 1', 'Motor axis 2']].values[0])\n    n_motors = label['Number of motors'].item()\n    if n_motors == 0:\n        return np.zeros(points.shape[0], )\n    d, h, w = label[['Array shape (axis 0)', 'Array shape (axis 1)', 'Array shape (axis 2)']].values[0]\n    loc = label[['Motor axis 0', 'Motor axis 1', 'Motor axis 2']].values[0]\n    tz, ty, tx = loc.astype('int32')\n    fd, fh, fw = feat.shape[0], feat.shape[2], feat.shape[3]\n    z = tz.astype('int32')\n    y = ((ty/h) * fh).astype('int32')\n    x = ((tx/w) * fw).astype('int32')\n    #sim\n    extract_feat = extract_feat / np.linalg.norm(extract_feat, axis=-1, keepdims=True)\n    target_feat = feat[z, :, y, x][None, :]\n    target_feat = target_feat / np.linalg.norm(target_feat, axis=-1, keepdims=True)\n    sim = extract_feat @ target_feat.T\n    bool1 = sim[:, 0]>thr_sim\n\n    #dist\n    dist = np.linalg.norm(points - loc[None, :], axis=-1) #(N, 3)\n    bool2 = dist<thr\n\n    #label\n    label = np.array(bool1 & bool2, dtype='float32')\n    return label\n```\n\n## Choice of Graphs \nI experimented with various graph structures and ultimately found that combining a k-NN graph with a [Delaunay graph](https://en.wikipedia.org/wiki/Delaunay_triangulation) yields the best results. The k-NN graph provides high density, while the Delaunay graph establishes connections between distant clusters. By leveraging both properties, the combined approach results in a more generalized model. The experiment results below clearly explain why someone only take 16 submission attemps and go to sleep for a while.\n\n|  | Public Score | Private Score | CV |\n| --- | --- |\n| Knn Graph | **0.835** | **0.835** | 0.968 |\n| Radius Graph | 0.826 | 0.830 | 0.966 |\n| Delaunay Graph | 0.801 | 0.806 | 0.962 |\n| Knn + Radius | 0.841 | 0.835 | 0.966 |\n| Knn + Delaunay | **0.856** | **0.843** | 0.966 |\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4310004%2F6627841c58b70434ff5c698f41fffe0a%2Fgraphs.png?generation=1749125582409642&alt=media)\n\n## Pipeline\nFirst training a YOLO on the competition dataset. I don't use external dataset because I think the annotator is not the same person so it may gives differenet judgement for the motor existence. Then training the two GraphSages for each fold with `Binary Cross Entropy` with `positive_weight=8`. The ensemble function is maximum in my selected submission, but HDBSCAN is better in some scenarios. The training data for GNN is collected from the points filtered by YOLO which contains both positive and negative data. No extra dataset being used here.\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4310004%2F82efe55faaa132866f20dbda41f235e7%2F123.png?generation=1749081901431650&alt=media)\n\n\n##YOLO Size \nUnfortunately I'm trapped from my experiments and cannot make the right choice in the final minutes of deadline. YOLO11L predicts a lot of false positives which is good in this competition. But I didin't select it. This means I still have to improve my mindset and experience as a data scientist.\n\n|  | Public Score | Private Score | CV | Sub Time |\n| --- | --- |\n| YOLO11L | 0.856 |**0.843** | 0.968 | 7HR |\n| YOLO11X | **0.865** | 0.831 | **0.983** | 11HR |\n\n\n## What didn't work\n* All the segmentation models \n* Pure keypoints generation \n* Larger model\n\n\n# Code Release\n\n[yolo11l inference notebook](https://www.kaggle.com/code/tom99763/21th-place-solution?scriptVersionId=240706129) (0.843 pb, best private score) \n[yolo11x interence notebook](https://www.kaggle.com/code/tom99763/21th-place-solution-yolo11x?scriptVersionId=243423147) (0.831 pb, final submission) \n[Github training code](https://github.com/tom99763/21th-place-solution-BYU)\n\n# Experiment Record \nMy pipeline is quite complicated, verifying those methods take me a lot of engineering effort. Here is all the experiment records in the past three months: [Google Sheet](https://docs.google.com/spreadsheets/d/1KCo8G_LF3T7FMXkVHWeKxo4vz24zoY7i697syACEaUY/edit?usp=sharing)\n\n\n\nIf you have any question please comment below or use [email](tom99763@gmail.com) & [linkin](www.linkedin.com/in/lin-chieh-huang-4b2231227) to contact me! \nYou're welcome to share my solution on your blog—it's a great honor for me! I hope others will find it helpful and refer to it in future competitions.",
      "votes": 38
    },
    {
      "id": 3217697,
      "postDate": "2025-06-05T10:18:47.550Z",
      "content": "<p>congrats, well designed approach with lot to learn from.</p>",
      "rawMarkdown": "congrats, well designed approach with lot to learn from.",
      "votes": 1,
      "replies": [
        {
          "id": 3217699,
          "postDate": "2025-06-05T10:20:09.353Z",
          "content": "<p><a href=\"https://www.kaggle.com/mahdiseddigh\" target=\"_blank\">@mahdiseddigh</a> Let's pray our big bird can show up.</p>",
          "rawMarkdown": "@mahdiseddigh Let's pray our big bird can show up.",
          "votes": 1
        }
      ]
    },
    {
      "id": 3217368,
      "postDate": "2025-06-05T00:31:25.433Z",
      "content": "<p>Awesome looking solution. I didn't use a gnn setup but similarly used Yolo as a sort of region proposal system and then had a 3d classifier to reassess regions</p>",
      "rawMarkdown": "Awesome looking solution. I didn't use a gnn setup but similarly used Yolo as a sort of region proposal system and then had a 3d classifier to reassess regions",
      "votes": 1,
      "replies": [
        {
          "id": 3217379,
          "postDate": "2025-06-05T00:56:00.700Z",
          "content": "<p><a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a> I've tried region proposal method, but my experiments strongly show that sampling + graphs outperform it.</p>",
          "rawMarkdown": "@ryches I've tried region proposal method, but my experiments strongly show that sampling + graphs outperform it.",
          "votes": 1
        }
      ]
    },
    {
      "id": 3217846,
      "postDate": "2025-06-05T12:57:34.827Z",
      "content": "<p>My first gold medal :(<br>\nNOOOOOOOOOOOOOOOOO</p>",
      "rawMarkdown": "My first gold medal :(\nNOOOOOOOOOOOOOOOOO",
      "votes": 2,
      "replies": [
        {
          "id": 3218173,
          "postDate": "2025-06-05T23:51:03.423Z",
          "content": "<p>💀💀💀💀💀</p>",
          "rawMarkdown": "💀💀💀💀💀"
        }
      ]
    },
    {
      "id": 3217519,
      "postDate": "2025-06-05T05:14:50.643Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 3217672,
      "postDate": "2025-06-05T09:37:24.797Z",
      "content": "<p>Thank you for your sharing <br>\nUpvoted</p>",
      "rawMarkdown": "Thank you for your sharing \nUpvoted",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 3217697,
      "author_name": "mhdaw",
      "author_url": "",
      "post_date": "2025-06-05T10:18:47.550000",
      "content": "<p>congrats, well designed approach with lot to learn from.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3217699,
          "author_name": "Tom",
          "author_url": "",
          "post_date": "2025-06-05T10:20:09.353000",
          "content": "<p><a href=\"https://www.kaggle.com/mahdiseddigh\" target=\"_blank\">@mahdiseddigh</a> Let's pray our big bird can show up.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3217368,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "2025-06-05T00:31:25.433000",
      "content": "<p>Awesome looking solution. I didn't use a gnn setup but similarly used Yolo as a sort of region proposal system and then had a 3d classifier to reassess regions</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3217379,
          "author_name": "Tom",
          "author_url": "",
          "post_date": "2025-06-05T00:56:00.700000",
          "content": "<p><a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a> I've tried region proposal method, but my experiments strongly show that sampling + graphs outperform it.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3217846,
      "author_name": "daphne",
      "author_url": "",
      "post_date": "2025-06-05T12:57:34.827000",
      "content": "<p>My first gold medal :(<br>\nNOOOOOOOOOOOOOOOOO</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3218173,
          "author_name": "Tom",
          "author_url": "",
          "post_date": "2025-06-05T23:51:03.423000",
          "content": "<p>💀💀💀💀💀</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3217519,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-06-05T05:14:50.643000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3217672,
      "author_name": "Yasir Hussein Shakir",
      "author_url": "",
      "post_date": "2025-06-05T09:37:24.797000",
      "content": "<p>Thank you for your sharing <br>\nUpvoted</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3217362": "#Acknowledgement\nAlthough I'm actually solo (with @daphne4sg ) in this competition, I want to thank all my teammates @goodcoder, @bigochampion and @cybersimar08 who help me in another LLM competition and share good things about personal life. I learn a lot from you guys. I also want to thank @andrewjdarley for hosting such amazing competition. This project gives a lot of helps for my job interview. I have high willingness to do further development because your goal is very interesting.  \n\n#Short Story About This Solution\nThis solution originates from the best project of my career, which I completed during my master’s degree a year ago. I collaborated with a director at Micron Technology through an industrial partnership. The project was extremely challenging—we were required to build a high-performance anomaly detector using only 88 HBM scanned images due to data privacy constraints.\n\nAdditionally, the labels were a mix of AOI-judged results and human expert annotations. Our objective was to model the preferences of the human expert. Some parts of my BYU solution were adapted from this HBM project, such as keypoint generation and label inference from graph structures.\n\nHowever, the most innovative aspect of my approach was leveraging few-shot prediction with a vision-language model (VLM) instead of traditional object detectors like YOLO. Essentially, I used prompts such as “this HBM is anomalous” or “HBM is OK” to represent the presence or absence of a target, allowing the model to make an initial classification guess.\n\nMy solution calibrates the text-image embedding space using graph-based methods to enable the VLM to perform effectively for our specific task. As a result, I was able to train the model with only about 20 images, and it successfully detected nearly all anomalies identified by human experts in the HBM dataset.\n\n# Solution \n\n## Initial Idea\nThe initial idea can refer to the [notebook](https://www.kaggle.com/code/tom99763/first-idea-to-extract-keypoints-byu/notebook) where I made it public earlier. I want to produce a lot of points then refining them by building their relationship with labels. My strategy is sampling around the bacteria since motor existed in the head of it. If I aggregate all the keypoints  and summarize them, the correct semantic might can be captured and get away from noisy labels. However, the biggest issue is how to represent the feature of each point, so I only use patch feature extracted from a simple CNN then building GNN to predict the location of motor. It turns out a really bad model. \n\n##Further Improvement -- YOLO as Feature Extractor and Keypoint Generator \nIt's funny that I almost leaked my solution in [this post](https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/573491) two month ago. I had short discussion with @hengck23 about this development. After this discussion, I start to implement it and the idea becomes the pipeline showing below.  So basically I use the C3K2 feature map of YOLO11 because the feature in each spatial position can represent refined object. Setting small confidence threshold (conf=0.05) can filter enough object locations determined from YOLO11. This design boosts the score, improving basic approach a lot.\n\n## Node Feature -- C3K2 + RandomWalkPE\nFeature makes very big difference for motor prediction. I've tried almost every types of features like HoG, Daisy, EfficientNet, ResNet and different YOLO versions, all of them do not work well except YOLO11, but the score is still not good enough. After several days, I found `AddRandomWalkPE` on the [PYG document](https://pytorch-geometric.readthedocs.io/en/2.5.1/generated/torch_geometric.transforms.AddRandomWalkPE.html). I suddenly realize this is the missing ingredient in 3d task. The score instantly boosts after I add it. So the node feature is C3K2 + `RanadomWalkPE` with `walk_length=8`. I think the main reason it improves the score is because it can represent the position of each point, implicitly modeling the position of motor.\n\n\n## Create More 3D points \nYOLO successfully generates a lot of points by setting low confidence threshold, but it is still sparse in 3d space and has high possibility missing object existence. So my strategy for solving this issue is using the predicted point as center then uniformly sampling other points within a radius. I set `radius` and `num_samples` to control the range and density, respectively. I found using this strategy for training GNN also boosts score. \n\n## Label Development -- Similarity within Radius\nSince annotating a single point as motor can not completely represent the actual semantic (that's why host gives so much radius tolerance in the metric), I label a point as positive by the following conditions:\n* Features similarity between current location and ground truth location larger than `thr_sim`.\n* The current location is within the 3d radius `thr` of ground truth location.\n\n```python\ndef feat_labeling(points, feat, extract_feat, tomo_id, train_labels, thr = 10, thr_sim=0.5):\n    label = train_labels[train_labels.tomo_id == tomo_id]\n    print('location:', label[['Motor axis 0', 'Motor axis 1', 'Motor axis 2']].values[0])\n    n_motors = label['Number of motors'].item()\n    if n_motors == 0:\n        return np.zeros(points.shape[0], )\n    d, h, w = label[['Array shape (axis 0)', 'Array shape (axis 1)', 'Array shape (axis 2)']].values[0]\n    loc = label[['Motor axis 0', 'Motor axis 1', 'Motor axis 2']].values[0]\n    tz, ty, tx = loc.astype('int32')\n    fd, fh, fw = feat.shape[0], feat.shape[2], feat.shape[3]\n    z = tz.astype('int32')\n    y = ((ty/h) * fh).astype('int32')\n    x = ((tx/w) * fw).astype('int32')\n    #sim\n    extract_feat = extract_feat / np.linalg.norm(extract_feat, axis=-1, keepdims=True)\n    target_feat = feat[z, :, y, x][None, :]\n    target_feat = target_feat / np.linalg.norm(target_feat, axis=-1, keepdims=True)\n    sim = extract_feat @ target_feat.T\n    bool1 = sim[:, 0]>thr_sim\n\n    #dist\n    dist = np.linalg.norm(points - loc[None, :], axis=-1) #(N, 3)\n    bool2 = dist<thr\n\n    #label\n    label = np.array(bool1 & bool2, dtype='float32')\n    return label\n```\n\n## Choice of Graphs \nI experimented with various graph structures and ultimately found that combining a k-NN graph with a [Delaunay graph](https://en.wikipedia.org/wiki/Delaunay_triangulation) yields the best results. The k-NN graph provides high density, while the Delaunay graph establishes connections between distant clusters. By leveraging both properties, the combined approach results in a more generalized model. The experiment results below clearly explain why someone only take 16 submission attemps and go to sleep for a while.\n\n|  | Public Score | Private Score | CV |\n| --- | --- |\n| Knn Graph | **0.835** | **0.835** | 0.968 |\n| Radius Graph | 0.826 | 0.830 | 0.966 |\n| Delaunay Graph | 0.801 | 0.806 | 0.962 |\n| Knn + Radius | 0.841 | 0.835 | 0.966 |\n| Knn + Delaunay | **0.856** | **0.843** | 0.966 |\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4310004%2F6627841c58b70434ff5c698f41fffe0a%2Fgraphs.png?generation=1749125582409642&alt=media)\n\n## Pipeline\nFirst training a YOLO on the competition dataset. I don't use external dataset because I think the annotator is not the same person so it may gives differenet judgement for the motor existence. Then training the two GraphSages for each fold with `Binary Cross Entropy` with `positive_weight=8`. The ensemble function is maximum in my selected submission, but HDBSCAN is better in some scenarios. The training data for GNN is collected from the points filtered by YOLO which contains both positive and negative data. No extra dataset being used here.\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4310004%2F82efe55faaa132866f20dbda41f235e7%2F123.png?generation=1749081901431650&alt=media)\n\n\n##YOLO Size \nUnfortunately I'm trapped from my experiments and cannot make the right choice in the final minutes of deadline. YOLO11L predicts a lot of false positives which is good in this competition. But I didin't select it. This means I still have to improve my mindset and experience as a data scientist.\n\n|  | Public Score | Private Score | CV | Sub Time |\n| --- | --- |\n| YOLO11L | 0.856 |**0.843** | 0.968 | 7HR |\n| YOLO11X | **0.865** | 0.831 | **0.983** | 11HR |\n\n\n## What didn't work\n* All the segmentation models \n* Pure keypoints generation \n* Larger model\n\n\n# Code Release\n\n[yolo11l inference notebook](https://www.kaggle.com/code/tom99763/21th-place-solution?scriptVersionId=240706129) (0.843 pb, best private score) \n[yolo11x interence notebook](https://www.kaggle.com/code/tom99763/21th-place-solution-yolo11x?scriptVersionId=243423147) (0.831 pb, final submission) \n[Github training code](https://github.com/tom99763/21th-place-solution-BYU)\n\n# Experiment Record \nMy pipeline is quite complicated, verifying those methods take me a lot of engineering effort. Here is all the experiment records in the past three months: [Google Sheet](https://docs.google.com/spreadsheets/d/1KCo8G_LF3T7FMXkVHWeKxo4vz24zoY7i697syACEaUY/edit?usp=sharing)\n\n\n\nIf you have any question please comment below or use [email](tom99763@gmail.com) & [linkin](www.linkedin.com/in/lin-chieh-huang-4b2231227) to contact me! \nYou're welcome to share my solution on your blog—it's a great honor for me! I hope others will find it helpful and refer to it in future competitions.",
    "3217697": "congrats, well designed approach with lot to learn from.",
    "3217368": "Awesome looking solution. I didn't use a gnn setup but similarly used Yolo as a sort of region proposal system and then had a 3d classifier to reassess regions",
    "3217846": "My first gold medal :(\nNOOOOOOOOOOOOOOOOO",
    "3217519": "",
    "3217672": "Thank you for your sharing \nUpvoted"
  }
}