{
  "id": 573491,
  "title": "[closed] lb0.744 my experiment results",
  "url": "/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/573491",
  "author_name": "hengck23",
  "post_date": "2025-04-15T23:25:46.410000",
  "votes": 46,
  "comment_count": 49,
  "views": 0,
  "content": "<p><strong>The test data has been relabelled and lb has been rescored.</strong><br>\n<strong>This thread is based on old test data and will be closed.</strong><br>\n<strong>Please wait for the new thread …</strong></p>\n<hr>\n<p>reference notebook:<br>\n<a href=\"https://www.kaggle.com/code/hengck23/yolo-experiments\" target=\"_blank\">https://www.kaggle.com/code/hengck23/yolo-experiments</a></p>\n<p>[1] apr-16:</p>\n<ul>\n<li>yolo11L, imgsize=960, train= pos images only : lb0.744 </li>\n</ul>\n<hr>\n<p>todo:</p>\n<ul>\n<li>transformer-based DETR object detector</li>\n<li>multi-slice yolo, etc (maybe add conv-lstm, 3d conv or transformer fuser?)</li>\n<li>bootstrap neg train images</li>\n<li>bootstrap pos train images </li>\n<li>external data</li>\n<li>synthetic data via video diffusion?</li>\n</ul>\n<hr>\n<p>to be updated as experiment proceeds …this thread will be cias</p>",
  "messages": [
    {
      "id": 3179924,
      "postDate": "2025-04-15T23:25:46.410Z",
      "content": "<p><strong>The test data has been relabelled and lb has been rescored.</strong><br>\n<strong>This thread is based on old test data and will be closed.</strong><br>\n<strong>Please wait for the new thread …</strong></p>\n<hr>\n<p>reference notebook:<br>\n<a href=\"https://www.kaggle.com/code/hengck23/yolo-experiments\" target=\"_blank\">https://www.kaggle.com/code/hengck23/yolo-experiments</a></p>\n<p>[1] apr-16:</p>\n<ul>\n<li>yolo11L, imgsize=960, train= pos images only : lb0.744 </li>\n</ul>\n<hr>\n<p>todo:</p>\n<ul>\n<li>transformer-based DETR object detector</li>\n<li>multi-slice yolo, etc (maybe add conv-lstm, 3d conv or transformer fuser?)</li>\n<li>bootstrap neg train images</li>\n<li>bootstrap pos train images </li>\n<li>external data</li>\n<li>synthetic data via video diffusion?</li>\n</ul>\n<hr>\n<p>to be updated as experiment proceeds …this thread will be cias</p>",
      "rawMarkdown": "**The test data has been relabelled and lb has been rescored.**\n**This thread is based on old test data and will be closed.**\n**Please wait for the new thread ...**\n\n\n---\n\nreference notebook:\nhttps://www.kaggle.com/code/hengck23/yolo-experiments\n\n[1] apr-16:\n- yolo11L, imgsize=960, train= pos images only : lb0.744 \n\n---\ntodo:\n- transformer-based DETR object detector\n- multi-slice yolo, etc (maybe add conv-lstm, 3d conv or transformer fuser?)\n- bootstrap neg train images\n- bootstrap pos train images \n- external data\n- synthetic data via video diffusion?\n\n\n\n---\n\nto be updated as experiment proceeds ...this thread will be cias",
      "votes": 46
    },
    {
      "id": 3179926,
      "postDate": "2025-04-15T23:28:17.487Z",
      "content": "<p>details for yolo11L, imgsize=960, train= pos images only : lb0.744<br>\ndataset</p>\n<pre><code>yaml_config = {\n    : {: },\n     : ,\n    : ,\n    : ,\n\n    \n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n}\n</code></pre>\n<p>trainer</p>\n<pre><code>    result = model.train(\n        =yaml_file,\n        =50,\n        =cfg.batch_size, #16\n        =cfg.imgsz, #960,\n        =,\n        =1e-4,\n        =0.1,\n        =0,\n        =0.1,\n        =out_dir,\n\n        =,\n        =cfg.experiment,\n\n        =100,   \n        =2,  \n        =,   \n        =  \n    )\n</code></pre>",
      "rawMarkdown": "details for yolo11L, imgsize=960, train= pos images only : lb0.744\ndataset\n```\nyaml_config = {\n    'names': {0: 'motor'},\n    'path' : f'{YOLO_DIR}',\n    'train': f'{YOLO_DIR}/train_list.fold{cfg.fold}.txt',\n    'val': f'{YOLO_DIR}/valid_list.fold{cfg.fold}.txt',\n\n    #augmentation ---\n    'mosaic': 1.0,\n    'close_mosaic': 10,\n    'mixup': 0.4,\n    'flipud': 0.5,\n    'scale': 0.25,\n    'degrees': 45,\n}#rest are default\n```\n\ntrainer\n\n```\n    result = model.train(\n        data=yaml_file,\n        epochs=50,\n        batch=cfg.batch_size, #16\n        imgsz=cfg.imgsz, #960,\n        optimizer='AdamW',\n        lr0=1e-4,\n        lrf=0.1,\n        warmup_epochs=0,\n        dropout=0.1,\n        project=out_dir,\n      \n        exist_ok=True,\n        name=cfg.experiment,\n     \n        patience=100,   \n        save_period=2,  \n        val=True,   \n        verbose=True  \n    )\n\n```",
      "votes": 9,
      "replies": [
        {
          "id": 3180159,
          "postDate": "2025-04-16T08:21:15.100Z",
          "content": "<p>As always informative and interesting. RTDETR is implemented in Ultralytics, in case anyone did not pay attention to it</p>",
          "rawMarkdown": "As always informative and interesting. RTDETR is implemented in Ultralytics, in case anyone did not pay attention to it",
          "votes": 2
        },
        {
          "id": 3180185,
          "postDate": "2025-04-16T09:09:59.493Z",
          "content": "<p>Thanks a lot for sharing this!</p>\n<p>Surprised to see such a low learning rate with a relatively small number of epochs. Are you willing to share some loss curves for your training?</p>",
          "rawMarkdown": "Thanks a lot for sharing this!\n\nSurprised to see such a low learning rate with a relatively small number of epochs. Are you willing to share some loss curves for your training?",
          "replies": [
            {
              "id": 3180213,
              "postDate": "2025-04-16T09:49:30.260Z",
              "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fffad3b8eef78760e3120e77f6b8462c8%2Fdfl_loss_curve.png?generation=1744796749619799&amp;alt=media\" alt=\"\"></p>\n<p>loss curve for  yolo11L</p>\n<p>local cv</p>\n<pre><code>        lb, more = compute_lb(solution, submission, =1000, =2)\n        (th, lb, more[], more[], more[])\n\n** pos tomo_id only **\n{: 1, : 16, : 0.1, : 24, : 0.5, : , : }\n0.1 0.9872611464968152 0.9919999999999999 1.0 0.9841269841269841\n0.2 0.9872611464968152 0.9919999999999999 1.0 0.9841269841269841\n0.3 0.9872611464968152 0.9919999999999999 1.0 0.9841269841269841\n0.4 0.9872611464968152 0.9919999999999999 1.0 0.9841269841269841\n0.5 0.9872611464968152 0.9919999999999999 1.0 0.9841269841269841\n0.6 0.9872611464968152 0.9919999999999999 1.0 0.9841269841269841\n0.7 0.97444089456869 0.9838709677419354 1.0 0.9682539682539683\n0.75 0.8823529411764707 0.923076923076923 1.0 0.8571428571428571\n0.8 0.29850746268656714 0.4050632911392405 1.0 0.25396825396825395\n0.85 0.0 0.0 0.0 0.0\n0.9 0.0 0.0 0.0 0.0\n\n** pos tomo_id + neg tomo_id = 120 tomo_id **\n{: 1, : 16, : 0.1, : 24, : 0.5, : , : }\n0.1 0.8516483516483516 0.7085714285714286 0.5535714285714286 0.9841269841269841\n0.2 0.8516483516483516 0.7085714285714286 0.5535714285714286 0.9841269841269841\n0.3 0.8516483516483516 0.7085714285714286 0.5535714285714286 0.9841269841269841\n0.4 0.861111111111111 0.7251461988304093 0.5740740740740741 0.9841269841269841\n0.5 0.8757062146892655 0.7515151515151515 0.6078431372549019 0.9841269841269841\n0.6 0.9198813056379821 0.8378378378378377 0.7294117647058823 0.9841269841269841\n0.7 0.9442724458204335 0.9104477611940299 0.8591549295774648 0.9682539682539683\n0.75 0.8653846153846153 0.8780487804878048 0.9 0.8571428571428571\n0.8 0.2952029520295203 0.3902439024390244 0.8421052631578947 0.25396825396825395\n0.85 0.0 0.0 0.0 0.0\n0.9 0.0 0.0 0.0 0.0\n</code></pre>",
              "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fffad3b8eef78760e3120e77f6b8462c8%2Fdfl_loss_curve.png?generation=1744796749619799&alt=media)\n\nloss curve for  yolo11L\n\nlocal cv\n\n```\n\n        lb, more = compute_lb(solution, submission, min_radius=1000, beta=2)\n        print(th, lb, more['fscore'], more['precision'], more['recall'])\n\n** pos tomo_id only **\n{'slice_step': 1, 'batch_size': 16, 'box_min_conf': 0.1, 'box_size': 24, 'iou_threshold': 0.5, 'device': 'cuda', 'checkpoint': '/media/hp/c30d34ed-0d55-4077-82dc-b56cd13dd548/2025/kaggle/bacterial-flagellar-motors/result/my-yolo/03/yolo11l-960-00/xxx/weights/best.pt'}\n0.1 0.9872611464968152 0.9919999999999999 1.0 0.9841269841269841\n0.2 0.9872611464968152 0.9919999999999999 1.0 0.9841269841269841\n0.3 0.9872611464968152 0.9919999999999999 1.0 0.9841269841269841\n0.4 0.9872611464968152 0.9919999999999999 1.0 0.9841269841269841\n0.5 0.9872611464968152 0.9919999999999999 1.0 0.9841269841269841\n0.6 0.9872611464968152 0.9919999999999999 1.0 0.9841269841269841\n0.7 0.97444089456869 0.9838709677419354 1.0 0.9682539682539683\n0.75 0.8823529411764707 0.923076923076923 1.0 0.8571428571428571\n0.8 0.29850746268656714 0.4050632911392405 1.0 0.25396825396825395\n0.85 0.0 0.0 0.0 0.0\n0.9 0.0 0.0 0.0 0.0\n\n** pos tomo_id + neg tomo_id = 120 tomo_id **\n{'slice_step': 1, 'batch_size': 16, 'box_min_conf': 0.1, 'box_size': 24, 'iou_threshold': 0.5, 'device': 'cuda', 'checkpoint': '/media/hp/c30d34ed-0d55-4077-82dc-b56cd13dd548/2025/kaggle/bacterial-flagellar-motors/result/my-yolo/03/yolo11l-960-00/xxx/weights/best.pt'}\n0.1 0.8516483516483516 0.7085714285714286 0.5535714285714286 0.9841269841269841\n0.2 0.8516483516483516 0.7085714285714286 0.5535714285714286 0.9841269841269841\n0.3 0.8516483516483516 0.7085714285714286 0.5535714285714286 0.9841269841269841\n0.4 0.861111111111111 0.7251461988304093 0.5740740740740741 0.9841269841269841\n0.5 0.8757062146892655 0.7515151515151515 0.6078431372549019 0.9841269841269841\n0.6 0.9198813056379821 0.8378378378378377 0.7294117647058823 0.9841269841269841\n0.7 0.9442724458204335 0.9104477611940299 0.8591549295774648 0.9682539682539683\n0.75 0.8653846153846153 0.8780487804878048 0.9 0.8571428571428571\n0.8 0.2952029520295203 0.3902439024390244 0.8421052631578947 0.25396825396825395\n0.85 0.0 0.0 0.0 0.0\n0.9 0.0 0.0 0.0 0.0\n```",
              "votes": 1
            },
            {
              "id": 3181166,
              "postDate": "2025-04-17T14:29:34.500Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        },
        {
          "id": 3180878,
          "postDate": "2025-04-17T08:29:32.560Z",
          "content": "<p>Hi, I'm wondering if “mosaic” and“ mixup” data enhancements have a positive effect on your model. In one of my few experiments, these two enhancements seem to degrade model performance🌹</p>",
          "rawMarkdown": "Hi, I'm wondering if “mosaic” and“ mixup” data enhancements have a positive effect on your model. In one of my few experiments, these two enhancements seem to degrade model performance🌹"
        },
        {
          "id": 3207273,
          "postDate": "2025-05-22T14:04:42.413Z",
          "content": "<p>With this image size what was the inference time? I think it will exceed than allocated time</p>",
          "rawMarkdown": "With this image size what was the inference time? I think it will exceed than allocated time"
        }
      ]
    },
    {
      "id": 3179939,
      "postDate": "2025-04-16T00:08:09.213Z",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> Same as my experiment. I have some results for the listed experiments:</p>\n<ul>\n<li><p>transformer-based DETR object detector: 0.74+-&gt;0.76+</p></li>\n<li><p>conv-lstm: 0.74+-&gt;0.72+</p></li>\n<li><p>bootstrap neg train images: 0.74+-&gt;0.72+</p></li>\n<li><p>bootstrap pos train images: 0.74+-&gt;0.75+</p></li>\n<li><p>external data: 0.74+ -&gt; 0.75+</p></li>\n</ul>",
      "rawMarkdown": "@hengck23 Same as my experiment. I have some results for the listed experiments:\n\n* transformer-based DETR object detector: 0.74+->0.76+\n\n* conv-lstm: 0.74+->0.72+\n\n* bootstrap neg train images: 0.74+->0.72+\n\n* bootstrap pos train images: 0.74+->0.75+\n\n* external data: 0.74+ -> 0.75+",
      "votes": 8,
      "replies": [
        {
          "id": 3179942,
          "postDate": "2025-04-16T00:26:09.100Z",
          "content": "<p>Thanks for the information. I will update my results later. Seems that I need to change my plan. </p>",
          "rawMarkdown": "Thanks for the information. I will update my results later. Seems that I need to change my plan. ",
          "votes": 1
        },
        {
          "id": 3182377,
          "postDate": "2025-04-19T07:53:24.393Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/tom99763\" target=\"_blank\">@tom99763</a>, thanks for your insights, what do you mean by bootstrap neg train images, and bootstrap pos images.</p>",
          "rawMarkdown": "Hi @tom99763, thanks for your insights, what do you mean by bootstrap neg train images, and bootstrap pos images."
        },
        {
          "id": 3185305,
          "postDate": "2025-04-23T06:09:42.197Z",
          "content": "<p>have you tried effenet</p>",
          "rawMarkdown": "have you tried effenet"
        }
      ]
    },
    {
      "id": 3182964,
      "postDate": "2025-04-20T07:14:44.837Z",
      "content": "<p>here is my updated next plan<br>\n1) 2d-to-3d image encoder (the below shows pvtv2 b2 is used). input is 192x384x384 and produces a feature map of 16x16x16</p>\n<p>2) DETR head of say, e.g. 1 to 3 query to predict:<br>\na. has motor or not (1 class)<br>\nb. location (x,y,z)</p>\n<p>in addition, I want to use softmax to rank the  query if there is more than 1 prediction</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F06cd8c52e33d42d0baa3941c727b444e%2FSelection_218.png?generation=1745133050681400&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fb9ea2dbdbfd7f6c666234ecf267bf429%2FSelection_219.png?generation=1745132930533655&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F201f0fe966dc3eef895937018afd4a18%2FSelection_221.png?generation=1745132954594223&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "here is my updated next plan\n1) 2d-to-3d image encoder (the below shows pvtv2 b2 is used). input is 192x384x384 and produces a feature map of 16x16x16\n\n2) DETR head of say, e.g. 1 to 3 query to predict:\na. has motor or not (1 class)\nb. location (x,y,z)\n\nin addition, I want to use softmax to rank the  query if there is more than 1 prediction\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F06cd8c52e33d42d0baa3941c727b444e%2FSelection_218.png?generation=1745133050681400&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fb9ea2dbdbfd7f6c666234ecf267bf429%2FSelection_219.png?generation=1745132930533655&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F201f0fe966dc3eef895937018afd4a18%2FSelection_221.png?generation=1745132954594223&alt=media)\n\n\n",
      "votes": 3,
      "replies": [
        {
          "id": 3182977,
          "postDate": "2025-04-20T07:49:04.123Z",
          "content": "<p>What if we use keypoint detection on 3d volume (point cloud, 3d sift, 3d surf?) and use GNN to predict whether a node is a motor. I think inference is really quick using this type of method.</p>",
          "rawMarkdown": "What if we use keypoint detection on 3d volume (point cloud, 3d sift, 3d surf?) and use GNN to predict whether a node is a motor. I think inference is really quick using this type of method.",
          "votes": 1,
          "replies": [
            {
              "id": 3182992,
              "postDate": "2025-04-20T08:11:48.970Z",
              "content": "<p>i don't recommend GNN. you can use a transformer on seqence of voxel, but i find it too computation expensive.<br>\ntransformer can also model graph.</p>\n<p>so i have it done in factorised way: 3d = 2d x 1d.<br>\ni will put a notebook soon.</p>",
              "rawMarkdown": "i don't recommend GNN. you can use a transformer on seqence of voxel, but i find it too computation expensive.\ntransformer can also model graph.\n\nso i have it done in factorised way: 3d = 2d x 1d.\ni will put a notebook soon.\n",
              "votes": 1
            },
            {
              "id": 3183021,
              "postDate": "2025-04-20T09:01:40.693Z",
              "content": "<p>example notebook here:<br>\n<a href=\"https://www.kaggle.com/code/hengck23/example-2d-3d-encoder-for-detr-head\" target=\"_blank\">https://www.kaggle.com/code/hengck23/example-2d-3d-encoder-for-detr-head</a></p>",
              "rawMarkdown": "example notebook here:\nhttps://www.kaggle.com/code/hengck23/example-2d-3d-encoder-for-detr-head",
              "votes": 2
            },
            {
              "id": 3183452,
              "postDate": "2025-04-20T22:56:17.637Z",
              "content": "<p>Seems a great approach. Good work. But I concern the limited effect of modeling the relationship between image and its annotation is suited for this comp. I mean instead of directly predicting the location, what if we treat the motor as an independent object and the positive is detected when motor and bacteria are overlapped.</p>",
              "rawMarkdown": "Seems a great approach. Good work. But I concern the limited effect of modeling the relationship between image and its annotation is suited for this comp. I mean instead of directly predicting the location, what if we treat the motor as an independent object and the positive is detected when motor and bacteria are overlapped.",
              "votes": 1
            },
            {
              "id": 3183476,
              "postDate": "2025-04-20T23:54:29.193Z",
              "content": "<p>How about Gnn on predictions of yolo over input volume. I now apply detr on stack 2d feature maps= 3d feature maps of yolo</p>",
              "rawMarkdown": "How about Gnn on predictions of yolo over input volume. I now apply detr on stack 2d feature maps= 3d feature maps of yolo",
              "votes": 1
            },
            {
              "id": 3183483,
              "postDate": "2025-04-20T23:59:19.160Z",
              "content": "<p>In some input the filament is not complete or having weak signals.  This is pos if it is the only motor in the volume. </p>\n<p>In some volume there are both complete and incomplete filaments and only the more complete and obviously is annotated.</p>\n<p>By using query and ranking, we can detect both cases ( the most likely motor, instead of just motor)</p>\n<p>Location prediction is just a by product.</p>\n<p>Note that query predict both label( presence of motor) and location</p>",
              "rawMarkdown": "In some input the filament is not complete or having weak signals.  This is pos if it is the only motor in the volume. \n\nIn some volume there are both complete and incomplete filaments and only the more complete and obviously is annotated.\n\nBy using query and ranking, we can detect both cases ( the most likely motor, instead of just motor)\n\nLocation prediction is just a by product.\n\nNote that query predict both label( presence of motor) and location"
            },
            {
              "id": 3183485,
              "postDate": "2025-04-21T00:00:47.770Z",
              "content": "<p>I'm working on that now haha. But my implementation is setting very low confidence threshold and treating the predictions from yolo as keypoint, then using GNN to aggregate them.</p>",
              "rawMarkdown": "I'm working on that now haha. But my implementation is setting very low confidence threshold and treating the predictions from yolo as keypoint, then using GNN to aggregate them.",
              "votes": 1
            },
            {
              "id": 3183498,
              "postDate": "2025-04-21T00:07:47.840Z",
              "content": "<p>You need to train yolo to make it less accurate. I.e high recall but less precise. You can take more frames from annotation, eg up to 20 frames. So the gnn becomes post processing to reduce fp. Post processing cannot improve recall, so yolo must have full recall </p>",
              "rawMarkdown": "You need to train yolo to make it less accurate. I.e high recall but less precise. You can take more frames from annotation, eg up to 20 frames. So the gnn becomes post processing to reduce fp. Post processing cannot improve recall, so yolo must have full recall ",
              "votes": 1
            },
            {
              "id": 3183501,
              "postDate": "2025-04-21T00:08:46.413Z",
              "content": "<p>I think yolo is great. It provides very good initial guess.</p>",
              "rawMarkdown": "I think yolo is great. It provides very good initial guess.",
              "votes": 1
            },
            {
              "id": 3183561,
              "postDate": "2025-04-21T02:11:42.343Z",
              "content": "<p>So here is my thought: Using very good yolo model (0.76+ lb) and produce predictions with very low confidence (my experiment it produces about 20-40 predictions in a single image), then expand a radius for each predicted point -&gt; Sampling some interested point within the radius and finally connect them into 3d graph after yolo has scanned a tomo.</p>",
              "rawMarkdown": "So here is my thought: Using very good yolo model (0.76+ lb) and produce predictions with very low confidence (my experiment it produces about 20-40 predictions in a single image), then expand a radius for each predicted point -> Sampling some interested point within the radius and finally connect them into 3d graph after yolo has scanned a tomo.",
              "votes": 1
            },
            {
              "id": 3183565,
              "postDate": "2025-04-21T02:21:44.463Z",
              "content": "<p>I would suggest once the model pipeline is fixed and stable, data augmentation, synthesis, or collection is more important. Yolo can go up to 0.80 with good data selection according to one post in forum</p>",
              "rawMarkdown": "I would suggest once the model pipeline is fixed and stable, data augmentation, synthesis, or collection is more important. Yolo can go up to 0.80 with good data selection according to one post in forum\n"
            },
            {
              "id": 3183753,
              "postDate": "2025-04-21T08:44:21.847Z",
              "content": "<p>Based on my experiment, SPPF feature in yolo10x gives huge boosting for modeling with GNN. </p>",
              "rawMarkdown": "Based on my experiment, SPPF feature in yolo10x gives huge boosting for modeling with GNN. ",
              "votes": 1
            },
            {
              "id": 3201833,
              "postDate": "2025-05-14T13:07:45.790Z",
              "content": "<p>In theory (If I am not wrong), using only sppf features for GNN (3d encoder here) would not be helpful as spatial information will be lost with 30*30 maps no?? u must using some other features as well then only you ll give more information to the GNN</p>",
              "rawMarkdown": "In theory (If I am not wrong), using only sppf features for GNN (3d encoder here) would not be helpful as spatial information will be lost with 30*30 maps no?? u must using some other features as well then only you ll give more information to the GNN\n    "
            },
            {
              "id": 3201851,
              "postDate": "2025-05-14T13:35:02.153Z",
              "content": "<p>Yeah I actually did massive engineering to verify which can work.</p>",
              "rawMarkdown": "Yeah I actually did massive engineering to verify which can work."
            },
            {
              "id": 3202093,
              "postDate": "2025-05-14T20:23:21.047Z",
              "content": "<p>also one more thing , I m a bit late , what problems (biggest) are with public yolo's (0.76 lb to 0.82+)</p>\n<p>-&gt;  are they missing tomos with motors  ( I don't think this is it )<br>\n-&gt;  are they predicting wrong coords which leads to FN ( possible )<br>\n-&gt;  or heavy predicting for tomos without motors??</p>\n<p>See All these ques are for me to work on a 2d+3d model because full 3D is very hard to train</p>\n<p>Bonus ques : Yolo's is suffering from domain shift more ?? (but with 3d models that can be covered (different orientations of bacteria in 3d space)) </p>\n<p>What were your number of epochs and training time and memory usage (If you are comfortable)</p>",
              "rawMarkdown": "also one more thing , I m a bit late , what problems (biggest) are with public yolo's (0.76 lb to 0.82+)\n\n->  are they missing tomos with motors  ( I don't think this is it )\n->  are they predicting wrong coords which leads to FN ( possible )\n->  or heavy predicting for tomos without motors??\n\nSee All these ques are for me to work on a 2d+3d model because full 3D is very hard to train\n\nBonus ques : Yolo's is suffering from domain shift more ?? (but with 3d models that can be covered (different orientations of bacteria in 3d space)) \n\nWhat were your number of epochs and training time and memory usage (If you are comfortable)"
            }
          ]
        }
      ]
    },
    {
      "id": 3180636,
      "postDate": "2025-04-16T22:14:14.167Z",
      "content": "<p><a href=\"https://www.embopress.org/doi/full/10.1038/emboj.2011.186\" target=\"_blank\">https://www.embopress.org/doi/full/10.1038/emboj.2011.186</a> <br>\nStructural diversity of bacterial flagellar motors</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fb5a174f6bf82edaab98f0cb78a783174%2FSelection_202.png?generation=1744841593756280&amp;alt=media\" alt=\"\"></p>\n<p>I have  a feeling that the filament is not always present</p>",
      "rawMarkdown": "https://www.embopress.org/doi/full/10.1038/emboj.2011.186 \nStructural diversity of bacterial flagellar motors\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fb5a174f6bf82edaab98f0cb78a783174%2FSelection_202.png?generation=1744841593756280&alt=media)\n\nI have  a feeling that the filament is not always present",
      "votes": 3,
      "replies": [
        {
          "id": 3180639,
          "postDate": "2025-04-16T22:23:29.553Z",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> So you mean there exists ambiguous problem?  The filament is not present but annotated as positive.</p>",
          "rawMarkdown": "@hengck23 So you mean there exists ambiguous problem?  The filament is not present but annotated as positive.",
          "votes": 1,
          "replies": [
            {
              "id": 3180642,
              "postDate": "2025-04-16T22:39:31.557Z",
              "content": "<p>Maybe we need to ask the host on this. When I goggle for lab tomography of bacteria motor, I do see motor assembly without filament . Alternatively, we can do a test probe</p>",
              "rawMarkdown": "Maybe we need to ask the host on this. When I goggle for lab tomography of bacteria motor, I do see motor assembly without filament . Alternatively, we can do a test probe"
            },
            {
              "id": 3180683,
              "postDate": "2025-04-17T02:09:02.290Z",
              "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F4ebcfa3abccbb08a6cf1b7f65d2eda74%2FSelection_999(8080).png?generation=1744855693512235&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F0b342b98893597d50b4306a2248b88b8%2FSelection_999(8079).png?generation=1744855703389117&amp;alt=media\" alt=\"\"></p>\n<p>i wonder if cchatgpt can do a websearch and compile a list of motor tomograhy for different bacteria</p>",
              "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F4ebcfa3abccbb08a6cf1b7f65d2eda74%2FSelection_999(8080).png?generation=1744855693512235&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F0b342b98893597d50b4306a2248b88b8%2FSelection_999(8079).png?generation=1744855703389117&alt=media)\n\ni wonder if cchatgpt can do a websearch and compile a list of motor tomograhy for different bacteria",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3180930,
      "postDate": "2025-04-17T09:26:29.033Z",
      "content": "<p>easy way to denoise. just average or apply median across frames (assuming motor object is aligned)<br>\nbut I haven't done experiment to see if this would improve results yet</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F809ca3181413a4f40dc8eeed98abec03%2FSelection_208.png?generation=1744882441568674&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Ff20176a2f8614b7c7e9c6f6cebf9b330%2FSelection_205.png?generation=1744881945202858&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F9105548f9a490f14c1acf9bd1db6301c%2FSelection_206.png?generation=1744881962894427&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fbed0953fcdb5307535efedf320cb68a5%2FSelection_207.png?generation=1744881974491683&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "easy way to denoise. just average or apply median across frames (assuming motor object is aligned)\nbut I haven't done experiment to see if this would improve results yet\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F809ca3181413a4f40dc8eeed98abec03%2FSelection_208.png?generation=1744882441568674&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Ff20176a2f8614b7c7e9c6f6cebf9b330%2FSelection_205.png?generation=1744881945202858&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F9105548f9a490f14c1acf9bd1db6301c%2FSelection_206.png?generation=1744881962894427&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fbed0953fcdb5307535efedf320cb68a5%2FSelection_207.png?generation=1744881974491683&alt=media)",
      "votes": 4,
      "replies": [
        {
          "id": 3180947,
          "postDate": "2025-04-17T10:05:51.027Z",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> But I notice removing some peak signals degrades the performance.</p>",
          "rawMarkdown": "@hengck23 But I notice removing some peak signals degrades the performance.",
          "votes": 1,
          "replies": [
            {
              "id": 3183574,
              "postDate": "2025-04-21T02:47:32.780Z",
              "content": "<p>Use 3 channel input, single slice, average over 5, average over 30</p>",
              "rawMarkdown": "Use 3 channel input, single slice, average over 5, average over 30"
            }
          ]
        }
      ]
    },
    {
      "id": 3183488,
      "postDate": "2025-04-21T00:04:29.727Z",
      "content": "<p>Another method is to use multi label, class1 for annotated z, class n for n slices from annotation. Then another model to combine all labels into single object</p>",
      "rawMarkdown": "Another method is to use multi label, class1 for annotated z, class n for n slices from annotation. Then another model to combine all labels into single object",
      "votes": 1
    },
    {
      "id": 3181104,
      "postDate": "2025-04-17T13:20:20.047Z",
      "content": "<p>Currently working on 2d seg models, and 0.76~ is my temp best</p>",
      "rawMarkdown": "Currently working on 2d seg models, and 0.76~ is my temp best",
      "votes": 1,
      "replies": [
        {
          "id": 3181503,
          "postDate": "2025-04-18T01:29:27.943Z",
          "content": "<p>thanks for the comment. performance is usually limited to data. so if yolo works, other model like 2d or 3d segmentation should also work. I think the trainer in ultralytics is good. despite only trained on limited pos images, it has very low false pos rate.</p>",
          "rawMarkdown": "thanks for the comment. performance is usually limited to data. so if yolo works, other model like 2d or 3d segmentation should also work. I think the trainer in ultralytics is good. despite only trained on limited pos images, it has very low false pos rate."
        }
      ]
    },
    {
      "id": 3180805,
      "postDate": "2025-04-17T06:14:16.143Z",
      "content": "<p>Brother Frog has finally come!</p>",
      "rawMarkdown": "Brother Frog has finally come!",
      "votes": 1
    },
    {
      "id": 3180167,
      "postDate": "2025-04-16T08:35:15.737Z",
      "content": "<p>Will u try using 2d or 3d segmentation models?</p>",
      "rawMarkdown": "Will u try using 2d or 3d segmentation models?",
      "votes": 1
    },
    {
      "id": 3185574,
      "postDate": "2025-04-23T14:16:06.443Z",
      "content": "<p>thanks for sharing such great results…</p>",
      "rawMarkdown": "thanks for sharing such great results..."
    },
    {
      "id": 3182924,
      "postDate": "2025-04-20T05:18:43.407Z",
      "content": "<p>感谢您的分享，这给我很多启示</p>",
      "rawMarkdown": "感谢您的分享，这给我很多启示"
    },
    {
      "id": 3180833,
      "postDate": "2025-04-17T07:05:17.203Z",
      "content": "<p>thank you for sharing your insight it is great!</p>",
      "rawMarkdown": "thank you for sharing your insight it is great!"
    },
    {
      "id": 3180280,
      "postDate": "2025-04-16T11:27:04.810Z",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> <br>\nThank you for sharing such great results.</p>\n<p>I have a question about the training data. Did you use the code (the one that uses images centered on the motor with ±4 frames) to create the training dataset, like the sample provided? Or did you come up with a different approach?</p>\n<p>I'm using the dataset , but I'm not getting scores that high like you..</p>",
      "rawMarkdown": "@hengck23 \nThank you for sharing such great results.\n\nI have a question about the training data. Did you use the code (the one that uses images centered on the motor with ±4 frames) to create the training dataset, like the sample provided? Or did you come up with a different approach?\n\nI'm using the dataset , but I'm not getting scores that high like you..",
      "replies": [
        {
          "id": 3180284,
          "postDate": "2025-04-16T11:34:14.663Z",
          "content": "<p>current results is based on TRUST=4. but i make the patch, see <a href=\"https://www.kaggle.com/code/andrewjdarley/parse-data/comments#3175684\" target=\"_blank\">https://www.kaggle.com/code/andrewjdarley/parse-data/comments#3175684</a>.</p>\n<p>The LB score is sensitive on data used. If any motor object (or motor like object) is trained as neg, the LB score will dropped much.</p>",
          "rawMarkdown": "current results is based on TRUST=4. but i make the patch, see https://www.kaggle.com/code/andrewjdarley/parse-data/comments#3175684.\n\nThe LB score is sensitive on data used. If any motor object (or motor like object) is trained as neg, the LB score will dropped much.",
          "votes": 2,
          "replies": [
            {
              "id": 3180330,
              "postDate": "2025-04-16T13:09:11.263Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 3180700,
              "postDate": "2025-04-17T02:44:43.620Z",
              "content": "<p>May I ask if this means you only used the slice with motor=1?</p>",
              "rawMarkdown": "May I ask if this means you only used the slice with motor=1?"
            }
          ]
        }
      ]
    },
    {
      "id": 3179951,
      "postDate": "2025-04-16T00:44:36.100Z",
      "content": "<p>newbie question, why can we change the image size in yolo model? <br>\nsince the default image size is 640x640, i thought those pre-trained models were trained with image size 640x640. So, i think i need to keep the training images with that size.</p>",
      "rawMarkdown": "newbie question, why can we change the image size in yolo model? \nsince the default image size is 640x640, i thought those pre-trained models were trained with image size 640x640. So, i think i need to keep the training images with that size.",
      "replies": [
        {
          "id": 3179953,
          "postDate": "2025-04-16T00:50:29.630Z",
          "content": "<p>The motor object is small. If u want to keep 640, u can crop. Else increase image size. Since we are fine running, yolo should adapt to new size. We can also try upsize, eg 1280, 2048</p>",
          "rawMarkdown": "The motor object is small. If u want to keep 640, u can crop. Else increase image size. Since we are fine running, yolo should adapt to new size. We can also try upsize, eg 1280, 2048",
          "votes": 2
        }
      ]
    },
    {
      "id": 3184284,
      "postDate": "2025-04-21T23:03:31.937Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3179926,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-04-15T23:28:17.487000",
      "content": "<p>details for yolo11L, imgsize=960, train= pos images only : lb0.744<br>\ndataset</p>\n<pre><code>yaml_config = {\n    : {: },\n     : ,\n    : ,\n    : ,\n\n    \n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n}\n</code></pre>\n<p>trainer</p>\n<pre><code>    result = model.train(\n        =yaml_file,\n        =50,\n        =cfg.batch_size, #16\n        =cfg.imgsz, #960,\n        =,\n        =1e-4,\n        =0.1,\n        =0,\n        =0.1,\n        =out_dir,\n\n        =,\n        =cfg.experiment,\n\n        =100,   \n        =2,  \n        =,   \n        =  \n    )\n</code></pre>",
      "votes": 9,
      "replies": [
        {
          "id": 3180159,
          "author_name": "Zaakcii Ru",
          "author_url": "",
          "post_date": "2025-04-16T08:21:15.100000",
          "content": "<p>As always informative and interesting. RTDETR is implemented in Ultralytics, in case anyone did not pay attention to it</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 3180185,
          "author_name": "Jeroen Cottaar",
          "author_url": "",
          "post_date": "2025-04-16T09:09:59.493000",
          "content": "<p>Thanks a lot for sharing this!</p>\n<p>Surprised to see such a low learning rate with a relatively small number of epochs. Are you willing to share some loss curves for your training?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3180213,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2025-04-16T09:49:30.260000",
              "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fffad3b8eef78760e3120e77f6b8462c8%2Fdfl_loss_curve.png?generation=1744796749619799&amp;alt=media\" alt=\"\"></p>\n<p>loss curve for  yolo11L</p>\n<p>local cv</p>\n<pre><code>        lb, more = compute_lb(solution, submission, =1000, =2)\n        (th, lb, more[], more[], more[])\n\n** pos tomo_id only **\n{: 1, : 16, : 0.1, : 24, : 0.5, : , : }\n0.1 0.9872611464968152 0.9919999999999999 1.0 0.9841269841269841\n0.2 0.9872611464968152 0.9919999999999999 1.0 0.9841269841269841\n0.3 0.9872611464968152 0.9919999999999999 1.0 0.9841269841269841\n0.4 0.9872611464968152 0.9919999999999999 1.0 0.9841269841269841\n0.5 0.9872611464968152 0.9919999999999999 1.0 0.9841269841269841\n0.6 0.9872611464968152 0.9919999999999999 1.0 0.9841269841269841\n0.7 0.97444089456869 0.9838709677419354 1.0 0.9682539682539683\n0.75 0.8823529411764707 0.923076923076923 1.0 0.8571428571428571\n0.8 0.29850746268656714 0.4050632911392405 1.0 0.25396825396825395\n0.85 0.0 0.0 0.0 0.0\n0.9 0.0 0.0 0.0 0.0\n\n** pos tomo_id + neg tomo_id = 120 tomo_id **\n{: 1, : 16, : 0.1, : 24, : 0.5, : , : }\n0.1 0.8516483516483516 0.7085714285714286 0.5535714285714286 0.9841269841269841\n0.2 0.8516483516483516 0.7085714285714286 0.5535714285714286 0.9841269841269841\n0.3 0.8516483516483516 0.7085714285714286 0.5535714285714286 0.9841269841269841\n0.4 0.861111111111111 0.7251461988304093 0.5740740740740741 0.9841269841269841\n0.5 0.8757062146892655 0.7515151515151515 0.6078431372549019 0.9841269841269841\n0.6 0.9198813056379821 0.8378378378378377 0.7294117647058823 0.9841269841269841\n0.7 0.9442724458204335 0.9104477611940299 0.8591549295774648 0.9682539682539683\n0.75 0.8653846153846153 0.8780487804878048 0.9 0.8571428571428571\n0.8 0.2952029520295203 0.3902439024390244 0.8421052631578947 0.25396825396825395\n0.85 0.0 0.0 0.0 0.0\n0.9 0.0 0.0 0.0 0.0\n</code></pre>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3181166,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-04-17T14:29:34.500000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 3180878,
          "author_name": "taokj",
          "author_url": "",
          "post_date": "2025-04-17T08:29:32.560000",
          "content": "<p>Hi, I'm wondering if “mosaic” and“ mixup” data enhancements have a positive effect on your model. In one of my few experiments, these two enhancements seem to degrade model performance🌹</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3207273,
          "author_name": "Rustam Bazarbayev",
          "author_url": "",
          "post_date": "2025-05-22T14:04:42.413000",
          "content": "<p>With this image size what was the inference time? I think it will exceed than allocated time</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3179939,
      "author_name": "Tom",
      "author_url": "",
      "post_date": "2025-04-16T00:08:09.213000",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> Same as my experiment. I have some results for the listed experiments:</p>\n<ul>\n<li><p>transformer-based DETR object detector: 0.74+-&gt;0.76+</p></li>\n<li><p>conv-lstm: 0.74+-&gt;0.72+</p></li>\n<li><p>bootstrap neg train images: 0.74+-&gt;0.72+</p></li>\n<li><p>bootstrap pos train images: 0.74+-&gt;0.75+</p></li>\n<li><p>external data: 0.74+ -&gt; 0.75+</p></li>\n</ul>",
      "votes": 8,
      "replies": [
        {
          "id": 3179942,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2025-04-16T00:26:09.100000",
          "content": "<p>Thanks for the information. I will update my results later. Seems that I need to change my plan. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 3182377,
          "author_name": "Gowri Shankar Penugonda",
          "author_url": "",
          "post_date": "2025-04-19T07:53:24.393000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/tom99763\" target=\"_blank\">@tom99763</a>, thanks for your insights, what do you mean by bootstrap neg train images, and bootstrap pos images.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3185305,
          "author_name": "thisArmin",
          "author_url": "",
          "post_date": "2025-04-23T06:09:42.197000",
          "content": "<p>have you tried effenet</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3182964,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-04-20T07:14:44.837000",
      "content": "<p>here is my updated next plan<br>\n1) 2d-to-3d image encoder (the below shows pvtv2 b2 is used). input is 192x384x384 and produces a feature map of 16x16x16</p>\n<p>2) DETR head of say, e.g. 1 to 3 query to predict:<br>\na. has motor or not (1 class)<br>\nb. location (x,y,z)</p>\n<p>in addition, I want to use softmax to rank the  query if there is more than 1 prediction</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F06cd8c52e33d42d0baa3941c727b444e%2FSelection_218.png?generation=1745133050681400&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fb9ea2dbdbfd7f6c666234ecf267bf429%2FSelection_219.png?generation=1745132930533655&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F201f0fe966dc3eef895937018afd4a18%2FSelection_221.png?generation=1745132954594223&amp;alt=media\" alt=\"\"></p>",
      "votes": 3,
      "replies": [
        {
          "id": 3182977,
          "author_name": "Tom",
          "author_url": "",
          "post_date": "2025-04-20T07:49:04.123000",
          "content": "<p>What if we use keypoint detection on 3d volume (point cloud, 3d sift, 3d surf?) and use GNN to predict whether a node is a motor. I think inference is really quick using this type of method.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3182992,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2025-04-20T08:11:48.970000",
              "content": "<p>i don't recommend GNN. you can use a transformer on seqence of voxel, but i find it too computation expensive.<br>\ntransformer can also model graph.</p>\n<p>so i have it done in factorised way: 3d = 2d x 1d.<br>\ni will put a notebook soon.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3183021,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2025-04-20T09:01:40.693000",
              "content": "<p>example notebook here:<br>\n<a href=\"https://www.kaggle.com/code/hengck23/example-2d-3d-encoder-for-detr-head\" target=\"_blank\">https://www.kaggle.com/code/hengck23/example-2d-3d-encoder-for-detr-head</a></p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3183452,
              "author_name": "Tom",
              "author_url": "",
              "post_date": "2025-04-20T22:56:17.637000",
              "content": "<p>Seems a great approach. Good work. But I concern the limited effect of modeling the relationship between image and its annotation is suited for this comp. I mean instead of directly predicting the location, what if we treat the motor as an independent object and the positive is detected when motor and bacteria are overlapped.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3183476,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2025-04-20T23:54:29.193000",
              "content": "<p>How about Gnn on predictions of yolo over input volume. I now apply detr on stack 2d feature maps= 3d feature maps of yolo</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3183483,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2025-04-20T23:59:19.160000",
              "content": "<p>In some input the filament is not complete or having weak signals.  This is pos if it is the only motor in the volume. </p>\n<p>In some volume there are both complete and incomplete filaments and only the more complete and obviously is annotated.</p>\n<p>By using query and ranking, we can detect both cases ( the most likely motor, instead of just motor)</p>\n<p>Location prediction is just a by product.</p>\n<p>Note that query predict both label( presence of motor) and location</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3183485,
              "author_name": "Tom",
              "author_url": "",
              "post_date": "2025-04-21T00:00:47.770000",
              "content": "<p>I'm working on that now haha. But my implementation is setting very low confidence threshold and treating the predictions from yolo as keypoint, then using GNN to aggregate them.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3183498,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2025-04-21T00:07:47.840000",
              "content": "<p>You need to train yolo to make it less accurate. I.e high recall but less precise. You can take more frames from annotation, eg up to 20 frames. So the gnn becomes post processing to reduce fp. Post processing cannot improve recall, so yolo must have full recall </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3183501,
              "author_name": "Tom",
              "author_url": "",
              "post_date": "2025-04-21T00:08:46.413000",
              "content": "<p>I think yolo is great. It provides very good initial guess.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3183561,
              "author_name": "Tom",
              "author_url": "",
              "post_date": "2025-04-21T02:11:42.343000",
              "content": "<p>So here is my thought: Using very good yolo model (0.76+ lb) and produce predictions with very low confidence (my experiment it produces about 20-40 predictions in a single image), then expand a radius for each predicted point -&gt; Sampling some interested point within the radius and finally connect them into 3d graph after yolo has scanned a tomo.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3183565,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2025-04-21T02:21:44.463000",
              "content": "<p>I would suggest once the model pipeline is fixed and stable, data augmentation, synthesis, or collection is more important. Yolo can go up to 0.80 with good data selection according to one post in forum</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3183753,
              "author_name": "Tom",
              "author_url": "",
              "post_date": "2025-04-21T08:44:21.847000",
              "content": "<p>Based on my experiment, SPPF feature in yolo10x gives huge boosting for modeling with GNN. </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3201833,
              "author_name": "Surya_trainer",
              "author_url": "",
              "post_date": "2025-05-14T13:07:45.790000",
              "content": "<p>In theory (If I am not wrong), using only sppf features for GNN (3d encoder here) would not be helpful as spatial information will be lost with 30*30 maps no?? u must using some other features as well then only you ll give more information to the GNN</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3201851,
              "author_name": "Tom",
              "author_url": "",
              "post_date": "2025-05-14T13:35:02.153000",
              "content": "<p>Yeah I actually did massive engineering to verify which can work.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3202093,
              "author_name": "Surya_trainer",
              "author_url": "",
              "post_date": "2025-05-14T20:23:21.047000",
              "content": "<p>also one more thing , I m a bit late , what problems (biggest) are with public yolo's (0.76 lb to 0.82+)</p>\n<p>-&gt;  are they missing tomos with motors  ( I don't think this is it )<br>\n-&gt;  are they predicting wrong coords which leads to FN ( possible )<br>\n-&gt;  or heavy predicting for tomos without motors??</p>\n<p>See All these ques are for me to work on a 2d+3d model because full 3D is very hard to train</p>\n<p>Bonus ques : Yolo's is suffering from domain shift more ?? (but with 3d models that can be covered (different orientations of bacteria in 3d space)) </p>\n<p>What were your number of epochs and training time and memory usage (If you are comfortable)</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3180636,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-04-16T22:14:14.167000",
      "content": "<p><a href=\"https://www.embopress.org/doi/full/10.1038/emboj.2011.186\" target=\"_blank\">https://www.embopress.org/doi/full/10.1038/emboj.2011.186</a> <br>\nStructural diversity of bacterial flagellar motors</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fb5a174f6bf82edaab98f0cb78a783174%2FSelection_202.png?generation=1744841593756280&amp;alt=media\" alt=\"\"></p>\n<p>I have  a feeling that the filament is not always present</p>",
      "votes": 3,
      "replies": [
        {
          "id": 3180639,
          "author_name": "Tom",
          "author_url": "",
          "post_date": "2025-04-16T22:23:29.553000",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> So you mean there exists ambiguous problem?  The filament is not present but annotated as positive.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3180642,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2025-04-16T22:39:31.557000",
              "content": "<p>Maybe we need to ask the host on this. When I goggle for lab tomography of bacteria motor, I do see motor assembly without filament . Alternatively, we can do a test probe</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3180683,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2025-04-17T02:09:02.290000",
              "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F4ebcfa3abccbb08a6cf1b7f65d2eda74%2FSelection_999(8080).png?generation=1744855693512235&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F0b342b98893597d50b4306a2248b88b8%2FSelection_999(8079).png?generation=1744855703389117&amp;alt=media\" alt=\"\"></p>\n<p>i wonder if cchatgpt can do a websearch and compile a list of motor tomograhy for different bacteria</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3180930,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-04-17T09:26:29.033000",
      "content": "<p>easy way to denoise. just average or apply median across frames (assuming motor object is aligned)<br>\nbut I haven't done experiment to see if this would improve results yet</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F809ca3181413a4f40dc8eeed98abec03%2FSelection_208.png?generation=1744882441568674&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Ff20176a2f8614b7c7e9c6f6cebf9b330%2FSelection_205.png?generation=1744881945202858&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F9105548f9a490f14c1acf9bd1db6301c%2FSelection_206.png?generation=1744881962894427&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fbed0953fcdb5307535efedf320cb68a5%2FSelection_207.png?generation=1744881974491683&amp;alt=media\" alt=\"\"></p>",
      "votes": 4,
      "replies": [
        {
          "id": 3180947,
          "author_name": "Tom",
          "author_url": "",
          "post_date": "2025-04-17T10:05:51.027000",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> But I notice removing some peak signals degrades the performance.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3183574,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2025-04-21T02:47:32.780000",
              "content": "<p>Use 3 channel input, single slice, average over 5, average over 30</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3183488,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-04-21T00:04:29.727000",
      "content": "<p>Another method is to use multi label, class1 for annotated z, class n for n slices from annotation. Then another model to combine all labels into single object</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3181104,
      "author_name": "Yksin Young",
      "author_url": "",
      "post_date": "2025-04-17T13:20:20.047000",
      "content": "<p>Currently working on 2d seg models, and 0.76~ is my temp best</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3181503,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2025-04-18T01:29:27.943000",
          "content": "<p>thanks for the comment. performance is usually limited to data. so if yolo works, other model like 2d or 3d segmentation should also work. I think the trainer in ultralytics is good. despite only trained on limited pos images, it has very low false pos rate.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3180805,
      "author_name": "Switch9527",
      "author_url": "",
      "post_date": "2025-04-17T06:14:16.143000",
      "content": "<p>Brother Frog has finally come!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3180167,
      "author_name": "eikyou",
      "author_url": "",
      "post_date": "2025-04-16T08:35:15.737000",
      "content": "<p>Will u try using 2d or 3d segmentation models?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3185574,
      "author_name": "Sarah Arshad",
      "author_url": "",
      "post_date": "2025-04-23T14:16:06.443000",
      "content": "<p>thanks for sharing such great results…</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3182924,
      "author_name": "majing2003",
      "author_url": "",
      "post_date": "2025-04-20T05:18:43.407000",
      "content": "<p>感谢您的分享，这给我很多启示</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3180833,
      "author_name": "boo chang gyu",
      "author_url": "",
      "post_date": "2025-04-17T07:05:17.203000",
      "content": "<p>thank you for sharing your insight it is great!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3180280,
      "author_name": "shiba-inu",
      "author_url": "",
      "post_date": "2025-04-16T11:27:04.810000",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> <br>\nThank you for sharing such great results.</p>\n<p>I have a question about the training data. Did you use the code (the one that uses images centered on the motor with ±4 frames) to create the training dataset, like the sample provided? Or did you come up with a different approach?</p>\n<p>I'm using the dataset , but I'm not getting scores that high like you..</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3180284,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2025-04-16T11:34:14.663000",
          "content": "<p>current results is based on TRUST=4. but i make the patch, see <a href=\"https://www.kaggle.com/code/andrewjdarley/parse-data/comments#3175684\" target=\"_blank\">https://www.kaggle.com/code/andrewjdarley/parse-data/comments#3175684</a>.</p>\n<p>The LB score is sensitive on data used. If any motor object (or motor like object) is trained as neg, the LB score will dropped much.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 3180330,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-04-16T13:09:11.263000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3180700,
              "author_name": "Heeler-Deer",
              "author_url": "",
              "post_date": "2025-04-17T02:44:43.620000",
              "content": "<p>May I ask if this means you only used the slice with motor=1?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3179951,
      "author_name": "MengYe",
      "author_url": "",
      "post_date": "2025-04-16T00:44:36.100000",
      "content": "<p>newbie question, why can we change the image size in yolo model? <br>\nsince the default image size is 640x640, i thought those pre-trained models were trained with image size 640x640. So, i think i need to keep the training images with that size.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3179953,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2025-04-16T00:50:29.630000",
          "content": "<p>The motor object is small. If u want to keep 640, u can crop. Else increase image size. Since we are fine running, yolo should adapt to new size. We can also try upsize, eg 1280, 2048</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 3184284,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-04-21T23:03:31.937000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3179924": "**The test data has been relabelled and lb has been rescored.**\n**This thread is based on old test data and will be closed.**\n**Please wait for the new thread ...**\n\n\n---\n\nreference notebook:\nhttps://www.kaggle.com/code/hengck23/yolo-experiments\n\n[1] apr-16:\n- yolo11L, imgsize=960, train= pos images only : lb0.744 \n\n---\ntodo:\n- transformer-based DETR object detector\n- multi-slice yolo, etc (maybe add conv-lstm, 3d conv or transformer fuser?)\n- bootstrap neg train images\n- bootstrap pos train images \n- external data\n- synthetic data via video diffusion?\n\n\n\n---\n\nto be updated as experiment proceeds ...this thread will be cias",
    "3179926": "details for yolo11L, imgsize=960, train= pos images only : lb0.744\ndataset\n```\nyaml_config = {\n    'names': {0: 'motor'},\n    'path' : f'{YOLO_DIR}',\n    'train': f'{YOLO_DIR}/train_list.fold{cfg.fold}.txt',\n    'val': f'{YOLO_DIR}/valid_list.fold{cfg.fold}.txt',\n\n    #augmentation ---\n    'mosaic': 1.0,\n    'close_mosaic': 10,\n    'mixup': 0.4,\n    'flipud': 0.5,\n    'scale': 0.25,\n    'degrees': 45,\n}#rest are default\n```\n\ntrainer\n\n```\n    result = model.train(\n        data=yaml_file,\n        epochs=50,\n        batch=cfg.batch_size, #16\n        imgsz=cfg.imgsz, #960,\n        optimizer='AdamW',\n        lr0=1e-4,\n        lrf=0.1,\n        warmup_epochs=0,\n        dropout=0.1,\n        project=out_dir,\n      \n        exist_ok=True,\n        name=cfg.experiment,\n     \n        patience=100,   \n        save_period=2,  \n        val=True,   \n        verbose=True  \n    )\n\n```",
    "3179939": "@hengck23 Same as my experiment. I have some results for the listed experiments:\n\n* transformer-based DETR object detector: 0.74+->0.76+\n\n* conv-lstm: 0.74+->0.72+\n\n* bootstrap neg train images: 0.74+->0.72+\n\n* bootstrap pos train images: 0.74+->0.75+\n\n* external data: 0.74+ -> 0.75+",
    "3182964": "here is my updated next plan\n1) 2d-to-3d image encoder (the below shows pvtv2 b2 is used). input is 192x384x384 and produces a feature map of 16x16x16\n\n2) DETR head of say, e.g. 1 to 3 query to predict:\na. has motor or not (1 class)\nb. location (x,y,z)\n\nin addition, I want to use softmax to rank the  query if there is more than 1 prediction\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F06cd8c52e33d42d0baa3941c727b444e%2FSelection_218.png?generation=1745133050681400&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fb9ea2dbdbfd7f6c666234ecf267bf429%2FSelection_219.png?generation=1745132930533655&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F201f0fe966dc3eef895937018afd4a18%2FSelection_221.png?generation=1745132954594223&alt=media)\n\n\n",
    "3180636": "https://www.embopress.org/doi/full/10.1038/emboj.2011.186 \nStructural diversity of bacterial flagellar motors\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fb5a174f6bf82edaab98f0cb78a783174%2FSelection_202.png?generation=1744841593756280&alt=media)\n\nI have  a feeling that the filament is not always present",
    "3180930": "easy way to denoise. just average or apply median across frames (assuming motor object is aligned)\nbut I haven't done experiment to see if this would improve results yet\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F809ca3181413a4f40dc8eeed98abec03%2FSelection_208.png?generation=1744882441568674&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Ff20176a2f8614b7c7e9c6f6cebf9b330%2FSelection_205.png?generation=1744881945202858&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F9105548f9a490f14c1acf9bd1db6301c%2FSelection_206.png?generation=1744881962894427&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fbed0953fcdb5307535efedf320cb68a5%2FSelection_207.png?generation=1744881974491683&alt=media)",
    "3183488": "Another method is to use multi label, class1 for annotated z, class n for n slices from annotation. Then another model to combine all labels into single object",
    "3181104": "Currently working on 2d seg models, and 0.76~ is my temp best",
    "3180805": "Brother Frog has finally come!",
    "3180167": "Will u try using 2d or 3d segmentation models?",
    "3185574": "thanks for sharing such great results...",
    "3182924": "感谢您的分享，这给我很多启示",
    "3180833": "thank you for sharing your insight it is great!",
    "3180280": "@hengck23 \nThank you for sharing such great results.\n\nI have a question about the training data. Did you use the code (the one that uses images centered on the motor with ±4 frames) to create the training dataset, like the sample provided? Or did you come up with a different approach?\n\nI'm using the dataset , but I'm not getting scores that high like you..",
    "3179951": "newbie question, why can we change the image size in yolo model? \nsince the default image size is 640x640, i thought those pre-trained models were trained with image size 640x640. So, i think i need to keep the training images with that size.",
    "3184284": ""
  }
}