{
  "id": 569921,
  "title": "More Motor Annotations",
  "url": "/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/569921",
  "author_name": "Bartley",
  "post_date": "2025-03-25T00:44:25.141000",
  "votes": 147,
  "comment_count": 45,
  "views": 0,
  "content": "<p>I wanted to share some external data that I have been collecting and annotating over the past week. This has improved CV/LB scores for my pipeline and hopefully it helps others too!</p>\n<p>The dataset contains 1617 flagellar motor annotations and 1288 tomograms, all sourced from 62 datasets from the CZII data portal. You can visualize the annotations and preprocessing pipeline at the links below.</p>\n<p>Dataset <a href=\"https://www.kaggle.com/datasets/brendanartley/cryoet-flagellar-motors-dataset\" target=\"_blank\">here</a>.<br>\nCode <a href=\"https://www.kaggle.com/code/brendanartley/flagellar-motors-dataset-code\" target=\"_blank\">here</a>.</p>\n<p>Happy Kaggling 😀</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F9e36df03f6bb44ee14dc22458f9240fc%2Fimg.jpg?generation=1742863339495044&amp;alt=media\" alt=\"image_sample\"></p>",
  "messages": [
    {
      "id": 3158828,
      "postDate": "2025-03-25T00:44:25.143Z",
      "content": "<p>I wanted to share some external data that I have been collecting and annotating over the past week. This has improved CV/LB scores for my pipeline and hopefully it helps others too!</p>\n<p>The dataset contains 1617 flagellar motor annotations and 1288 tomograms, all sourced from 62 datasets from the CZII data portal. You can visualize the annotations and preprocessing pipeline at the links below.</p>\n<p>Dataset <a href=\"https://www.kaggle.com/datasets/brendanartley/cryoet-flagellar-motors-dataset\" target=\"_blank\">here</a>.<br>\nCode <a href=\"https://www.kaggle.com/code/brendanartley/flagellar-motors-dataset-code\" target=\"_blank\">here</a>.</p>\n<p>Happy Kaggling 😀</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F9e36df03f6bb44ee14dc22458f9240fc%2Fimg.jpg?generation=1742863339495044&amp;alt=media\" alt=\"image_sample\"></p>",
      "rawMarkdown": "I wanted to share some external data that I have been collecting and annotating over the past week. This has improved CV/LB scores for my pipeline and hopefully it helps others too!\n\nThe dataset contains 1617 flagellar motor annotations and 1288 tomograms, all sourced from 62 datasets from the CZII data portal. You can visualize the annotations and preprocessing pipeline at the links below.\n\nDataset [here](https://www.kaggle.com/datasets/brendanartley/cryoet-flagellar-motors-dataset).\nCode [here](https://www.kaggle.com/code/brendanartley/flagellar-motors-dataset-code).\n\nHappy Kaggling 😀\n\n![image_sample](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F9e36df03f6bb44ee14dc22458f9240fc%2Fimg.jpg?generation=1742863339495044&alt=media)\n\n",
      "votes": 146
    },
    {
      "id": 3160527,
      "postDate": "2025-03-26T22:05:33.307Z",
      "content": "<p>Appreciate all your work and sharing</p>",
      "rawMarkdown": "Appreciate all your work and sharing",
      "votes": 11
    },
    {
      "id": 3173157,
      "postDate": "2025-04-07T15:42:34.973Z",
      "content": "<p><strong>Update:</strong></p>\n<p>After some feedback, I uploaded a <code>jpgs</code> folder to this dataset to match the competition data format. This should make integration with existing pipelines easier.</p>\n<p>In addition, if anyone finds more tomograms to annotate, please let me know!</p>",
      "rawMarkdown": "**Update:**\n\nAfter some feedback, I uploaded a `jpgs` folder to this dataset to match the competition data format. This should make integration with existing pipelines easier.\n\nIn addition, if anyone finds more tomograms to annotate, please let me know!",
      "votes": 7,
      "replies": [
        {
          "id": 3173177,
          "postDate": "2025-04-07T16:06:34.683Z",
          "content": "<p>Where are the JPGs? I thought with the original dataset, but I see only NPYs there.</p>",
          "rawMarkdown": "Where are the JPGs? I thought with the original dataset, but I see only NPYs there.",
          "votes": 1,
          "replies": [
            {
              "id": 3173216,
              "postDate": "2025-04-07T16:51:43.367Z",
              "content": "<p> </p>\n<p><a href=\"https://www.kaggle.com/tilii7\" target=\"_blank\">@tilii7</a>, done.</p>",
              "rawMarkdown": "~~The upload is processing, I will update this comment when it finishes.~~ \n\n@tilii7, done.",
              "votes": 4
            },
            {
              "id": 3173233,
              "postDate": "2025-04-07T17:13:48.567Z",
              "content": "<p>Great, thank you for providing these files.</p>",
              "rawMarkdown": "Great, thank you for providing these files.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3176421,
      "postDate": "2025-04-11T10:24:59.140Z",
      "content": "<p>thank  you very much for sharing.  </p>\n<p>Anyone could share the PARSE DATA of  combined competition dataset and extradata?  <br>\nSince most of us don't have enough resources to process big amount of data.</p>",
      "rawMarkdown": "thank  you very much for sharing.  \n\nAnyone could share the PARSE DATA of  combined competition dataset and extradata?  \nSince most of us don't have enough resources to process big amount of data.",
      "votes": 4,
      "replies": [
        {
          "id": 3193532,
          "postDate": "2025-05-04T14:22:48.080Z",
          "content": "<p>were you able to \"parse data\" the combined set? if so can you share it?</p>",
          "rawMarkdown": "were you able to \"parse data\" the combined set? if so can you share it?"
        }
      ]
    },
    {
      "id": 3214190,
      "postDate": "2025-05-31T04:57:02.817Z",
      "content": "<p>Thanks for sharing, this work is very helpful for me</p>",
      "rawMarkdown": "Thanks for sharing, this work is very helpful for me",
      "votes": 1
    },
    {
      "id": 3159398,
      "postDate": "2025-03-25T14:52:59.393Z",
      "content": "<p><a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a>, Thanks a lot. Did you manually annotate them all after preprocessing or some of them have annotations already and you rescale it for new size?</p>",
      "rawMarkdown": "@brendanartley, Thanks a lot. Did you manually annotate them all after preprocessing or some of them have annotations already and you rescale it for new size?",
      "votes": 3,
      "replies": [
        {
          "id": 3159453,
          "postDate": "2025-03-25T15:58:00.773Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/fautei\" target=\"_blank\">@fautei</a>, the annotations were manually annotated after preprocessing. If you see any errors, please let me know!</p>",
          "rawMarkdown": "Hi @fautei, the annotations were manually annotated after preprocessing. If you see any errors, please let me know!",
          "votes": 3,
          "replies": [
            {
              "id": 3159551,
              "postDate": "2025-03-25T18:06:59.630Z",
              "content": "<p>I share my notebook that create extra data in competition dataset format using your labels <a href=\"https://www.kaggle.com/code/fautei/extra-data-from-brendanartley-in-competition-data\" target=\"_blank\">https://www.kaggle.com/code/fautei/extra-data-from-brendanartley-in-competition-data</a></p>",
              "rawMarkdown": "I share my notebook that create extra data in competition dataset format using your labels https://www.kaggle.com/code/fautei/extra-data-from-brendanartley-in-competition-data",
              "votes": 17
            }
          ]
        }
      ]
    },
    {
      "id": 3158869,
      "postDate": "2025-03-25T02:27:16.880Z",
      "content": "<p>Thank you for providing this dataset.</p>\n<p>To my eye, some of the dots are close to or at the internal membrane, meaning they are \"lower\" than the location of flagella in the competition dataset. It probably works when a large enough bounding box is used.</p>",
      "rawMarkdown": "Thank you for providing this dataset.\n\nTo my eye, some of the dots are close to or at the internal membrane, meaning they are \"lower\" than the location of flagella in the competition dataset. It probably works when a large enough bounding box is used.",
      "votes": 3,
      "replies": [
        {
          "id": 3158874,
          "postDate": "2025-03-25T02:35:20.997Z",
          "content": "<p>Thanks for the feedback <a href=\"https://www.kaggle.com/tilii7\" target=\"_blank\">@tilii7</a>. Hopefully they are close enough for most pipelines!</p>",
          "rawMarkdown": "Thanks for the feedback @tilii7. Hopefully they are close enough for most pipelines!",
          "votes": 1
        }
      ]
    },
    {
      "id": 3186200,
      "postDate": "2025-04-24T10:41:28.023Z",
      "content": "<p><a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> Hi. Thank you for sharing this nice dataset.</p>\n<p>I found tiny amount of irregular pixel data in this dataset (~4.4%).<br>\nIt seems because of normalizing pixel value in <code>uint8</code> datatype. <br>\nMaybe normalizing with float32 data type will fix this issue.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F50f64b4e583486c15e68b96d01ccb45c%2Fvisualization_002.jpeg?generation=1745491086858616&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F79920aa53aa4d108e5404632bd3c5e5d%2Fvisualization.jpeg?generation=1745491101892394&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F6664d640b06875ef787cc2fc8f355807%2Fpixel_dist.jpeg?generation=1745491115831111&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F264812db967341b3b3918a18e38dd5aa%2Fhistogram_irregular.jpeg?generation=1745491137864416&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "@brendanartley Hi. Thank you for sharing this nice dataset.\n\nI found tiny amount of irregular pixel data in this dataset (~4.4%).\nIt seems because of normalizing pixel value in `uint8` datatype. \nMaybe normalizing with float32 data type will fix this issue.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F50f64b4e583486c15e68b96d01ccb45c%2Fvisualization_002.jpeg?generation=1745491086858616&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F79920aa53aa4d108e5404632bd3c5e5d%2Fvisualization.jpeg?generation=1745491101892394&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F6664d640b06875ef787cc2fc8f355807%2Fpixel_dist.jpeg?generation=1745491115831111&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F264812db967341b3b3918a18e38dd5aa%2Fhistogram_irregular.jpeg?generation=1745491137864416&alt=media)",
      "votes": 1,
      "replies": [
        {
          "id": 3186227,
          "postDate": "2025-04-24T11:15:28.923Z",
          "content": "<p>I checked the originals and they are just super noisy, also some of the samples are inverted.</p>",
          "rawMarkdown": "I checked the originals and they are just super noisy, also some of the samples are inverted.",
          "votes": 4
        },
        {
          "id": 3207700,
          "postDate": "2025-05-23T06:33:30.113Z",
          "content": "<p>Very useful work! Thanks!</p>",
          "rawMarkdown": "Very useful work! Thanks!",
          "votes": 2
        }
      ]
    },
    {
      "id": 3191797,
      "postDate": "2025-05-02T06:43:12.740Z",
      "content": "<p>Hi BARTLEY,</p>\n<p>Thank you so much for sharing your data and work!</p>\n<p>I have integrated the extra data into my training. My strategy was to add your data to my training set while keeping the validation data unchanged (the same as before, when I was only using the competition data).</p>\n<p>However, I ran into an issue: with exactly the same settings, I achieved a <strong>0.612</strong> score using only the competition data, but after integrating your data, the score dropped to <strong>0.555</strong>. In other words, adding more training data actually degraded the LB score.</p>\n<p>Below, I’ve attached the training/validation loss curves for (1) using only competition data and (2) using the integrated dataset.<br>\nDo you have any insights on how to interpret this behavior? And do you have suggestions on how to use external data properly?<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2580499%2F43242816d5585f0661b6e7080ee21142%2Fdfl_loss_curve%20(1).png?generation=1746168033304680&amp;alt=media\" alt=\"only competition data\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2580499%2F5e350cb277b55aaeb67835d1c31463d1%2Fdfl_loss_curve.png?generation=1746168058709572&amp;alt=media\" alt=\"integrated extra data into trainining\"></p>\n<p>Best <br>\nLeo</p>",
      "rawMarkdown": "Hi BARTLEY,\n\nThank you so much for sharing your data and work!\n\nI have integrated the extra data into my training. My strategy was to add your data to my training set while keeping the validation data unchanged (the same as before, when I was only using the competition data).\n\nHowever, I ran into an issue: with exactly the same settings, I achieved a **0.612** score using only the competition data, but after integrating your data, the score dropped to **0.555**. In other words, adding more training data actually degraded the LB score.\n\nBelow, I’ve attached the training/validation loss curves for (1) using only competition data and (2) using the integrated dataset.\nDo you have any insights on how to interpret this behavior? And do you have suggestions on how to use external data properly?\n![only competition data](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2580499%2F43242816d5585f0661b6e7080ee21142%2Fdfl_loss_curve%20(1).png?generation=1746168033304680&alt=media)\n![integrated extra data into trainining](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2580499%2F5e350cb277b55aaeb67835d1c31463d1%2Fdfl_loss_curve.png?generation=1746168058709572&alt=media)\n\nBest \nLeo\n\n",
      "votes": 2,
      "replies": [
        {
          "id": 3192128,
          "postDate": "2025-05-02T13:36:43.037Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/yxyyxy\" target=\"_blank\">@yxyyxy</a>, thanks for the comment.</p>\n<p>Have you preprocessed each dataset in the same way? </p>\n<p>For example, this dataset resizes all tomograms to <code>(128, 512, 512)</code> which distorts the aspect ratio + voxel spacing - unlike the competition data.</p>",
          "rawMarkdown": "Hi @yxyyxy, thanks for the comment.\n\nHave you preprocessed each dataset in the same way? \n\nFor example, this dataset resizes all tomograms to `(128, 512, 512)` which distorts the aspect ratio + voxel spacing - unlike the competition data.",
          "votes": 1,
          "replies": [
            {
              "id": 3192527,
              "postDate": "2025-05-03T02:02:20.703Z",
              "content": "<p>Hi Bartley</p>\n<p>Thx for replying! I do preprocess like this:<br>\n1) normalize the image values<br>\n2) prepare  bounding box annotation by dividing the \"z,x,y\" in your csv label by image size.<br>\nIn my code, the aspect ratio and voxel spacing are <strong>not needed</strong>. So how should I use them ? And also I think these info are not avaliable in your database.</p>\n<p>Below is my core code for preprocessing. </p>\n<pre><code> ():\n    motor_counts = []\n     tomo_id  tomogram_ids:\n        \n        tomo_motors = labels_df[labels_df[] == tomo_id]\n         _, motor  tomo_motors.iterrows():\n             pd.isna(motor[]):\n                \n            motor_counts.append(\n                (tomo_id, \n                 (motor[]), \n                 (motor[]), \n                 (motor[]))\n            )\n\n    ()\n    processed_slices = \n\n    \n    z_max = \n     tomo_id, z_center, y_center, x_center  tqdm(motor_counts, desc=):\n        z_min = (, z_center - trust)\n        z_max_bound = (z_max - , z_center + trust)\n         z  (z_min, z_max_bound + ):\n            \n            slice_filename = \n            src_path = os.path.join(train_dir, tomo_id, slice_filename)\n              os.path.exists(src_path):\n                ()\n                \n\n            \n            img = Image.(src_path)\n            img_array = np.array(img)\n            normalized_img = normalize_slice(img_array)\n            dest_filename = \n            dest_path = os.path.join(images_dir, dest_filename)\n            Image.fromarray(normalized_img).save(dest_path)\n\n            \n            img_width, img_height = img.size\n            x_center_norm = x_center / img_width\n            y_center_norm = y_center / img_height\n            box_width_norm = BOX_SIZE / img_width\n            box_height_norm = BOX_SIZE / img_height\n            label_path = os.path.join(labels_dir, dest_filename.replace(, ))\n             (label_path, )  f:\n                f.write()\n\n            processed_slices += \n\n     processed_slices, (motor_counts)\n</code></pre>",
              "rawMarkdown": "Hi Bartley\n\nThx for replying! I do preprocess like this:\n1) normalize the image values\n2) prepare  bounding box annotation by dividing the \"z,x,y\" in your csv label by image size.\nIn my code, the aspect ratio and voxel spacing are **not needed**. So how should I use them ? And also I think these info are not avaliable in your database.\n\n\nBelow is my core code for preprocessing. \n```python\ndef process_tomogram_set_extra(tomogram_ids, images_dir, labels_dir, train_dir, labels_df, set_name):\n    motor_counts = []\n    for tomo_id in tomogram_ids:\n        # Get motor annotations for the current tomogram\n        tomo_motors = labels_df[labels_df['tomo_id'] == tomo_id]\n        for _, motor in tomo_motors.iterrows():\n            if pd.isna(motor['z']):\n                continue\n            motor_counts.append(\n                (tomo_id, \n                 int(motor['z']), \n                 int(motor['y']), \n                 int(motor['x']))\n            )\n    \n    print(f\"Will process approximately {len(motor_counts) * (2 * trust + 1)} slices for {set_name}\")\n    processed_slices = 0\n    \n    # Loop over each motor annotation\n    z_max = 128\n    for tomo_id, z_center, y_center, x_center in tqdm(motor_counts, desc=f\"Processing {set_name} motors\"):\n        z_min = max(0, z_center - trust)\n        z_max_bound = min(z_max - 1, z_center + trust)\n        for z in range(z_min, z_max_bound + 1):\n            # Create the slice filename and source path\n            slice_filename = f\"slice_{z:04d}.jpg\"\n            src_path = os.path.join(train_dir, tomo_id, slice_filename)\n            if not os.path.exists(src_path):\n                print(f\"Warning: {src_path} does not exist, skipping.\")\n                continue\n            \n            # Load, normalize, and save the image slice\n            img = Image.open(src_path)\n            img_array = np.array(img)\n            normalized_img = normalize_slice(img_array)\n            dest_filename = f\"{tomo_id}_z{z:04d}_y{y_center:04d}_x{x_center:04d}.jpg\"\n            dest_path = os.path.join(images_dir, dest_filename)\n            Image.fromarray(normalized_img).save(dest_path)\n            \n            # Prepare YOLO bounding box annotation (normalized values)\n            img_width, img_height = img.size\n            x_center_norm = x_center / img_width\n            y_center_norm = y_center / img_height\n            box_width_norm = BOX_SIZE / img_width\n            box_height_norm = BOX_SIZE / img_height\n            label_path = os.path.join(labels_dir, dest_filename.replace('.jpg', '.txt'))\n            with open(label_path, 'w') as f:\n                f.write(f\"0 {x_center_norm} {y_center_norm} {box_width_norm} {box_height_norm}\\n\")\n\n            processed_slices += 1\n    \n    return processed_slices, len(motor_counts)\n```",
              "votes": -1
            },
            {
              "id": 3192990,
              "postDate": "2025-05-03T16:33:11.423Z",
              "content": "<p>Hello, I highly believe the cause is that you use the same trust while in the extraction the 128 depth is not the same as in the competition data</p>\n<p>it is extracted with an adaptive step, example for a volume of depth 800 the step is 800 div 128 = 8 so with a trust of 4 it would be like choosing a trust of 4*8 in the original competition data</p>",
              "rawMarkdown": "Hello, I highly believe the cause is that you use the same trust while in the extraction the 128 depth is not the same as in the competition data\n\nit is extracted with an adaptive step, example for a volume of depth 800 the step is 800 div 128 = 8 so with a trust of 4 it would be like choosing a trust of 4*8 in the original competition data",
              "votes": 4
            },
            {
              "id": 3194190,
              "postDate": "2025-05-05T14:09:10.097Z",
              "content": "<p>Hi Youssef Ouertani,</p>\n<p>Thank you so much !<br>\nI changed my trust for extra data from 4 to 2. The LB score improved from 0.555 to 0.619, with other settings unchanged. Note that without the extra training data, the LB score was 0.612. Finally in my case, I make use of extra data and get a little improvement.</p>\n<p>Leo</p>",
              "rawMarkdown": "Hi Youssef Ouertani,\n\nThank you so much !\nI changed my trust for extra data from 4 to 2. The LB score improved from 0.555 to 0.619, with other settings unchanged. Note that without the extra training data, the LB score was 0.612. Finally in my case, I make use of extra data and get a little improvement.\n\nLeo",
              "votes": 3
            }
          ]
        }
      ]
    },
    {
      "id": 3168412,
      "postDate": "2025-04-02T12:01:27.733Z",
      "content": "<p>感谢你所有的工作和分享</p>",
      "rawMarkdown": "感谢你所有的工作和分享",
      "votes": 1
    },
    {
      "id": 3180527,
      "postDate": "2025-04-16T18:27:35.600Z",
      "content": "<p>Thank you for sharing extra data! I am trying to add partial or all data on top of original data but the performance seems worse. Can anyone share some tips or tricks how to leverage this data properly? Thank you. </p>",
      "rawMarkdown": "Thank you for sharing extra data! I am trying to add partial or all data on top of original data but the performance seems worse. Can anyone share some tips or tricks how to leverage this data properly? Thank you. ",
      "votes": 2,
      "replies": [
        {
          "id": 3181314,
          "postDate": "2025-04-17T17:57:31.953Z",
          "content": "<p>i had the same effect</p>",
          "rawMarkdown": "i had the same effect"
        }
      ]
    },
    {
      "id": 3159197,
      "postDate": "2025-03-25T11:32:23.417Z",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a>, thank you so much for sharing!</p>\n<p>I have a question regarding the preprocessing steps applied, especially resizing. Did you do sub tomogram averaging for reducing the z dimension to 128?</p>\n<p>EDIT: Just looked at the code and noticed you include everything there, ignore the question! Thanks!</p>",
      "rawMarkdown": "Hello @brendanartley, thank you so much for sharing!\n\nI have a question regarding the preprocessing steps applied, especially resizing. Did you do sub tomogram averaging for reducing the z dimension to 128?\n\nEDIT: Just looked at the code and noticed you include everything there, ignore the question! Thanks!",
      "votes": 2,
      "replies": [
        {
          "id": 3161063,
          "postDate": "2025-03-27T13:21:35Z",
          "content": "<p><a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> if that does not constitute too much, could you detail the way you've labelled these? Was it manual annotating? How long did this take you?</p>",
          "rawMarkdown": "@brendanartley if that does not constitute too much, could you detail the way you've labelled these? Was it manual annotating? How long did this take you?",
          "votes": 1,
          "replies": [
            {
              "id": 3161105,
              "postDate": "2025-03-27T14:16:09.663Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/andreizamfir\" target=\"_blank\">@andreizamfir</a>, thanks for the comment. The points were manually annotated using <a href=\"https://napari.org/dev/index.html\" target=\"_blank\">Napari</a>. As for how long it took, I am not sure.</p>\n<p>To assist in labelling, I first trained a model on the competition data to make predictions. Then, I overlaid  the predictions on the volumes. This helped to guide annotation for ~10% of the images, but the rest were different enough that the predictions were not useful. </p>\n<p>Here is a sample of how that looks below (the arrow was drawn after haha).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F4ab78479ddb9f91f96b6fb51593481b5%2FCapture.JPG?generation=1743087591590431&amp;alt=media\" alt=\"\"></p>",
              "rawMarkdown": "Hi @andreizamfir, thanks for the comment. The points were manually annotated using [Napari](https://napari.org/dev/index.html). As for how long it took, I am not sure.\n\nTo assist in labelling, I first trained a model on the competition data to make predictions. Then, I overlaid  the predictions on the volumes. This helped to guide annotation for ~10% of the images, but the rest were different enough that the predictions were not useful. \n\nHere is a sample of how that looks below (the arrow was drawn after haha).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F4ab78479ddb9f91f96b6fb51593481b5%2FCapture.JPG?generation=1743087591590431&alt=media)",
              "votes": 8
            }
          ]
        }
      ]
    },
    {
      "id": 3158878,
      "postDate": "2025-03-25T02:44:06.170Z",
      "content": "<p>Appreciate all your work and sharing, this is magnificent!</p>",
      "rawMarkdown": "Appreciate all your work and sharing, this is magnificent!",
      "votes": 2
    },
    {
      "id": 3158844,
      "postDate": "2025-03-25T01:18:29.810Z",
      "content": "<p>if you don't mind sharing to how much extent did extra data effect your lb socre</p>",
      "rawMarkdown": "if you don't mind sharing to how much extent did extra data effect your lb socre",
      "votes": 2,
      "replies": [
        {
          "id": 3158870,
          "postDate": "2025-03-25T02:31:52.127Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/tahaalshatiri\" target=\"_blank\">@tahaalshatiri</a>. </p>\n<p>Improves LB by ~0.07.</p>",
          "rawMarkdown": "Hi @tahaalshatiri. ~~I have not trained my pipeline on the complete dataset yet, but with ~800 of them it LB improves by about 0.05.~~\n\nImproves LB by ~0.07.",
          "votes": 7,
          "replies": [
            {
              "id": 3188090,
              "postDate": "2025-04-27T03:33:14.200Z",
              "content": "<p>Thanks for your insights! This seems to help reduce false positives while maintaining good recall on motors.</p>",
              "rawMarkdown": "Thanks for your insights! This seems to help reduce false positives while maintaining good recall on motors."
            }
          ]
        }
      ]
    },
    {
      "id": 3161260,
      "postDate": "2025-03-27T17:41:30.620Z",
      "content": "<p>you have done nice work</p>",
      "rawMarkdown": "you have done nice work\n"
    },
    {
      "id": 3160472,
      "postDate": "2025-03-26T20:06:14.233Z",
      "content": "<p>wow! great job!</p>",
      "rawMarkdown": "wow! great job!",
      "votes": 1
    },
    {
      "id": 3236336,
      "postDate": "2025-06-30T09:04:17.660Z",
      "content": "<p>Thank you. I’ll use this as a reference.</p>",
      "rawMarkdown": "Thank you. I’ll use this as a reference."
    },
    {
      "id": 3183082,
      "postDate": "2025-04-20T11:12:35.847Z",
      "content": "<p>thanks a lot of the extra data [im just in my early steps on that field btw]😄, so i'm just wondering if your rate in the Leaderboard  is because you do it manually or you have a strong model with that Fb_score? </p>",
      "rawMarkdown": "thanks a lot of the extra data [im just in my early steps on that field btw]😄, so i'm just wondering if your rate in the Leaderboard  is because you do it manually or you have a strong model with that Fb_score? "
    },
    {
      "id": 3160894,
      "postDate": "2025-03-27T08:58:54.950Z",
      "rawMarkdown": "",
      "votes": 10,
      "isDeleted": true,
      "replies": [
        {
          "id": 3207699,
          "postDate": "2025-05-23T06:32:54.483Z",
          "content": "<p>Hahaha😄 that's for sure</p>",
          "rawMarkdown": "Hahaha😄 that's for sure",
          "votes": 3
        }
      ]
    },
    {
      "id": 3160892,
      "postDate": "2025-03-27T08:57:42.710Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 3218027,
      "postDate": "2025-06-05T17:11:01.120Z",
      "content": "<p>thanks for sharing….</p>",
      "rawMarkdown": "thanks for sharing....\n",
      "votes": 1
    },
    {
      "id": 3216164,
      "postDate": "2025-06-03T07:30:14.563Z",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing",
      "votes": 1
    },
    {
      "id": 3159042,
      "postDate": "2025-03-25T07:38:53.787Z",
      "content": "<p>Thanks a lot for sharing this!</p>",
      "rawMarkdown": "Thanks a lot for sharing this!",
      "votes": 3
    },
    {
      "id": 3191748,
      "postDate": "2025-05-02T05:28:28.370Z",
      "content": "<p>thanks you very much for your help </p>",
      "rawMarkdown": "thanks you very much for your help ",
      "votes": 1
    },
    {
      "id": 3178198,
      "postDate": "2025-04-13T20:36:55.533Z",
      "content": "<p>It was very helpful, thank you.</p>",
      "rawMarkdown": "It was very helpful, thank you.",
      "votes": 1
    },
    {
      "id": 3176753,
      "postDate": "2025-04-11T17:17:32.587Z",
      "content": "<p>Thank you very much for sharing </p>",
      "rawMarkdown": "Thank you very much for sharing ",
      "votes": 1
    },
    {
      "id": 3185587,
      "postDate": "2025-04-23T14:28:11.760Z",
      "content": "<p>thank you very much.</p>",
      "rawMarkdown": "thank you very much.",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3160527,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-03-26T22:05:33.307000",
      "content": "<p>Appreciate all your work and sharing</p>",
      "votes": 11,
      "replies": []
    },
    {
      "id": 3173157,
      "author_name": "Bartley",
      "author_url": "",
      "post_date": "2025-04-07T15:42:34.973000",
      "content": "<p><strong>Update:</strong></p>\n<p>After some feedback, I uploaded a <code>jpgs</code> folder to this dataset to match the competition data format. This should make integration with existing pipelines easier.</p>\n<p>In addition, if anyone finds more tomograms to annotate, please let me know!</p>",
      "votes": 7,
      "replies": [
        {
          "id": 3173177,
          "author_name": "Tilii",
          "author_url": "",
          "post_date": "2025-04-07T16:06:34.683000",
          "content": "<p>Where are the JPGs? I thought with the original dataset, but I see only NPYs there.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3173216,
              "author_name": "Bartley",
              "author_url": "",
              "post_date": "2025-04-07T16:51:43.367000",
              "content": "<p> </p>\n<p><a href=\"https://www.kaggle.com/tilii7\" target=\"_blank\">@tilii7</a>, done.</p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 3173233,
              "author_name": "Tilii",
              "author_url": "",
              "post_date": "2025-04-07T17:13:48.567000",
              "content": "<p>Great, thank you for providing these files.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3176421,
      "author_name": "dragon zhang",
      "author_url": "",
      "post_date": "2025-04-11T10:24:59.140000",
      "content": "<p>thank  you very much for sharing.  </p>\n<p>Anyone could share the PARSE DATA of  combined competition dataset and extradata?  <br>\nSince most of us don't have enough resources to process big amount of data.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 3193532,
          "author_name": "Sushlok Ajay Shah",
          "author_url": "",
          "post_date": "2025-05-04T14:22:48.080000",
          "content": "<p>were you able to \"parse data\" the combined set? if so can you share it?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3214190,
      "author_name": "youlemei",
      "author_url": "",
      "post_date": "2025-05-31T04:57:02.817000",
      "content": "<p>Thanks for sharing, this work is very helpful for me</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3159398,
      "author_name": "Maxim Ilyin",
      "author_url": "",
      "post_date": "2025-03-25T14:52:59.393000",
      "content": "<p><a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a>, Thanks a lot. Did you manually annotate them all after preprocessing or some of them have annotations already and you rescale it for new size?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 3159453,
          "author_name": "Bartley",
          "author_url": "",
          "post_date": "2025-03-25T15:58:00.773000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/fautei\" target=\"_blank\">@fautei</a>, the annotations were manually annotated after preprocessing. If you see any errors, please let me know!</p>",
          "votes": 3,
          "replies": [
            {
              "id": 3159551,
              "author_name": "Maxim Ilyin",
              "author_url": "",
              "post_date": "2025-03-25T18:06:59.630000",
              "content": "<p>I share my notebook that create extra data in competition dataset format using your labels <a href=\"https://www.kaggle.com/code/fautei/extra-data-from-brendanartley-in-competition-data\" target=\"_blank\">https://www.kaggle.com/code/fautei/extra-data-from-brendanartley-in-competition-data</a></p>",
              "votes": 17,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3158869,
      "author_name": "Tilii",
      "author_url": "",
      "post_date": "2025-03-25T02:27:16.880000",
      "content": "<p>Thank you for providing this dataset.</p>\n<p>To my eye, some of the dots are close to or at the internal membrane, meaning they are \"lower\" than the location of flagella in the competition dataset. It probably works when a large enough bounding box is used.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 3158874,
          "author_name": "Bartley",
          "author_url": "",
          "post_date": "2025-03-25T02:35:20.997000",
          "content": "<p>Thanks for the feedback <a href=\"https://www.kaggle.com/tilii7\" target=\"_blank\">@tilii7</a>. Hopefully they are close enough for most pipelines!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3186200,
      "author_name": "Bilzard",
      "author_url": "",
      "post_date": "2025-04-24T10:41:28.023000",
      "content": "<p><a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> Hi. Thank you for sharing this nice dataset.</p>\n<p>I found tiny amount of irregular pixel data in this dataset (~4.4%).<br>\nIt seems because of normalizing pixel value in <code>uint8</code> datatype. <br>\nMaybe normalizing with float32 data type will fix this issue.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F50f64b4e583486c15e68b96d01ccb45c%2Fvisualization_002.jpeg?generation=1745491086858616&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F79920aa53aa4d108e5404632bd3c5e5d%2Fvisualization.jpeg?generation=1745491101892394&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F6664d640b06875ef787cc2fc8f355807%2Fpixel_dist.jpeg?generation=1745491115831111&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F264812db967341b3b3918a18e38dd5aa%2Fhistogram_irregular.jpeg?generation=1745491137864416&amp;alt=media\" alt=\"\"></p>",
      "votes": 1,
      "replies": [
        {
          "id": 3186227,
          "author_name": "DennisSakva",
          "author_url": "",
          "post_date": "2025-04-24T11:15:28.923000",
          "content": "<p>I checked the originals and they are just super noisy, also some of the samples are inverted.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 3207700,
          "author_name": "",
          "author_url": "",
          "post_date": "2025-05-23T06:33:30.113000",
          "content": "<p>Very useful work! Thanks!</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 3191797,
      "author_name": "Leo Yang",
      "author_url": "",
      "post_date": "2025-05-02T06:43:12.740000",
      "content": "<p>Hi BARTLEY,</p>\n<p>Thank you so much for sharing your data and work!</p>\n<p>I have integrated the extra data into my training. My strategy was to add your data to my training set while keeping the validation data unchanged (the same as before, when I was only using the competition data).</p>\n<p>However, I ran into an issue: with exactly the same settings, I achieved a <strong>0.612</strong> score using only the competition data, but after integrating your data, the score dropped to <strong>0.555</strong>. In other words, adding more training data actually degraded the LB score.</p>\n<p>Below, I’ve attached the training/validation loss curves for (1) using only competition data and (2) using the integrated dataset.<br>\nDo you have any insights on how to interpret this behavior? And do you have suggestions on how to use external data properly?<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2580499%2F43242816d5585f0661b6e7080ee21142%2Fdfl_loss_curve%20(1).png?generation=1746168033304680&amp;alt=media\" alt=\"only competition data\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2580499%2F5e350cb277b55aaeb67835d1c31463d1%2Fdfl_loss_curve.png?generation=1746168058709572&amp;alt=media\" alt=\"integrated extra data into trainining\"></p>\n<p>Best <br>\nLeo</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3192128,
          "author_name": "Bartley",
          "author_url": "",
          "post_date": "2025-05-02T13:36:43.037000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/yxyyxy\" target=\"_blank\">@yxyyxy</a>, thanks for the comment.</p>\n<p>Have you preprocessed each dataset in the same way? </p>\n<p>For example, this dataset resizes all tomograms to <code>(128, 512, 512)</code> which distorts the aspect ratio + voxel spacing - unlike the competition data.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3192527,
              "author_name": "Leo Yang",
              "author_url": "",
              "post_date": "2025-05-03T02:02:20.703000",
              "content": "<p>Hi Bartley</p>\n<p>Thx for replying! I do preprocess like this:<br>\n1) normalize the image values<br>\n2) prepare  bounding box annotation by dividing the \"z,x,y\" in your csv label by image size.<br>\nIn my code, the aspect ratio and voxel spacing are <strong>not needed</strong>. So how should I use them ? And also I think these info are not avaliable in your database.</p>\n<p>Below is my core code for preprocessing. </p>\n<pre><code> ():\n    motor_counts = []\n     tomo_id  tomogram_ids:\n        \n        tomo_motors = labels_df[labels_df[] == tomo_id]\n         _, motor  tomo_motors.iterrows():\n             pd.isna(motor[]):\n                \n            motor_counts.append(\n                (tomo_id, \n                 (motor[]), \n                 (motor[]), \n                 (motor[]))\n            )\n\n    ()\n    processed_slices = \n\n    \n    z_max = \n     tomo_id, z_center, y_center, x_center  tqdm(motor_counts, desc=):\n        z_min = (, z_center - trust)\n        z_max_bound = (z_max - , z_center + trust)\n         z  (z_min, z_max_bound + ):\n            \n            slice_filename = \n            src_path = os.path.join(train_dir, tomo_id, slice_filename)\n              os.path.exists(src_path):\n                ()\n                \n\n            \n            img = Image.(src_path)\n            img_array = np.array(img)\n            normalized_img = normalize_slice(img_array)\n            dest_filename = \n            dest_path = os.path.join(images_dir, dest_filename)\n            Image.fromarray(normalized_img).save(dest_path)\n\n            \n            img_width, img_height = img.size\n            x_center_norm = x_center / img_width\n            y_center_norm = y_center / img_height\n            box_width_norm = BOX_SIZE / img_width\n            box_height_norm = BOX_SIZE / img_height\n            label_path = os.path.join(labels_dir, dest_filename.replace(, ))\n             (label_path, )  f:\n                f.write()\n\n            processed_slices += \n\n     processed_slices, (motor_counts)\n</code></pre>",
              "votes": -1,
              "replies": []
            },
            {
              "id": 3192990,
              "author_name": "Youssef Ouertani",
              "author_url": "",
              "post_date": "2025-05-03T16:33:11.423000",
              "content": "<p>Hello, I highly believe the cause is that you use the same trust while in the extraction the 128 depth is not the same as in the competition data</p>\n<p>it is extracted with an adaptive step, example for a volume of depth 800 the step is 800 div 128 = 8 so with a trust of 4 it would be like choosing a trust of 4*8 in the original competition data</p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 3194190,
              "author_name": "Leo Yang",
              "author_url": "",
              "post_date": "2025-05-05T14:09:10.097000",
              "content": "<p>Hi Youssef Ouertani,</p>\n<p>Thank you so much !<br>\nI changed my trust for extra data from 4 to 2. The LB score improved from 0.555 to 0.619, with other settings unchanged. Note that without the extra training data, the LB score was 0.612. Finally in my case, I make use of extra data and get a little improvement.</p>\n<p>Leo</p>",
              "votes": 3,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3168412,
      "author_name": "Hide on bush",
      "author_url": "",
      "post_date": "2025-04-02T12:01:27.733000",
      "content": "<p>感谢你所有的工作和分享</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3180527,
      "author_name": "HongweiLuan",
      "author_url": "",
      "post_date": "2025-04-16T18:27:35.600000",
      "content": "<p>Thank you for sharing extra data! I am trying to add partial or all data on top of original data but the performance seems worse. Can anyone share some tips or tricks how to leverage this data properly? Thank you. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 3181314,
          "author_name": "Vasilis",
          "author_url": "",
          "post_date": "2025-04-17T17:57:31.953000",
          "content": "<p>i had the same effect</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3159197,
      "author_name": "Andrei Zamfir",
      "author_url": "",
      "post_date": "2025-03-25T11:32:23.417000",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a>, thank you so much for sharing!</p>\n<p>I have a question regarding the preprocessing steps applied, especially resizing. Did you do sub tomogram averaging for reducing the z dimension to 128?</p>\n<p>EDIT: Just looked at the code and noticed you include everything there, ignore the question! Thanks!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3161063,
          "author_name": "Andrei Zamfir",
          "author_url": "",
          "post_date": "2025-03-27T13:21:35",
          "content": "<p><a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> if that does not constitute too much, could you detail the way you've labelled these? Was it manual annotating? How long did this take you?</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3161105,
              "author_name": "Bartley",
              "author_url": "",
              "post_date": "2025-03-27T14:16:09.663000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/andreizamfir\" target=\"_blank\">@andreizamfir</a>, thanks for the comment. The points were manually annotated using <a href=\"https://napari.org/dev/index.html\" target=\"_blank\">Napari</a>. As for how long it took, I am not sure.</p>\n<p>To assist in labelling, I first trained a model on the competition data to make predictions. Then, I overlaid  the predictions on the volumes. This helped to guide annotation for ~10% of the images, but the rest were different enough that the predictions were not useful. </p>\n<p>Here is a sample of how that looks below (the arrow was drawn after haha).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F4ab78479ddb9f91f96b6fb51593481b5%2FCapture.JPG?generation=1743087591590431&amp;alt=media\" alt=\"\"></p>",
              "votes": 8,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3158878,
      "author_name": "Yksin Young",
      "author_url": "",
      "post_date": "2025-03-25T02:44:06.170000",
      "content": "<p>Appreciate all your work and sharing, this is magnificent!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3158844,
      "author_name": "Taha_Alshatiri",
      "author_url": "",
      "post_date": "2025-03-25T01:18:29.810000",
      "content": "<p>if you don't mind sharing to how much extent did extra data effect your lb socre</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3158870,
          "author_name": "Bartley",
          "author_url": "",
          "post_date": "2025-03-25T02:31:52.127000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/tahaalshatiri\" target=\"_blank\">@tahaalshatiri</a>. </p>\n<p>Improves LB by ~0.07.</p>",
          "votes": 7,
          "replies": [
            {
              "id": 3188090,
              "author_name": "Fae Gaze",
              "author_url": "",
              "post_date": "2025-04-27T03:33:14.200000",
              "content": "<p>Thanks for your insights! This seems to help reduce false positives while maintaining good recall on motors.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3161260,
      "author_name": "Fahim Sarker",
      "author_url": "",
      "post_date": "2025-03-27T17:41:30.620000",
      "content": "<p>you have done nice work</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3160472,
      "author_name": "Shohinur Pervez Shohan",
      "author_url": "",
      "post_date": "2025-03-26T20:06:14.233000",
      "content": "<p>wow! great job!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3236336,
      "author_name": "tngmb08",
      "author_url": "",
      "post_date": "2025-06-30T09:04:17.660000",
      "content": "<p>Thank you. I’ll use this as a reference.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3183082,
      "author_name": "ZAID",
      "author_url": "",
      "post_date": "2025-04-20T11:12:35.847000",
      "content": "<p>thanks a lot of the extra data [im just in my early steps on that field btw]😄, so i'm just wondering if your rate in the Leaderboard  is because you do it manually or you have a strong model with that Fb_score? </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3160894,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-03-27T08:58:54.950000",
      "content": "",
      "votes": 10,
      "replies": [
        {
          "id": 3207699,
          "author_name": "",
          "author_url": "",
          "post_date": "2025-05-23T06:32:54.483000",
          "content": "<p>Hahaha😄 that's for sure</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 3160892,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-03-27T08:57:42.710000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3218027,
      "author_name": "Sarah Arshad",
      "author_url": "",
      "post_date": "2025-06-05T17:11:01.120000",
      "content": "<p>thanks for sharing….</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3216164,
      "author_name": "yunhehebupt",
      "author_url": "",
      "post_date": "2025-06-03T07:30:14.563000",
      "content": "<p>Thanks for sharing</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3159042,
      "author_name": "Jeroen Cottaar",
      "author_url": "",
      "post_date": "2025-03-25T07:38:53.787000",
      "content": "<p>Thanks a lot for sharing this!</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 3191748,
      "author_name": "Abdul Wasay",
      "author_url": "",
      "post_date": "2025-05-02T05:28:28.370000",
      "content": "<p>thanks you very much for your help </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3178198,
      "author_name": "Ali Kabirzadeh",
      "author_url": "",
      "post_date": "2025-04-13T20:36:55.533000",
      "content": "<p>It was very helpful, thank you.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3176753,
      "author_name": "MEHUL MRIDUL",
      "author_url": "",
      "post_date": "2025-04-11T17:17:32.587000",
      "content": "<p>Thank you very much for sharing </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3185587,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-04-23T14:28:11.760000",
      "content": "<p>thank you very much.</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3158828": "I wanted to share some external data that I have been collecting and annotating over the past week. This has improved CV/LB scores for my pipeline and hopefully it helps others too!\n\nThe dataset contains 1617 flagellar motor annotations and 1288 tomograms, all sourced from 62 datasets from the CZII data portal. You can visualize the annotations and preprocessing pipeline at the links below.\n\nDataset [here](https://www.kaggle.com/datasets/brendanartley/cryoet-flagellar-motors-dataset).\nCode [here](https://www.kaggle.com/code/brendanartley/flagellar-motors-dataset-code).\n\nHappy Kaggling 😀\n\n![image_sample](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F9e36df03f6bb44ee14dc22458f9240fc%2Fimg.jpg?generation=1742863339495044&alt=media)\n\n",
    "3160527": "Appreciate all your work and sharing",
    "3173157": "**Update:**\n\nAfter some feedback, I uploaded a `jpgs` folder to this dataset to match the competition data format. This should make integration with existing pipelines easier.\n\nIn addition, if anyone finds more tomograms to annotate, please let me know!",
    "3176421": "thank  you very much for sharing.  \n\nAnyone could share the PARSE DATA of  combined competition dataset and extradata?  \nSince most of us don't have enough resources to process big amount of data.",
    "3214190": "Thanks for sharing, this work is very helpful for me",
    "3159398": "@brendanartley, Thanks a lot. Did you manually annotate them all after preprocessing or some of them have annotations already and you rescale it for new size?",
    "3158869": "Thank you for providing this dataset.\n\nTo my eye, some of the dots are close to or at the internal membrane, meaning they are \"lower\" than the location of flagella in the competition dataset. It probably works when a large enough bounding box is used.",
    "3186200": "@brendanartley Hi. Thank you for sharing this nice dataset.\n\nI found tiny amount of irregular pixel data in this dataset (~4.4%).\nIt seems because of normalizing pixel value in `uint8` datatype. \nMaybe normalizing with float32 data type will fix this issue.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F50f64b4e583486c15e68b96d01ccb45c%2Fvisualization_002.jpeg?generation=1745491086858616&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F79920aa53aa4d108e5404632bd3c5e5d%2Fvisualization.jpeg?generation=1745491101892394&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F6664d640b06875ef787cc2fc8f355807%2Fpixel_dist.jpeg?generation=1745491115831111&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F264812db967341b3b3918a18e38dd5aa%2Fhistogram_irregular.jpeg?generation=1745491137864416&alt=media)",
    "3191797": "Hi BARTLEY,\n\nThank you so much for sharing your data and work!\n\nI have integrated the extra data into my training. My strategy was to add your data to my training set while keeping the validation data unchanged (the same as before, when I was only using the competition data).\n\nHowever, I ran into an issue: with exactly the same settings, I achieved a **0.612** score using only the competition data, but after integrating your data, the score dropped to **0.555**. In other words, adding more training data actually degraded the LB score.\n\nBelow, I’ve attached the training/validation loss curves for (1) using only competition data and (2) using the integrated dataset.\nDo you have any insights on how to interpret this behavior? And do you have suggestions on how to use external data properly?\n![only competition data](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2580499%2F43242816d5585f0661b6e7080ee21142%2Fdfl_loss_curve%20(1).png?generation=1746168033304680&alt=media)\n![integrated extra data into trainining](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2580499%2F5e350cb277b55aaeb67835d1c31463d1%2Fdfl_loss_curve.png?generation=1746168058709572&alt=media)\n\nBest \nLeo\n\n",
    "3168412": "感谢你所有的工作和分享",
    "3180527": "Thank you for sharing extra data! I am trying to add partial or all data on top of original data but the performance seems worse. Can anyone share some tips or tricks how to leverage this data properly? Thank you. ",
    "3159197": "Hello @brendanartley, thank you so much for sharing!\n\nI have a question regarding the preprocessing steps applied, especially resizing. Did you do sub tomogram averaging for reducing the z dimension to 128?\n\nEDIT: Just looked at the code and noticed you include everything there, ignore the question! Thanks!",
    "3158878": "Appreciate all your work and sharing, this is magnificent!",
    "3158844": "if you don't mind sharing to how much extent did extra data effect your lb socre",
    "3161260": "you have done nice work\n",
    "3160472": "wow! great job!",
    "3236336": "Thank you. I’ll use this as a reference.",
    "3183082": "thanks a lot of the extra data [im just in my early steps on that field btw]😄, so i'm just wondering if your rate in the Leaderboard  is because you do it manually or you have a strong model with that Fb_score? ",
    "3160894": "",
    "3160892": "",
    "3218027": "thanks for sharing....\n",
    "3216164": "Thanks for sharing",
    "3159042": "Thanks a lot for sharing this!",
    "3191748": "thanks you very much for your help ",
    "3178198": "It was very helpful, thank you.",
    "3176753": "Thank you very much for sharing ",
    "3185587": "thank you very much."
  }
}