{
  "id": 574340,
  "title": "Extra data harm my LB by a lot",
  "url": "/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/574340",
  "author_name": "",
  "post_date": "2025-04-21T12:10:32.478896400Z",
  "votes": 5,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hey guys, i have incorporated the extra data from <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a>. I use the public notebook of yolo V8 and i notice a huge drop in the LB after using the extra data. Do i miss something? Anybody with a similar problem or sugestions how to solve?</p>\n<p>this is how i transformed them into jpg:</p>\n<pre><code> numpy  np\n os\n PIL  Image\n\ndirectory_path = \noutput_folder = \n\nfile_names = os.listdir(directory_path)\n num, file  (file_names):\n    (num)\n    tomograph = np.load()\n    (tomograph.shape)\n    os.makedirs(, exist_ok=)\n     idx,   (tomograph):\n        img = Image.fromarray()\n        img.save(os.path.join(, ))\n\n\n()\n(, (file_names))\n</code></pre>\n<p>And this is how i concat the labels</p>\n<pre><code>labels_df = pd(os(DATA_DIR, ))\n train_file == :\n    extra_labels = pd()\n    extra_labels = extra_labels(={\n        : ,\n        : ,\n        : ,\n    })\n\n    extra_labels = \n    extra_labels = \n    extra_labels = \n    extra_labels = \n\n    labels_df = pd(, ignore_index=True)\n</code></pre>",
  "messages": [
    {
      "id": "3183893",
      "postDate": "04/21/2025 12:10:32",
      "content": "<p>Hey guys, i have incorporated the extra data from <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a>. I use the public notebook of yolo V8 and i notice a huge drop in the LB after using the extra data. Do i miss something? Anybody with a similar problem or sugestions how to solve?</p>\n<p>this is how i transformed them into jpg:</p>\n<pre><code> numpy  np\n os\n PIL  Image\n\ndirectory_path = \noutput_folder = \n\nfile_names = os.listdir(directory_path)\n num, file  (file_names):\n    (num)\n    tomograph = np.load()\n    (tomograph.shape)\n    os.makedirs(, exist_ok=)\n     idx,   (tomograph):\n        img = Image.fromarray()\n        img.save(os.path.join(, ))\n\n\n()\n(, (file_names))\n</code></pre>\n<p>And this is how i concat the labels</p>\n<pre><code>labels_df = pd(os(DATA_DIR, ))\n train_file == :\n    extra_labels = pd()\n    extra_labels = extra_labels(={\n        : ,\n        : ,\n        : ,\n    })\n\n    extra_labels = \n    extra_labels = \n    extra_labels = \n    extra_labels = \n\n    labels_df = pd(, ignore_index=True)\n</code></pre>",
      "rawMarkdown": "Hey guys, i have incorporated the extra data from @brendanartley. I use the public notebook of yolo V8 and i notice a huge drop in the LB after using the extra data. Do i miss something? Anybody with a similar problem or sugestions how to solve?\n\nthis is how i transformed them into jpg:\n```\nimport numpy as np\nimport os\nfrom PIL import Image\n\ndirectory_path = 'archive/volumes/'\noutput_folder = 'kaggle/input/byu-locating-bacterial-flagellar-motors-2025/train_big/'\n\nfile_names = os.listdir(directory_path)\nfor num, file in enumerate(file_names):\n    print(num)\n    tomograph = np.load(f\"{directory_path}{file}\")\n    print(tomograph.shape)\n    os.makedirs(f\"{output_folder}{file[:-4]}\", exist_ok=True)\n    for idx, slice in enumerate(tomograph):\n        img = Image.fromarray(slice)\n        img.save(os.path.join(f\"{output_folder}{file[:-4]}\", f'slice_{idx:04d}.jpg'))\n\n\nprint()\nprint(\"Number of files:\", len(file_names))\n\n```\n\nAnd this is how i concat the labels\n\n```\nlabels_df = pd.read_csv(os.path.join(DATA_DIR, \"train_labels.csv\"))\nif train_file == \"train_big\":\n    extra_labels = pd.read_csv(\"../archive/labels.csv\")\n    extra_labels = extra_labels.rename(columns={\n        \"z\": \"Motor axis 0\",\n        \"y\": \"Motor axis 1\",\n        \"x\": \"Motor axis 2\",\n    })\n\n    extra_labels['Number of motors'] = 1\n    extra_labels['Array shape (axis 0)'] = 128\n    extra_labels['Array shape (axis 1)'] = 512\n    extra_labels['Array shape (axis 2)'] = 512\n\n    labels_df = pd.concat([labels_df, extra_labels], ignore_index=True)\n\n```",
      "votes": null
    },
    {
      "id": "3184253",
      "postDate": "04/21/2025 20:58:08",
      "content": "<p>can you give public notebook with this code?</p>",
      "rawMarkdown": "can you give public notebook with this code?",
      "votes": null
    },
    {
      "id": "3185572",
      "postDate": "04/23/2025 14:15:31",
      "content": "<p>I tested my model with an external-only dataset, and it achieved 0.44 on the LB.</p>",
      "rawMarkdown": "I tested my model with an external-only dataset, and it achieved 0.44 on the LB.",
      "votes": null
    },
    {
      "id": "3187546",
      "postDate": "04/26/2025 07:15:58",
      "content": "<p>So with a different train_val / test split the extra data help raise the score (in the test set), while in my first train_val / test split they were deteriorating the test score but in my first split i noticed that there were very few slices with 1 motor.</p>",
      "rawMarkdown": "So with a different train_val / test split the extra data help raise the score (in the test set), while in my first train_val / test split they were deteriorating the test score but in my first split i noticed that there were very few slices with 1 motor.",
      "votes": null
    },
    {
      "id": "3188081",
      "postDate": "04/27/2025 02:47:05",
      "content": "<p>Hi everyone, Thanks for sharing your experience!<br>\nCurious if others tried weighting or filtering external slices more carefully? Thanks again</p>",
      "rawMarkdown": "Hi everyone, Thanks for sharing your experience!\nCurious if others tried weighting or filtering external slices more carefully? Thanks again",
      "votes": null
    },
    {
      "id": "3231589",
      "postDate": "06/24/2025 16:07:42",
      "content": "<p>No need in it at all</p>",
      "rawMarkdown": "No need in it at all",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3184253,
      "author_name": "",
      "author_url": "",
      "post_date": "04/21/2025 20:58:08",
      "content": "<p>can you give public notebook with this code?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3231589,
          "author_name": "",
          "author_url": "",
          "post_date": "06/24/2025 16:07:42",
          "content": "<p>No need in it at all</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3185572,
      "author_name": "seeingtimes",
      "author_url": "",
      "post_date": "04/23/2025 14:15:31",
      "content": "<p>I tested my model with an external-only dataset, and it achieved 0.44 on the LB.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3187546,
      "author_name": "vasileioscharatsidis",
      "author_url": "",
      "post_date": "04/26/2025 07:15:58",
      "content": "<p>So with a different train_val / test split the extra data help raise the score (in the test set), while in my first train_val / test split they were deteriorating the test score but in my first split i noticed that there were very few slices with 1 motor.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3188081,
      "author_name": "achievement",
      "author_url": "",
      "post_date": "04/27/2025 02:47:05",
      "content": "<p>Hi everyone, Thanks for sharing your experience!<br>\nCurious if others tried weighting or filtering external slices more carefully? Thanks again</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3183893": "Hey guys, i have incorporated the extra data from @brendanartley. I use the public notebook of yolo V8 and i notice a huge drop in the LB after using the extra data. Do i miss something? Anybody with a similar problem or sugestions how to solve?\n\nthis is how i transformed them into jpg:\n```\nimport numpy as np\nimport os\nfrom PIL import Image\n\ndirectory_path = 'archive/volumes/'\noutput_folder = 'kaggle/input/byu-locating-bacterial-flagellar-motors-2025/train_big/'\n\nfile_names = os.listdir(directory_path)\nfor num, file in enumerate(file_names):\n    print(num)\n    tomograph = np.load(f\"{directory_path}{file}\")\n    print(tomograph.shape)\n    os.makedirs(f\"{output_folder}{file[:-4]}\", exist_ok=True)\n    for idx, slice in enumerate(tomograph):\n        img = Image.fromarray(slice)\n        img.save(os.path.join(f\"{output_folder}{file[:-4]}\", f'slice_{idx:04d}.jpg'))\n\n\nprint()\nprint(\"Number of files:\", len(file_names))\n\n```\n\nAnd this is how i concat the labels\n\n```\nlabels_df = pd.read_csv(os.path.join(DATA_DIR, \"train_labels.csv\"))\nif train_file == \"train_big\":\n    extra_labels = pd.read_csv(\"../archive/labels.csv\")\n    extra_labels = extra_labels.rename(columns={\n        \"z\": \"Motor axis 0\",\n        \"y\": \"Motor axis 1\",\n        \"x\": \"Motor axis 2\",\n    })\n\n    extra_labels['Number of motors'] = 1\n    extra_labels['Array shape (axis 0)'] = 128\n    extra_labels['Array shape (axis 1)'] = 512\n    extra_labels['Array shape (axis 2)'] = 512\n\n    labels_df = pd.concat([labels_df, extra_labels], ignore_index=True)\n\n```",
    "3184253": "can you give public notebook with this code?",
    "3185572": "I tested my model with an external-only dataset, and it achieved 0.44 on the LB.",
    "3187546": "So with a different train_val / test split the extra data help raise the score (in the test set), while in my first train_val / test split they were deteriorating the test score but in my first split i noticed that there were very few slices with 1 motor.",
    "3188081": "Hi everyone, Thanks for sharing your experience!\nCurious if others tried weighting or filtering external slices more carefully? Thanks again",
    "3231589": "No need in it at all"
  },
  "source": "meta"
}