{
  "id": 575503,
  "title": "As the dataset size is large.",
  "url": "/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/575503",
  "author_name": "",
  "post_date": "2025-04-29T06:05:30.801103200Z",
  "votes": 2,
  "comment_count": 4,
  "views": 0,
  "content": "<p>As dataset is large. exceeding RAM and output-disk limit<br>\nWhat techniques are getting used to preprocess the dataset in kaggle-environment?<br>\nTo train the model on it later.</p>\n<p>Also does increasing data via augmentation seems feasible?<br>\nIt may make the data whooping to 150GB+</p>",
  "messages": [
    {
      "id": "3189384",
      "postDate": "04/29/2025 06:05:30",
      "content": "<p>As dataset is large. exceeding RAM and output-disk limit<br>\nWhat techniques are getting used to preprocess the dataset in kaggle-environment?<br>\nTo train the model on it later.</p>\n<p>Also does increasing data via augmentation seems feasible?<br>\nIt may make the data whooping to 150GB+</p>",
      "rawMarkdown": "As dataset is large. exceeding RAM and output-disk limit\nWhat techniques are getting used to preprocess the dataset in kaggle-environment?\nTo train the model on it later.\n\nAlso does increasing data via augmentation seems feasible?\nIt may make the data whooping to 150GB+",
      "votes": null
    },
    {
      "id": "3189456",
      "postDate": "04/29/2025 08:19:46",
      "content": "<p>You can reduce image size or use only subsets of the tomograms at a time. Depending on your choices and RAM limitations, you might need lazy loading. </p>",
      "rawMarkdown": "You can reduce image size or use only subsets of the tomograms at a time. Depending on your choices and RAM limitations, you might need lazy loading.",
      "votes": null
    },
    {
      "id": "3189467",
      "postDate": "04/29/2025 08:43:53",
      "content": "<p>ok, what techniques you are using, If any?</p>",
      "rawMarkdown": "ok, what techniques you are using, If any?",
      "votes": null
    },
    {
      "id": "3189469",
      "postDate": "04/29/2025 08:45:56",
      "content": "<p>Resize into 640 or 960 and only use positive samples</p>",
      "rawMarkdown": "Resize into 640 or 960 and only use positive samples",
      "votes": null
    },
    {
      "id": "3190432",
      "postDate": "04/30/2025 16:57:01",
      "content": "<p>You can consider resizing the inputs or even tiling them out.  I am a fan of lazy loading for this scenario where you want to load the entire tomo at once so you can load them in batches so while your model trains on one set of tomos the next set is being loaded via a second process.</p>",
      "rawMarkdown": "You can consider resizing the inputs or even tiling them out.  I am a fan of lazy loading for this scenario where you want to load the entire tomo at once so you can load them in batches so while your model trains on one set of tomos the next set is being loaded via a second process.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3189456,
      "author_name": "tennogh",
      "author_url": "",
      "post_date": "04/29/2025 08:19:46",
      "content": "<p>You can reduce image size or use only subsets of the tomograms at a time. Depending on your choices and RAM limitations, you might need lazy loading. </p>",
      "votes": null,
      "replies": [
        {
          "id": 3189467,
          "author_name": "arpit1bansal",
          "author_url": "",
          "post_date": "04/29/2025 08:43:53",
          "content": "<p>ok, what techniques you are using, If any?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3189469,
      "author_name": "seeingtimes",
      "author_url": "",
      "post_date": "04/29/2025 08:45:56",
      "content": "<p>Resize into 640 or 960 and only use positive samples</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3190432,
      "author_name": "connorjd",
      "author_url": "",
      "post_date": "04/30/2025 16:57:01",
      "content": "<p>You can consider resizing the inputs or even tiling them out.  I am a fan of lazy loading for this scenario where you want to load the entire tomo at once so you can load them in batches so while your model trains on one set of tomos the next set is being loaded via a second process.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3189384": "As dataset is large. exceeding RAM and output-disk limit\nWhat techniques are getting used to preprocess the dataset in kaggle-environment?\nTo train the model on it later.\n\nAlso does increasing data via augmentation seems feasible?\nIt may make the data whooping to 150GB+",
    "3189456": "You can reduce image size or use only subsets of the tomograms at a time. Depending on your choices and RAM limitations, you might need lazy loading.",
    "3189467": "ok, what techniques you are using, If any?",
    "3189469": "Resize into 640 or 960 and only use positive samples",
    "3190432": "You can consider resizing the inputs or even tiling them out.  I am a fan of lazy loading for this scenario where you want to load the entire tomo at once so you can load them in batches so while your model trains on one set of tomos the next set is being loaded via a second process."
  },
  "source": "meta"
}