{
  "id": 672961,
  "title": "How were the images processed for the pre-trained models?",
  "url": "/competitions/plantclef-2026/discussion/672961",
  "author_name": "",
  "post_date": "2026-02-11T15:14:22.507999800Z",
  "votes": 4,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I read the overview paper for the 2025 plantclef competition but it did specify how the images were processed. </p>\n<p>I want to know if for example the max side 800 or min side 800 data was used, and how this was transformed to 518x518 (assuming this was input size), was it cropped square then resized, just resized, or was augmentation of some sort applied. </p>\n<p>I want to fine tune some more heads for the model but would like to know how the data was processed to match training as it might have an impact. </p>\n<p>If anyone knows who the appropriate person to @ is let me know. Thanks!</p>",
  "messages": [
    {
      "id": "3404921",
      "postDate": "02/11/2026 15:14:22",
      "content": "<p>I read the overview paper for the 2025 plantclef competition but it did specify how the images were processed. </p>\n<p>I want to know if for example the max side 800 or min side 800 data was used, and how this was transformed to 518x518 (assuming this was input size), was it cropped square then resized, just resized, or was augmentation of some sort applied. </p>\n<p>I want to fine tune some more heads for the model but would like to know how the data was processed to match training as it might have an impact. </p>\n<p>If anyone knows who the appropriate person to @ is let me know. Thanks!</p>",
      "rawMarkdown": "I read the overview paper for the 2025 plantclef competition but it did specify how the images were processed. \n\nI want to know if for example the max side 800 or min side 800 data was used, and how this was transformed to 518x518 (assuming this was input size), was it cropped square then resized, just resized, or was augmentation of some sort applied. \n\nI want to fine tune some more heads for the model but would like to know how the data was processed to match training as it might have an impact. \n\nIf anyone knows who the appropriate person to @ is let me know. Thanks!",
      "votes": null
    },
    {
      "id": "3405241",
      "postDate": "02/12/2026 13:56:02",
      "content": "<p><a href=\"https://www.kaggle.com/juliostat\" target=\"_blank\">@juliostat</a> Do you know how the images were processed for pretrained models (or who knows?). Thanks!</p>",
      "rawMarkdown": "juliostat Do you know how the images were processed for pretrained models (or who knows?). Thanks!",
      "votes": null
    },
    {
      "id": "3405258",
      "postDate": "02/12/2026 14:52:49",
      "content": "<p>Hi,</p>\n<p>Based on the training configuration file args.yml that you can find on the original <a href=\"http://zenodo.org/records/10848263\" target=\"_blank\">http://zenodo.org/records/10848263</a>, here are some technical details regarding the image preprocessing and augmentation:</p>\n<ul>\n<li>Cropping &amp; Resizing Strategy The training did not use a simple resize or fixed center crop. The presence of the following parameters confirms the use of a RandomResizedCrop strategy:</li>\n</ul>\n<p>scale: [0.08, 1.0]: The model extracts a random patch covering 8% to 100% of the original image area.</p>\n<p>ratio: [0.75, 1.33]: The aspect ratio of this patch varies between 3/4 and 4/3.</p>\n<p>This random patch is then resized to the model input size (typically 518x518 for a ViT-L/14 DINOv2 architecture to align with patch size 14).</p>\n<ul>\n<li>Augmentation Pipeline The pipeline applies heavy regularization using the following specific parameters:</li>\n</ul>\n<p>RandAugment: aa: rand-m9-mstd0.5-inc1. This applies a sequence of random transformations with magnitude 9 and noise (std 0.5).</p>\n<p>Interpolation: train_interpolation: random. The resizing method cycles (likely between bicubic, bilinear, etc.) during training.</p>\n<p>Mixing: Both mixup: 0.8 and cutmix: 1.0 were active, meaning the model was trained on mixed image inputs rather than single pure images.</p>\n<p>Color Jitter: Set to 0.4.</p>",
      "rawMarkdown": "Hi,\n\nBased on the training configuration file args.yml that you can find on the original http://zenodo.org/records/10848263, here are some technical details regarding the image preprocessing and augmentation:\n\n- Cropping & Resizing Strategy The training did not use a simple resize or fixed center crop. The presence of the following parameters confirms the use of a RandomResizedCrop strategy:\n\nscale: [0.08, 1.0]: The model extracts a random patch covering 8% to 100% of the original image area.\n\nratio: [0.75, 1.33]: The aspect ratio of this patch varies between 3/4 and 4/3.\n\nThis random patch is then resized to the model input size (typically 518x518 for a ViT-L/14 DINOv2 architecture to align with patch size 14).\n\n- Augmentation Pipeline The pipeline applies heavy regularization using the following specific parameters:\n\nRandAugment: aa: rand-m9-mstd0.5-inc1. This applies a sequence of random transformations with magnitude 9 and noise (std 0.5).\n\nInterpolation: train_interpolation: random. The resizing method cycles (likely between bicubic, bilinear, etc.) during training.\n\nMixing: Both mixup: 0.8 and cutmix: 1.0 were active, meaning the model was trained on mixed image inputs rather than single pure images.\n\nColor Jitter: Set to 0.4.",
      "votes": null
    },
    {
      "id": "3405264",
      "postDate": "02/12/2026 15:05:36",
      "content": "<p>Thank so much! I didn't see that file.</p>",
      "rawMarkdown": "Thank so much! I didn't see that file.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3405241,
      "author_name": "devinanzelmo",
      "author_url": "",
      "post_date": "02/12/2026 13:56:02",
      "content": "<p><a href=\"https://www.kaggle.com/juliostat\" target=\"_blank\">@juliostat</a> Do you know how the images were processed for pretrained models (or who knows?). Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 3405258,
          "author_name": "hgoeau",
          "author_url": "",
          "post_date": "02/12/2026 14:52:49",
          "content": "<p>Hi,</p>\n<p>Based on the training configuration file args.yml that you can find on the original <a href=\"http://zenodo.org/records/10848263\" target=\"_blank\">http://zenodo.org/records/10848263</a>, here are some technical details regarding the image preprocessing and augmentation:</p>\n<ul>\n<li>Cropping &amp; Resizing Strategy The training did not use a simple resize or fixed center crop. The presence of the following parameters confirms the use of a RandomResizedCrop strategy:</li>\n</ul>\n<p>scale: [0.08, 1.0]: The model extracts a random patch covering 8% to 100% of the original image area.</p>\n<p>ratio: [0.75, 1.33]: The aspect ratio of this patch varies between 3/4 and 4/3.</p>\n<p>This random patch is then resized to the model input size (typically 518x518 for a ViT-L/14 DINOv2 architecture to align with patch size 14).</p>\n<ul>\n<li>Augmentation Pipeline The pipeline applies heavy regularization using the following specific parameters:</li>\n</ul>\n<p>RandAugment: aa: rand-m9-mstd0.5-inc1. This applies a sequence of random transformations with magnitude 9 and noise (std 0.5).</p>\n<p>Interpolation: train_interpolation: random. The resizing method cycles (likely between bicubic, bilinear, etc.) during training.</p>\n<p>Mixing: Both mixup: 0.8 and cutmix: 1.0 were active, meaning the model was trained on mixed image inputs rather than single pure images.</p>\n<p>Color Jitter: Set to 0.4.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3405264,
              "author_name": "devinanzelmo",
              "author_url": "",
              "post_date": "02/12/2026 15:05:36",
              "content": "<p>Thank so much! I didn't see that file.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3404921": "I read the overview paper for the 2025 plantclef competition but it did specify how the images were processed. \n\nI want to know if for example the max side 800 or min side 800 data was used, and how this was transformed to 518x518 (assuming this was input size), was it cropped square then resized, just resized, or was augmentation of some sort applied. \n\nI want to fine tune some more heads for the model but would like to know how the data was processed to match training as it might have an impact. \n\nIf anyone knows who the appropriate person to @ is let me know. Thanks!",
    "3405241": "juliostat Do you know how the images were processed for pretrained models (or who knows?). Thanks!",
    "3405258": "Hi,\n\nBased on the training configuration file args.yml that you can find on the original http://zenodo.org/records/10848263, here are some technical details regarding the image preprocessing and augmentation:\n\n- Cropping & Resizing Strategy The training did not use a simple resize or fixed center crop. The presence of the following parameters confirms the use of a RandomResizedCrop strategy:\n\nscale: [0.08, 1.0]: The model extracts a random patch covering 8% to 100% of the original image area.\n\nratio: [0.75, 1.33]: The aspect ratio of this patch varies between 3/4 and 4/3.\n\nThis random patch is then resized to the model input size (typically 518x518 for a ViT-L/14 DINOv2 architecture to align with patch size 14).\n\n- Augmentation Pipeline The pipeline applies heavy regularization using the following specific parameters:\n\nRandAugment: aa: rand-m9-mstd0.5-inc1. This applies a sequence of random transformations with magnitude 9 and noise (std 0.5).\n\nInterpolation: train_interpolation: random. The resizing method cycles (likely between bicubic, bilinear, etc.) during training.\n\nMixing: Both mixup: 0.8 and cutmix: 1.0 were active, meaning the model was trained on mixed image inputs rather than single pure images.\n\nColor Jitter: Set to 0.4.",
    "3405264": "Thank so much! I didn't see that file."
  },
  "source": "meta"
}