{
  "id": 698089,
  "title": "4th place solution — Self-training on test tiles + temporal/cluster post-process",
  "url": "/competitions/plantclef-2026/writeups/4th-place-solution-self-training-on-test-tiles",
  "author_name": "",
  "post_date": "2026-05-08T13:04:49Z",
  "votes": 6,
  "comment_count": 4,
  "views": 0,
  "content": "<h1>4th place solution — Self-training on test tiles + temporal/cluster post-processing</h1>\n<p>Thanks to the organizers and everyone who shared insights. Here is a short writeup of our 4th-place solution (final Public LB <strong>0.4611</strong>, submission <code>submission_r2_full_b10.0_yr3.0.csv</code>).</p>\n<ul>\n<li><strong>Code &amp; full experiment log</strong>: <a href=\"https://github.com/sugupoko/2026_kaggle_PlantCLEF2026\" target=\"_blank\">https://github.com/sugupoko/2026_kaggle_PlantCLEF2026</a></li>\n<li><strong>Worked entirely with <a href=\"https://www.anthropic.com/claude-code\" target=\"_blank\">Claude Code</a></strong>, on top of my own Kaggle starter template: <a href=\"https://github.com/sugupoko/xxxx_kaggle_starterRepository\" target=\"_blank\">https://github.com/sugupoko/xxxx_kaggle_starterRepository</a></li>\n</ul>\n<p>The repository is published as-is — daily reports, session notes, and exploratory dead-ends are all there, so you can see the actual day-to-day process, not a polished retrospective.</p>\n<h2>Overview</h2>\n<p>The core ideas:</p>\n<ol>\n<li><strong>Domain-adapt the official <code>ViTD2PC24All</code> (DINOv2 ViT-B/14) on test-tile pseudo-labels</strong> — head-only, FixMatch-style.</li>\n<li><strong>Iterate the pseudo-labeling once</strong> (Round 2).</li>\n<li><strong>Stack four post-processing priors at inference</strong> — consensus / year / cluster / geo.</li>\n</ol>\n<p>The single biggest jump came from (1)+(2). Post-processing then gave the last few points.</p>\n<h2>Pre-processing &amp; tiling</h2>\n<p>Reproduced the 2025 winner's pipeline:</p>\n<ul>\n<li>Lanczos resize + JPEG re-compression (q=85, YCbCr 4:2:2)</li>\n<li>Multi-scale tiling, scales 1..6 → 91 tiles / image, all at 518 px (matching ViT input)</li>\n<li>Tile aggregation: max-pool over species</li>\n</ul>\n<h2>Self-training (the main contribution)</h2>\n<table>\n<thead>\n<tr>\n<th>step</th>\n<th>data</th>\n<th>model</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>baseline</td>\n<td>—</td>\n<td>official ViTD2PC24All</td>\n<td>0.405</td>\n</tr>\n<tr>\n<td>Round 1</td>\n<td>17,676 test tiles / 340 spp (top1&gt;0.7, top2&lt;0.2), hard label</td>\n<td>head-only, 5 ep</td>\n<td>0.460</td>\n</tr>\n<tr>\n<td><strong>Round 2</strong></td>\n<td>re-inferred with R1 model → new pseudo-labels</td>\n<td>head-only, 5 ep</td>\n<td><strong>0.461</strong></td>\n</tr>\n</tbody>\n</table>\n<p>Choices that mattered:</p>\n<ul>\n<li><strong>Hard labels beat soft labels</strong> by +0.015–0.018. With teacher = student, soft uncertainty carries no information; hard + high threshold = entropy minimization (FixMatch).</li>\n<li><strong>30% pseudo-label sampler</strong>. Without oversampling (the natural ratio is ~2%) the adaptation does nothing.</li>\n<li><strong>Use the <em>last</em> checkpoint, not the <em>best</em></strong>. Val accuracy on PlantCLEF-style single-plant images keeps decreasing as the model adapts to plot-style images, but LB keeps going up. Best-on-val ≠ best-on-test under domain shift.</li>\n<li>Lowering the confidence threshold (top1&gt;0.6) added more pseudo-labels but no gain — extra noise canceled extra coverage.</li>\n</ul>\n<h2>Post-processing stack</h2>\n<p>Applied on top of the per-tile species scores:</p>\n<ol>\n<li><strong>Consensus boost (b=10)</strong> — boost species seen across many quadrats inside the same cluster.</li>\n<li><strong>Year boost (yr=3)</strong> — boost species detected at the same <code>quadrat_id</code> across different years (this is the \"temporal\" prior).</li>\n<li><strong>Cluster prior (α=3)</strong> — per-cluster species prior.</li>\n<li><strong>Geo filter</strong> — drop species impossible at the location.</li>\n</ol>\n<p>The interesting observation: after domain adaptation, the optimal <strong>year boost moved from b=2 → b=3</strong>, while consensus and cluster stayed the same. Adapting to the plot-image domain made cross-year detections of the same species more consistent, so year boost stops generating false positives and can be pushed harder.</p>\n<h2>Ablation (post-processing on the R2 model)</h2>\n<table>\n<thead>\n<tr>\n<th>post-proc</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>raw max-pool</td>\n<td>0.4496</td>\n</tr>\n<tr>\n<td>+ cluster + geo</td>\n<td>0.4514</td>\n</tr>\n<tr>\n<td>+ consensus (b=8)</td>\n<td>0.4588</td>\n</tr>\n<tr>\n<td>+ year (b=2)</td>\n<td>0.4601</td>\n</tr>\n<tr>\n<td>+ year (b=3)</td>\n<td><strong>0.4611</strong></td>\n</tr>\n</tbody>\n</table>\n<h2>What didn't work / didn't help</h2>\n<ul>\n<li>Soft pseudo-labels</li>\n<li>Lower confidence threshold (top1 &gt; 0.6)</li>\n<li>Best-on-val checkpoint</li>\n<li>Adding more pseudo-label rounds beyond Round 2 (no time to fully verify, but Round 2 was already at noise floor)</li>\n</ul>\n<h2>Things I wish I had tried</h2>\n<ul>\n<li>Using the 212K LUCAS unlabeled images for pseudo-labeling (largest unused signal in this competition)</li>\n<li>Unfreezing more than the head in later rounds</li>\n<li>Ensembling R1 and R2 score tensors instead of replacing</li>\n</ul>\n<p>Thanks again — congrats to the top teams!</p>",
  "messages": [
    {
      "id": "3455033",
      "postDate": "05/08/2026 13:03:39",
      "content": "<h1>4th place solution — Self-training on test tiles + temporal/cluster post-processing</h1>\n<p>Thanks to the organizers and everyone who shared insights. Here is a short writeup of our 4th-place solution (final Public LB <strong>0.4611</strong>, submission <code>submission_r2_full_b10.0_yr3.0.csv</code>).</p>\n<ul>\n<li><strong>Code &amp; full experiment log</strong>: <a href=\"https://github.com/sugupoko/2026_kaggle_PlantCLEF2026\" target=\"_blank\">https://github.com/sugupoko/2026_kaggle_PlantCLEF2026</a></li>\n<li><strong>Worked entirely with <a href=\"https://www.anthropic.com/claude-code\" target=\"_blank\">Claude Code</a></strong>, on top of my own Kaggle starter template: <a href=\"https://github.com/sugupoko/xxxx_kaggle_starterRepository\" target=\"_blank\">https://github.com/sugupoko/xxxx_kaggle_starterRepository</a></li>\n</ul>\n<p>The repository is published as-is — daily reports, session notes, and exploratory dead-ends are all there, so you can see the actual day-to-day process, not a polished retrospective.</p>\n<h2>Overview</h2>\n<p>The core ideas:</p>\n<ol>\n<li><strong>Domain-adapt the official <code>ViTD2PC24All</code> (DINOv2 ViT-B/14) on test-tile pseudo-labels</strong> — head-only, FixMatch-style.</li>\n<li><strong>Iterate the pseudo-labeling once</strong> (Round 2).</li>\n<li><strong>Stack four post-processing priors at inference</strong> — consensus / year / cluster / geo.</li>\n</ol>\n<p>The single biggest jump came from (1)+(2). Post-processing then gave the last few points.</p>\n<h2>Pre-processing &amp; tiling</h2>\n<p>Reproduced the 2025 winner's pipeline:</p>\n<ul>\n<li>Lanczos resize + JPEG re-compression (q=85, YCbCr 4:2:2)</li>\n<li>Multi-scale tiling, scales 1..6 → 91 tiles / image, all at 518 px (matching ViT input)</li>\n<li>Tile aggregation: max-pool over species</li>\n</ul>\n<h2>Self-training (the main contribution)</h2>\n<table>\n<thead>\n<tr>\n<th>step</th>\n<th>data</th>\n<th>model</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>baseline</td>\n<td>—</td>\n<td>official ViTD2PC24All</td>\n<td>0.405</td>\n</tr>\n<tr>\n<td>Round 1</td>\n<td>17,676 test tiles / 340 spp (top1&gt;0.7, top2&lt;0.2), hard label</td>\n<td>head-only, 5 ep</td>\n<td>0.460</td>\n</tr>\n<tr>\n<td><strong>Round 2</strong></td>\n<td>re-inferred with R1 model → new pseudo-labels</td>\n<td>head-only, 5 ep</td>\n<td><strong>0.461</strong></td>\n</tr>\n</tbody>\n</table>\n<p>Choices that mattered:</p>\n<ul>\n<li><strong>Hard labels beat soft labels</strong> by +0.015–0.018. With teacher = student, soft uncertainty carries no information; hard + high threshold = entropy minimization (FixMatch).</li>\n<li><strong>30% pseudo-label sampler</strong>. Without oversampling (the natural ratio is ~2%) the adaptation does nothing.</li>\n<li><strong>Use the <em>last</em> checkpoint, not the <em>best</em></strong>. Val accuracy on PlantCLEF-style single-plant images keeps decreasing as the model adapts to plot-style images, but LB keeps going up. Best-on-val ≠ best-on-test under domain shift.</li>\n<li>Lowering the confidence threshold (top1&gt;0.6) added more pseudo-labels but no gain — extra noise canceled extra coverage.</li>\n</ul>\n<h2>Post-processing stack</h2>\n<p>Applied on top of the per-tile species scores:</p>\n<ol>\n<li><strong>Consensus boost (b=10)</strong> — boost species seen across many quadrats inside the same cluster.</li>\n<li><strong>Year boost (yr=3)</strong> — boost species detected at the same <code>quadrat_id</code> across different years (this is the \"temporal\" prior).</li>\n<li><strong>Cluster prior (α=3)</strong> — per-cluster species prior.</li>\n<li><strong>Geo filter</strong> — drop species impossible at the location.</li>\n</ol>\n<p>The interesting observation: after domain adaptation, the optimal <strong>year boost moved from b=2 → b=3</strong>, while consensus and cluster stayed the same. Adapting to the plot-image domain made cross-year detections of the same species more consistent, so year boost stops generating false positives and can be pushed harder.</p>\n<h2>Ablation (post-processing on the R2 model)</h2>\n<table>\n<thead>\n<tr>\n<th>post-proc</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>raw max-pool</td>\n<td>0.4496</td>\n</tr>\n<tr>\n<td>+ cluster + geo</td>\n<td>0.4514</td>\n</tr>\n<tr>\n<td>+ consensus (b=8)</td>\n<td>0.4588</td>\n</tr>\n<tr>\n<td>+ year (b=2)</td>\n<td>0.4601</td>\n</tr>\n<tr>\n<td>+ year (b=3)</td>\n<td><strong>0.4611</strong></td>\n</tr>\n</tbody>\n</table>\n<h2>What didn't work / didn't help</h2>\n<ul>\n<li>Soft pseudo-labels</li>\n<li>Lower confidence threshold (top1 &gt; 0.6)</li>\n<li>Best-on-val checkpoint</li>\n<li>Adding more pseudo-label rounds beyond Round 2 (no time to fully verify, but Round 2 was already at noise floor)</li>\n</ul>\n<h2>Things I wish I had tried</h2>\n<ul>\n<li>Using the 212K LUCAS unlabeled images for pseudo-labeling (largest unused signal in this competition)</li>\n<li>Unfreezing more than the head in later rounds</li>\n<li>Ensembling R1 and R2 score tensors instead of replacing</li>\n</ul>\n<p>Thanks again — congrats to the top teams!</p>",
      "rawMarkdown": "# 4th place solution — Self-training on test tiles + temporal/cluster post-processing\n\nThanks to the organizers and everyone who shared insights. Here is a short writeup of our 4th-place solution (final Public LB **0.4611**, submission `submission_r2_full_b10.0_yr3.0.csv`).\n\n- **Code & full experiment log**: https://github.com/sugupoko/2026_kaggle_PlantCLEF2026\n- **Worked entirely with [Claude Code](https://www.anthropic.com/claude-code)**, on top of my own Kaggle starter template: https://github.com/sugupoko/xxxx_kaggle_starterRepository\n\nThe repository is published as-is — daily reports, session notes, and exploratory dead-ends are all there, so you can see the actual day-to-day process, not a polished retrospective.\n\n## Overview\n\nThe core ideas:\n\n1. **Domain-adapt the official `ViTD2PC24All` (DINOv2 ViT-B/14) on test-tile pseudo-labels** — head-only, FixMatch-style.\n2. **Iterate the pseudo-labeling once** (Round 2).\n3. **Stack four post-processing priors at inference** — consensus / year / cluster / geo.\n\nThe single biggest jump came from (1)+(2). Post-processing then gave the last few points.\n\n## Pre-processing & tiling\n\nReproduced the 2025 winner's pipeline:\n\n- Lanczos resize + JPEG re-compression (q=85, YCbCr 4:2:2)\n- Multi-scale tiling, scales 1..6 → 91 tiles / image, all at 518 px (matching ViT input)\n- Tile aggregation: max-pool over species\n\n## Self-training (the main contribution)\n\n| step | data | model | LB |\n|---|---|---|---|\n| baseline | — | official ViTD2PC24All | 0.405 |\n| Round 1 | 17,676 test tiles / 340 spp (top1>0.7, top2<0.2), hard label | head-only, 5 ep | 0.460 |\n| **Round 2** | re-inferred with R1 model → new pseudo-labels | head-only, 5 ep | **0.461** |\n\nChoices that mattered:\n\n- **Hard labels beat soft labels** by +0.015–0.018. With teacher = student, soft uncertainty carries no information; hard + high threshold = entropy minimization (FixMatch).\n- **30% pseudo-label sampler**. Without oversampling (the natural ratio is ~2%) the adaptation does nothing.\n- **Use the *last* checkpoint, not the *best***. Val accuracy on PlantCLEF-style single-plant images keeps decreasing as the model adapts to plot-style images, but LB keeps going up. Best-on-val ≠ best-on-test under domain shift.\n- Lowering the confidence threshold (top1>0.6) added more pseudo-labels but no gain — extra noise canceled extra coverage.\n\n## Post-processing stack\n\nApplied on top of the per-tile species scores:\n\n1. **Consensus boost (b=10)** — boost species seen across many quadrats inside the same cluster.\n2. **Year boost (yr=3)** — boost species detected at the same `quadrat_id` across different years (this is the \"temporal\" prior).\n3. **Cluster prior (α=3)** — per-cluster species prior.\n4. **Geo filter** — drop species impossible at the location.\n\nThe interesting observation: after domain adaptation, the optimal **year boost moved from b=2 → b=3**, while consensus and cluster stayed the same. Adapting to the plot-image domain made cross-year detections of the same species more consistent, so year boost stops generating false positives and can be pushed harder.\n\n## Ablation (post-processing on the R2 model)\n\n| post-proc | LB |\n|---|---|\n| raw max-pool | 0.4496 |\n| + cluster + geo | 0.4514 |\n| + consensus (b=8) | 0.4588 |\n| + year (b=2) | 0.4601 |\n| + year (b=3) | **0.4611** |\n\n## What didn't work / didn't help\n\n- Soft pseudo-labels\n- Lower confidence threshold (top1 > 0.6)\n- Best-on-val checkpoint\n- Adding more pseudo-label rounds beyond Round 2 (no time to fully verify, but Round 2 was already at noise floor)\n\n## Things I wish I had tried\n\n- Using the 212K LUCAS unlabeled images for pseudo-labeling (largest unused signal in this competition)\n- Unfreezing more than the head in later rounds\n- Ensembling R1 and R2 score tensors instead of replacing\n\nThanks again — congrats to the top teams!",
      "votes": null
    },
    {
      "id": "3455246",
      "postDate": "05/08/2026 22:36:21",
      "content": "<p>Cool solution, thanks for sharing! As a note, we experimented with semi-supervised learning (FixMatch-style) on the LUCAS dataset, but the image distribution seems to differ significantly from the test set. As a result, the pseudo-labels were too noisy to provide a useful learning signal. Would be interesting to see if others can make it work.</p>",
      "rawMarkdown": "Cool solution, thanks for sharing! As a note, we experimented with semi-supervised learning (FixMatch-style) on the LUCAS dataset, but the image distribution seems to differ significantly from the test set. As a result, the pseudo-labels were too noisy to provide a useful learning signal. Would be interesting to see if others can make it work.",
      "votes": null
    },
    {
      "id": "3456211",
      "postDate": "05/11/2026 13:24:57",
      "content": "<p>Thank you for sharing this detailed writeup and for your efforts throughout the competition. Congratulations on achieving 4th place!  Your observations regarding the self-training approach, the domain shift, and the post-processing rules are highly relevant. We hope you will consider writing these findings and sharing them by submitting a working note for the LifeCLEF evaluation campaign :) </p>",
      "rawMarkdown": "Thank you for sharing this detailed writeup and for your efforts throughout the competition. Congratulations on achieving 4th place!  Your observations regarding the self-training approach, the domain shift, and the post-processing rules are highly relevant. We hope you will consider writing these findings and sharing them by submitting a working note for the LifeCLEF evaluation campaign :)",
      "votes": null
    },
    {
      "id": "3456669",
      "postDate": "05/12/2026 11:49:12",
      "content": "<p>Wow, this is a really strong solution. I especially like the self-training on test tiles and the temporal/cluster post-processing. Congratulations on 4th place and thank you for sharing</p>",
      "rawMarkdown": "Wow, this is a really strong solution. I especially like the self-training on test tiles and the temporal/cluster post-processing. Congratulations on 4th place and thank you for sharing",
      "votes": null
    },
    {
      "id": "3456969",
      "postDate": "05/12/2026 20:40:44",
      "content": "<p>Thanks for sharing with such important knowledge and experience. Congratulations on 4th place, great job!</p>",
      "rawMarkdown": "Thanks for sharing with such important knowledge and experience. Congratulations on 4th place, great job!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3455246,
      "author_name": "alpere",
      "author_url": "",
      "post_date": "05/08/2026 22:36:21",
      "content": "<p>Cool solution, thanks for sharing! As a note, we experimented with semi-supervised learning (FixMatch-style) on the LUCAS dataset, but the image distribution seems to differ significantly from the test set. As a result, the pseudo-labels were too noisy to provide a useful learning signal. Would be interesting to see if others can make it work.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3456211,
      "author_name": "hgoeau",
      "author_url": "",
      "post_date": "05/11/2026 13:24:57",
      "content": "<p>Thank you for sharing this detailed writeup and for your efforts throughout the competition. Congratulations on achieving 4th place!  Your observations regarding the self-training approach, the domain shift, and the post-processing rules are highly relevant. We hope you will consider writing these findings and sharing them by submitting a working note for the LifeCLEF evaluation campaign :) </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3456669,
      "author_name": "stevefffff",
      "author_url": "",
      "post_date": "05/12/2026 11:49:12",
      "content": "<p>Wow, this is a really strong solution. I especially like the self-training on test tiles and the temporal/cluster post-processing. Congratulations on 4th place and thank you for sharing</p>",
      "votes": null,
      "replies": [
        {
          "id": 3456969,
          "author_name": "sardorrazikov",
          "author_url": "",
          "post_date": "05/12/2026 20:40:44",
          "content": "<p>Thanks for sharing with such important knowledge and experience. Congratulations on 4th place, great job!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3455033": "# 4th place solution — Self-training on test tiles + temporal/cluster post-processing\n\nThanks to the organizers and everyone who shared insights. Here is a short writeup of our 4th-place solution (final Public LB **0.4611**, submission `submission_r2_full_b10.0_yr3.0.csv`).\n\n- **Code & full experiment log**: https://github.com/sugupoko/2026_kaggle_PlantCLEF2026\n- **Worked entirely with [Claude Code](https://www.anthropic.com/claude-code)**, on top of my own Kaggle starter template: https://github.com/sugupoko/xxxx_kaggle_starterRepository\n\nThe repository is published as-is — daily reports, session notes, and exploratory dead-ends are all there, so you can see the actual day-to-day process, not a polished retrospective.\n\n## Overview\n\nThe core ideas:\n\n1. **Domain-adapt the official `ViTD2PC24All` (DINOv2 ViT-B/14) on test-tile pseudo-labels** — head-only, FixMatch-style.\n2. **Iterate the pseudo-labeling once** (Round 2).\n3. **Stack four post-processing priors at inference** — consensus / year / cluster / geo.\n\nThe single biggest jump came from (1)+(2). Post-processing then gave the last few points.\n\n## Pre-processing & tiling\n\nReproduced the 2025 winner's pipeline:\n\n- Lanczos resize + JPEG re-compression (q=85, YCbCr 4:2:2)\n- Multi-scale tiling, scales 1..6 → 91 tiles / image, all at 518 px (matching ViT input)\n- Tile aggregation: max-pool over species\n\n## Self-training (the main contribution)\n\n| step | data | model | LB |\n|---|---|---|---|\n| baseline | — | official ViTD2PC24All | 0.405 |\n| Round 1 | 17,676 test tiles / 340 spp (top1>0.7, top2<0.2), hard label | head-only, 5 ep | 0.460 |\n| **Round 2** | re-inferred with R1 model → new pseudo-labels | head-only, 5 ep | **0.461** |\n\nChoices that mattered:\n\n- **Hard labels beat soft labels** by +0.015–0.018. With teacher = student, soft uncertainty carries no information; hard + high threshold = entropy minimization (FixMatch).\n- **30% pseudo-label sampler**. Without oversampling (the natural ratio is ~2%) the adaptation does nothing.\n- **Use the *last* checkpoint, not the *best***. Val accuracy on PlantCLEF-style single-plant images keeps decreasing as the model adapts to plot-style images, but LB keeps going up. Best-on-val ≠ best-on-test under domain shift.\n- Lowering the confidence threshold (top1>0.6) added more pseudo-labels but no gain — extra noise canceled extra coverage.\n\n## Post-processing stack\n\nApplied on top of the per-tile species scores:\n\n1. **Consensus boost (b=10)** — boost species seen across many quadrats inside the same cluster.\n2. **Year boost (yr=3)** — boost species detected at the same `quadrat_id` across different years (this is the \"temporal\" prior).\n3. **Cluster prior (α=3)** — per-cluster species prior.\n4. **Geo filter** — drop species impossible at the location.\n\nThe interesting observation: after domain adaptation, the optimal **year boost moved from b=2 → b=3**, while consensus and cluster stayed the same. Adapting to the plot-image domain made cross-year detections of the same species more consistent, so year boost stops generating false positives and can be pushed harder.\n\n## Ablation (post-processing on the R2 model)\n\n| post-proc | LB |\n|---|---|\n| raw max-pool | 0.4496 |\n| + cluster + geo | 0.4514 |\n| + consensus (b=8) | 0.4588 |\n| + year (b=2) | 0.4601 |\n| + year (b=3) | **0.4611** |\n\n## What didn't work / didn't help\n\n- Soft pseudo-labels\n- Lower confidence threshold (top1 > 0.6)\n- Best-on-val checkpoint\n- Adding more pseudo-label rounds beyond Round 2 (no time to fully verify, but Round 2 was already at noise floor)\n\n## Things I wish I had tried\n\n- Using the 212K LUCAS unlabeled images for pseudo-labeling (largest unused signal in this competition)\n- Unfreezing more than the head in later rounds\n- Ensembling R1 and R2 score tensors instead of replacing\n\nThanks again — congrats to the top teams!",
    "3455246": "Cool solution, thanks for sharing! As a note, we experimented with semi-supervised learning (FixMatch-style) on the LUCAS dataset, but the image distribution seems to differ significantly from the test set. As a result, the pseudo-labels were too noisy to provide a useful learning signal. Would be interesting to see if others can make it work.",
    "3456211": "Thank you for sharing this detailed writeup and for your efforts throughout the competition. Congratulations on achieving 4th place!  Your observations regarding the self-training approach, the domain shift, and the post-processing rules are highly relevant. We hope you will consider writing these findings and sharing them by submitting a working note for the LifeCLEF evaluation campaign :)",
    "3456669": "Wow, this is a really strong solution. I especially like the self-training on test tiles and the temporal/cluster post-processing. Congratulations on 4th place and thank you for sharing",
    "3456969": "Thanks for sharing with such important knowledge and experience. Congratulations on 4th place, great job!"
  },
  "source": "meta"
}