{
  "id": 666251,
  "title": "CV vs Public LB",
  "url": "/competitions/recodai-luc-scientific-image-forgery-detection/discussion/666251",
  "author_name": "Kh0a",
  "post_date": "2026-01-06T03:26:03.028000",
  "votes": 0,
  "comment_count": 3,
  "views": 0,
  "content": "<p>My DinoV3 pipeline makes good use of the training dataset, but it is hard to evaluate performance using the public leaderboard. </p>\n<p>According to the dataset description</p>\n<blockquote>\n  <p>The competition will proceed in two phases:</p>\n  <p>1.A model training phase with a public leaderboard test set of roughly 1,100 images. Because these images are from publicly available research papers leaderboard scores during this phase are not meaningful.</p>\n  <p>2.A forecasting phase will add a private leaderboard test set to be collected after submissions close. Expect the additional images to roughly double the size of the test set.</p>\n</blockquote>\n<p>Also, the host cannot confirm whether the private test set will contain publicly available datasets or data with the same distribution as the training set.</p>\n<p>Table (examples of runs / models)</p>\n<table>\n<thead>\n<tr>\n<th>CV</th>\n<th>AuthF1</th>\n<th>ForgedF1</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.4592</td>\n<td>1.0000</td>\n<td>0.0000</td>\n<td>0.303</td>\n</tr>\n<tr>\n<td>0.7636</td>\n<td>0.9700</td>\n<td>0.5800</td>\n<td>0.223</td>\n</tr>\n<tr>\n<td>0.7577</td>\n<td>0.9663</td>\n<td>0.5806</td>\n<td>0.204</td>\n</tr>\n<tr>\n<td>0.6515</td>\n<td>0.9558</td>\n<td>0.3930</td>\n<td>0.186</td>\n</tr>\n<tr>\n<td>0.6691</td>\n<td>0.9651</td>\n<td>0.4178</td>\n<td>0.242</td>\n</tr>\n</tbody>\n</table>\n<p>I've seen many people do well on the public leaderboard — do your CV/OOF scores match your public LB scores?</p>",
  "messages": [
    {
      "id": 3392090,
      "postDate": "2026-01-16T09:20:10.177Z",
      "content": "<p>What did you get on supplemented images? As described \"These new images were acquired and labeled using the exact process that will be used for the final test.\" so maybe that score will be important. I got 0.2 using the official competition F1 metric with 1/5 fold + 1 authentic image in an extra sub validation during training. Easier to segment the corn images vs supplemented images from research papers! :)</p>",
      "rawMarkdown": "What did you get on supplemented images? As described \"These new images were acquired and labeled using the exact process that will be used for the final test.\" so maybe that score will be important. I got 0.2 using the official competition F1 metric with 1/5 fold + 1 authentic image in an extra sub validation during training. Easier to segment the corn images vs supplemented images from research papers! :)"
    },
    {
      "id": 3386956,
      "postDate": "2026-01-06T06:25:25.367Z",
      "content": "<p>I think the images and masks in supply_img should be informative (and useful for guidance). When I score on supply_img, I can even reach around 0.65, but on the test set it’s only about 0.33.</p>",
      "rawMarkdown": "I think the images and masks in supply_img should be informative (and useful for guidance). When I score on supply_img, I can even reach around 0.65, but on the test set it’s only about 0.33.\n",
      "replies": [
        {
          "id": 3387694,
          "postDate": "2026-01-07T13:31:45.600Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 3386900,
      "postDate": "2026-01-06T03:26:03.027Z",
      "content": "<p>My DinoV3 pipeline makes good use of the training dataset, but it is hard to evaluate performance using the public leaderboard. </p>\n<p>According to the dataset description</p>\n<blockquote>\n  <p>The competition will proceed in two phases:</p>\n  <p>1.A model training phase with a public leaderboard test set of roughly 1,100 images. Because these images are from publicly available research papers leaderboard scores during this phase are not meaningful.</p>\n  <p>2.A forecasting phase will add a private leaderboard test set to be collected after submissions close. Expect the additional images to roughly double the size of the test set.</p>\n</blockquote>\n<p>Also, the host cannot confirm whether the private test set will contain publicly available datasets or data with the same distribution as the training set.</p>\n<p>Table (examples of runs / models)</p>\n<table>\n<thead>\n<tr>\n<th>CV</th>\n<th>AuthF1</th>\n<th>ForgedF1</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.4592</td>\n<td>1.0000</td>\n<td>0.0000</td>\n<td>0.303</td>\n</tr>\n<tr>\n<td>0.7636</td>\n<td>0.9700</td>\n<td>0.5800</td>\n<td>0.223</td>\n</tr>\n<tr>\n<td>0.7577</td>\n<td>0.9663</td>\n<td>0.5806</td>\n<td>0.204</td>\n</tr>\n<tr>\n<td>0.6515</td>\n<td>0.9558</td>\n<td>0.3930</td>\n<td>0.186</td>\n</tr>\n<tr>\n<td>0.6691</td>\n<td>0.9651</td>\n<td>0.4178</td>\n<td>0.242</td>\n</tr>\n</tbody>\n</table>\n<p>I've seen many people do well on the public leaderboard — do your CV/OOF scores match your public LB scores?</p>",
      "rawMarkdown": "My DinoV3 pipeline makes good use of the training dataset, but it is hard to evaluate performance using the public leaderboard. \n\nAccording to the dataset description\n\n>The competition will proceed in two phases:\n\n>1.A model training phase with a public leaderboard test set of roughly 1,100 images. Because these images are from publicly available research papers leaderboard scores during this phase are not meaningful.\n\n\n>2.A forecasting phase will add a private leaderboard test set to be collected after submissions close. Expect the additional images to roughly double the size of the test set.\n\nAlso, the host cannot confirm whether the private test set will contain publicly available datasets or data with the same distribution as the training set.\n\nTable (examples of runs / models)\n\n| CV     | AuthF1 | ForgedF1 | LB    |\n|-------:|:------:|:--------:|:-----:|\n| 0.4592 | 1.0000 | 0.0000   | 0.303 |\n| 0.7636 | 0.9700 | 0.5800   | 0.223 |\n| 0.7577 | 0.9663 | 0.5806   | 0.204 |\n| 0.6515 | 0.9558 | 0.3930   | 0.186 |\n| 0.6691 | 0.9651 | 0.4178   | 0.242 |\n\nI've seen many people do well on the public leaderboard — do your CV/OOF scores match your public LB scores?"
    }
  ],
  "comments": [
    {
      "id": 3392090,
      "author_name": "Kirderf",
      "author_url": "",
      "post_date": "2026-01-16T09:20:10.177000",
      "content": "<p>What did you get on supplemented images? As described \"These new images were acquired and labeled using the exact process that will be used for the final test.\" so maybe that score will be important. I got 0.2 using the official competition F1 metric with 1/5 fold + 1 authentic image in an extra sub validation during training. Easier to segment the corn images vs supplemented images from research papers! :)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3386956,
      "author_name": "",
      "author_url": "",
      "post_date": "2026-01-06T06:25:25.367000",
      "content": "<p>I think the images and masks in supply_img should be informative (and useful for guidance). When I score on supply_img, I can even reach around 0.65, but on the test set it’s only about 0.33.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3387694,
          "author_name": "",
          "author_url": "",
          "post_date": "2026-01-07T13:31:45.600000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3392090": "What did you get on supplemented images? As described \"These new images were acquired and labeled using the exact process that will be used for the final test.\" so maybe that score will be important. I got 0.2 using the official competition F1 metric with 1/5 fold + 1 authentic image in an extra sub validation during training. Easier to segment the corn images vs supplemented images from research papers! :)",
    "3386956": "I think the images and masks in supply_img should be informative (and useful for guidance). When I score on supply_img, I can even reach around 0.65, but on the test set it’s only about 0.33.\n",
    "3386900": "My DinoV3 pipeline makes good use of the training dataset, but it is hard to evaluate performance using the public leaderboard. \n\nAccording to the dataset description\n\n>The competition will proceed in two phases:\n\n>1.A model training phase with a public leaderboard test set of roughly 1,100 images. Because these images are from publicly available research papers leaderboard scores during this phase are not meaningful.\n\n\n>2.A forecasting phase will add a private leaderboard test set to be collected after submissions close. Expect the additional images to roughly double the size of the test set.\n\nAlso, the host cannot confirm whether the private test set will contain publicly available datasets or data with the same distribution as the training set.\n\nTable (examples of runs / models)\n\n| CV     | AuthF1 | ForgedF1 | LB    |\n|-------:|:------:|:--------:|:-----:|\n| 0.4592 | 1.0000 | 0.0000   | 0.303 |\n| 0.7636 | 0.9700 | 0.5800   | 0.223 |\n| 0.7577 | 0.9663 | 0.5806   | 0.204 |\n| 0.6515 | 0.9558 | 0.3930   | 0.186 |\n| 0.6691 | 0.9651 | 0.4178   | 0.242 |\n\nI've seen many people do well on the public leaderboard — do your CV/OOF scores match your public LB scores?"
  }
}