{
  "id": 679224,
  "title": "Probing Results (Train and Test Scroll IDs)",
  "url": "/competitions/vesuvius-challenge-surface-detection/discussion/679224",
  "author_name": "yu4u",
  "post_date": "2026-02-28T01:43:35.776000",
  "votes": 4,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I probed the test dataset and found the following:</p>\n<hr>\n<h3>Scroll IDs Present in Test</h3>\n<table>\n<thead>\n<tr>\n<th>Scroll ID</th>\n<th>Source</th>\n<th># Samples</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>34117</td>\n<td>Train</td>\n<td>42</td>\n</tr>\n<tr>\n<td>35360</td>\n<td>Train</td>\n<td>19</td>\n</tr>\n<tr>\n<td>44430</td>\n<td>Train</td>\n<td>4</td>\n</tr>\n<tr>\n<td>26010</td>\n<td>Train</td>\n<td>13</td>\n</tr>\n<tr>\n<td>test_1</td>\n<td>Test-only</td>\n<td>30</td>\n</tr>\n<tr>\n<td>test_2</td>\n<td>Test-only</td>\n<td>2</td>\n</tr>\n<tr>\n<td><strong>Total</strong></td>\n<td></td>\n<td><strong>110</strong></td>\n</tr>\n</tbody>\n</table>\n<hr>\n<h3>Scroll IDs <strong>NOT</strong> Present in Test</h3>\n<ul>\n<li>26002</li>\n<li>53997</li>\n</ul>\n<hr>\n<h3>Summary</h3>\n<ul>\n<li><strong>Total test samples</strong>: 110</li>\n<li><strong>Total unique scroll IDs in test</strong>: 6<ul>\n<li>4 scroll IDs overlap with training data (34117, 35360, 44430, 26010)</li>\n<li>2 scroll IDs are test-only and not seen during training (test_1, test_2)</li></ul></li>\n<li><strong>78 out of 110 test samples (≈71%) come from scroll IDs also present in the training data</strong>, meaning a large portion of the test set shares scroll IDs with the training set. This suggests that models may benefit significantly from in-distribution generalization rather than purely out-of-distribution generalization.</li>\n</ul>",
  "messages": [
    {
      "id": 3414948,
      "postDate": "2026-02-28T01:43:35.777Z",
      "content": "<p>I probed the test dataset and found the following:</p>\n<hr>\n<h3>Scroll IDs Present in Test</h3>\n<table>\n<thead>\n<tr>\n<th>Scroll ID</th>\n<th>Source</th>\n<th># Samples</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>34117</td>\n<td>Train</td>\n<td>42</td>\n</tr>\n<tr>\n<td>35360</td>\n<td>Train</td>\n<td>19</td>\n</tr>\n<tr>\n<td>44430</td>\n<td>Train</td>\n<td>4</td>\n</tr>\n<tr>\n<td>26010</td>\n<td>Train</td>\n<td>13</td>\n</tr>\n<tr>\n<td>test_1</td>\n<td>Test-only</td>\n<td>30</td>\n</tr>\n<tr>\n<td>test_2</td>\n<td>Test-only</td>\n<td>2</td>\n</tr>\n<tr>\n<td><strong>Total</strong></td>\n<td></td>\n<td><strong>110</strong></td>\n</tr>\n</tbody>\n</table>\n<hr>\n<h3>Scroll IDs <strong>NOT</strong> Present in Test</h3>\n<ul>\n<li>26002</li>\n<li>53997</li>\n</ul>\n<hr>\n<h3>Summary</h3>\n<ul>\n<li><strong>Total test samples</strong>: 110</li>\n<li><strong>Total unique scroll IDs in test</strong>: 6<ul>\n<li>4 scroll IDs overlap with training data (34117, 35360, 44430, 26010)</li>\n<li>2 scroll IDs are test-only and not seen during training (test_1, test_2)</li></ul></li>\n<li><strong>78 out of 110 test samples (≈71%) come from scroll IDs also present in the training data</strong>, meaning a large portion of the test set shares scroll IDs with the training set. This suggests that models may benefit significantly from in-distribution generalization rather than purely out-of-distribution generalization.</li>\n</ul>",
      "rawMarkdown": "I probed the test dataset and found the following:\n\n---\n\n### Scroll IDs Present in Test\n\n| Scroll ID | Source | # Samples |\n|-----------|--------|-----------|\n| 34117 | Train | 42 |\n| 35360 | Train | 19 |\n| 44430 | Train | 4 |\n| 26010 | Train | 13 |\n| test_1 | Test-only | 30 |\n| test_2 | Test-only | 2 |\n| **Total** | | **110** |\n\n---\n\n### Scroll IDs **NOT** Present in Test\n\n- 26002\n- 53997\n\n---\n\n### Summary\n\n- **Total test samples**: 110\n- **Total unique scroll IDs in test**: 6\n  - 4 scroll IDs overlap with training data (34117, 35360, 44430, 26010)\n  - 2 scroll IDs are test-only and not seen during training (test_1, test_2)\n- **78 out of 110 test samples (≈71%) come from scroll IDs also present in the training data**, meaning a large portion of the test set shares scroll IDs with the training set. This suggests that models may benefit significantly from in-distribution generalization rather than purely out-of-distribution generalization.",
      "votes": 4
    },
    {
      "id": 3414970,
      "postDate": "2026-02-28T02:37:45.003Z",
      "content": "<p>Interesting — how were you able to infer the test scroll IDs and sample counts? Kaggle submissions usually only report an overall score, so I’m curious about the methodology. If possible, could you share the approach, even at a high level? I’d really appreciate it. Thanks.</p>",
      "rawMarkdown": "Interesting — how were you able to infer the test scroll IDs and sample counts? Kaggle submissions usually only report an overall score, so I’m curious about the methodology. If possible, could you share the approach, even at a high level? I’d really appreciate it. Thanks.",
      "votes": -1,
      "replies": [
        {
          "id": 3414973,
          "postDate": "2026-02-28T02:42:02.043Z",
          "content": "<p>if \"bla bla bla\":</p>\n<p>raise error</p>",
          "rawMarkdown": "if \"bla bla bla\":\n\nraise error\n\n",
          "votes": 1,
          "replies": [
            {
              "id": 3414983,
              "postDate": "2026-02-28T02:52:57.827Z",
              "content": "<p>smarter way:</p>\n<pre><code>if ...:\n  b1 = ...\nif ...:\n  b2 = ...\nif ...:\n  b3 = ...\n\n\nbinary_encode = b1,b2,b3\n\nmatch binary_encode:\n    case 1:\n        submit with score = 0.1\n    case 2:\n        submit with score = 0.2\n    case ..._:\n        submit with score = 0.3\n</code></pre>\n<p>use sleep to control the timing of submission</p>\n<pre><code>use code to control gpu utilisation .... the frequency of generated heat change  then can be measured by very sensitive instrucmen: aka butterfly effect\n</code></pre>",
              "rawMarkdown": "smarter way:\n```\nif ...:\n  b1 = ...\nif ...:\n  b2 = ...\nif ...:\n  b3 = ...\n\n\nbinary_encode = b1,b2,b3\n\nmatch binary_encode:\n    case 1:\n        submit with score = 0.1\n    case 2:\n        submit with score = 0.2\n    case ..._:\n        submit with score = 0.3\n\n```\n\nuse sleep to control the timing of submission\n\n```\nuse code to control gpu utilisation .... the frequency of generated heat change  then can be measured by very sensitive instrucmen: aka butterfly effect\n\n```",
              "votes": 2
            }
          ]
        },
        {
          "id": 3415056,
          "postDate": "2026-02-28T05:44:22.423Z",
          "content": "<p>Prepare multiple submissions whose scores are already known. In my case, I created several versions by changing the post-processing threshold th that sets the region where the predicted z-coordinate exceeds th to 0.\nSwitch between submissions (i.e., different th values) depending on the information you want to probe.\nIf you prepare N submissions, you can obtain log₂ N bits of information from a single probing attempt.</p>",
          "rawMarkdown": "Prepare multiple submissions whose scores are already known. In my case, I created several versions by changing the post-processing threshold th that sets the region where the predicted z-coordinate exceeds th to 0.\nSwitch between submissions (i.e., different th values) depending on the information you want to probe.\nIf you prepare N submissions, you can obtain log₂ N bits of information from a single probing attempt.",
          "votes": 1
        }
      ]
    },
    {
      "id": 3415102,
      "postDate": "2026-02-28T07:06:08.400Z",
      "content": "<p>What about the dataset? <a href=\"https://www.kaggle.com/ren4yu\" target=\"_blank\">@ren4yu</a> </p>",
      "rawMarkdown": "What about the dataset? @ren4yu "
    },
    {
      "id": 3414974,
      "postDate": "2026-02-28T02:42:41.713Z",
      "content": "<p>How may one obtain the test dataset?</p>",
      "rawMarkdown": "How may one obtain the test dataset?",
      "replies": [
        {
          "id": 3415506,
          "postDate": "2026-03-01T00:57:14.657Z",
          "content": "<p>Only the submitted notebook can access the hidden test data.\nTherefore, by changing the processing inside the submitted notebook according to the information of the hidden test data and observing the resulting submission score, you can extract information about the hidden test set.</p>",
          "rawMarkdown": "Only the submitted notebook can access the hidden test data.\nTherefore, by changing the processing inside the submitted notebook according to the information of the hidden test data and observing the resulting submission score, you can extract information about the hidden test set.",
          "votes": 1
        }
      ]
    },
    {
      "id": 3415683,
      "postDate": "2026-03-01T06:19:42.920Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3414970,
      "author_name": "Tony Li",
      "author_url": "",
      "post_date": "2026-02-28T02:37:45.003000",
      "content": "<p>Interesting — how were you able to infer the test scroll IDs and sample counts? Kaggle submissions usually only report an overall score, so I’m curious about the methodology. If possible, could you share the approach, even at a high level? I’d really appreciate it. Thanks.</p>",
      "votes": -1,
      "replies": [
        {
          "id": 3414973,
          "author_name": "Taha_Alshatiri",
          "author_url": "",
          "post_date": "2026-02-28T02:42:02.043000",
          "content": "<p>if \"bla bla bla\":</p>\n<p>raise error</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3414983,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2026-02-28T02:52:57.827000",
              "content": "<p>smarter way:</p>\n<pre><code>if ...:\n  b1 = ...\nif ...:\n  b2 = ...\nif ...:\n  b3 = ...\n\n\nbinary_encode = b1,b2,b3\n\nmatch binary_encode:\n    case 1:\n        submit with score = 0.1\n    case 2:\n        submit with score = 0.2\n    case ..._:\n        submit with score = 0.3\n</code></pre>\n<p>use sleep to control the timing of submission</p>\n<pre><code>use code to control gpu utilisation .... the frequency of generated heat change  then can be measured by very sensitive instrucmen: aka butterfly effect\n</code></pre>",
              "votes": 2,
              "replies": []
            }
          ]
        },
        {
          "id": 3415056,
          "author_name": "yu4u",
          "author_url": "",
          "post_date": "2026-02-28T05:44:22.423000",
          "content": "<p>Prepare multiple submissions whose scores are already known. In my case, I created several versions by changing the post-processing threshold th that sets the region where the predicted z-coordinate exceeds th to 0.\nSwitch between submissions (i.e., different th values) depending on the information you want to probe.\nIf you prepare N submissions, you can obtain log₂ N bits of information from a single probing attempt.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3415102,
      "author_name": "Navneet",
      "author_url": "",
      "post_date": "2026-02-28T07:06:08.400000",
      "content": "<p>What about the dataset? <a href=\"https://www.kaggle.com/ren4yu\" target=\"_blank\">@ren4yu</a> </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3414974,
      "author_name": "huoxu",
      "author_url": "",
      "post_date": "2026-02-28T02:42:41.713000",
      "content": "<p>How may one obtain the test dataset?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3415506,
          "author_name": "yu4u",
          "author_url": "",
          "post_date": "2026-03-01T00:57:14.657000",
          "content": "<p>Only the submitted notebook can access the hidden test data.\nTherefore, by changing the processing inside the submitted notebook according to the information of the hidden test data and observing the resulting submission score, you can extract information about the hidden test set.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3415683,
      "author_name": "",
      "author_url": "",
      "post_date": "2026-03-01T06:19:42.920000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3414948": "I probed the test dataset and found the following:\n\n---\n\n### Scroll IDs Present in Test\n\n| Scroll ID | Source | # Samples |\n|-----------|--------|-----------|\n| 34117 | Train | 42 |\n| 35360 | Train | 19 |\n| 44430 | Train | 4 |\n| 26010 | Train | 13 |\n| test_1 | Test-only | 30 |\n| test_2 | Test-only | 2 |\n| **Total** | | **110** |\n\n---\n\n### Scroll IDs **NOT** Present in Test\n\n- 26002\n- 53997\n\n---\n\n### Summary\n\n- **Total test samples**: 110\n- **Total unique scroll IDs in test**: 6\n  - 4 scroll IDs overlap with training data (34117, 35360, 44430, 26010)\n  - 2 scroll IDs are test-only and not seen during training (test_1, test_2)\n- **78 out of 110 test samples (≈71%) come from scroll IDs also present in the training data**, meaning a large portion of the test set shares scroll IDs with the training set. This suggests that models may benefit significantly from in-distribution generalization rather than purely out-of-distribution generalization.",
    "3414970": "Interesting — how were you able to infer the test scroll IDs and sample counts? Kaggle submissions usually only report an overall score, so I’m curious about the methodology. If possible, could you share the approach, even at a high level? I’d really appreciate it. Thanks.",
    "3415102": "What about the dataset? @ren4yu ",
    "3414974": "How may one obtain the test dataset?",
    "3415683": ""
  }
}