{
  "id": 665682,
  "title": "Test Set Size & Private Leaderboard Time Limit?",
  "url": "/competitions/vesuvius-challenge-surface-detection/discussion/665682",
  "author_name": "",
  "post_date": "2026-01-03T05:26:54.860170200Z",
  "votes": 2,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Hi organizers and fellow competitors,</p>\n<p>I have two questions regarding the test set and submission evaluation:</p>\n<h3>1. Test Set Size</h3>\n<p>How many samples are in the test set? Specifically:</p>\n<ul>\n<li>Public leaderboard test set size?</li>\n<li>Private leaderboard test set size?</li>\n</ul>\n<p>This information would help us better plan our inference pipeline</p>\n<h3>2. Private Leaderboard Time Limit</h3>\n<p>The current notebook runtime limit is <strong>9 hours</strong> for GPU/CPU notebooks.\nWill the same 9-hour limit apply during <strong>private leaderboard evaluation</strong>?</p>\n<p>This is important for planning inference strategies - if the private set is significantly larger, we need to ensure our solution scales within the time constraint.</p>",
  "messages": [
    {
      "id": "3385354",
      "postDate": "01/03/2026 05:26:54",
      "content": "<p>Hi organizers and fellow competitors,</p>\n<p>I have two questions regarding the test set and submission evaluation:</p>\n<h3>1. Test Set Size</h3>\n<p>How many samples are in the test set? Specifically:</p>\n<ul>\n<li>Public leaderboard test set size?</li>\n<li>Private leaderboard test set size?</li>\n</ul>\n<p>This information would help us better plan our inference pipeline</p>\n<h3>2. Private Leaderboard Time Limit</h3>\n<p>The current notebook runtime limit is <strong>9 hours</strong> for GPU/CPU notebooks.\nWill the same 9-hour limit apply during <strong>private leaderboard evaluation</strong>?</p>\n<p>This is important for planning inference strategies - if the private set is significantly larger, we need to ensure our solution scales within the time constraint.</p>",
      "rawMarkdown": "Hi organizers and fellow competitors,\n\nI have two questions regarding the test set and submission evaluation:\n\n### 1. Test Set Size\nHow many samples are in the test set? Specifically:\n- Public leaderboard test set size?\n- Private leaderboard test set size?\n\nThis information would help us better plan our inference pipeline\n\n### 2. Private Leaderboard Time Limit\nThe current notebook runtime limit is **9 hours** for GPU/CPU notebooks.\nWill the same 9-hour limit apply during **private leaderboard evaluation**?\n\nThis is important for planning inference strategies - if the private set is significantly larger, we need to ensure our solution scales within the time constraint.",
      "votes": null
    },
    {
      "id": "3385365",
      "postDate": "01/03/2026 06:02:50",
      "content": "<p>I Too have the second question, because with my current pipeline 5x the data would tle on 9 hour limit.</p>",
      "rawMarkdown": "I Too have the second question, because with my current pipeline 5x the data would tle on 9 hour limit.",
      "votes": null
    },
    {
      "id": "3385374",
      "postDate": "01/03/2026 06:27:43",
      "content": "<p>The notebook runs against the public + private dataset, which means that if you have got public score then it is ok.</p>",
      "rawMarkdown": "The notebook runs against the public + private dataset, which means that if you have got public score then it is ok.",
      "votes": null
    },
    {
      "id": "3385414",
      "postDate": "01/03/2026 08:58:30",
      "content": "<p>5-fold ensemble with TTA will time out.</p>",
      "rawMarkdown": "5-fold ensemble with TTA will time out.",
      "votes": null
    },
    {
      "id": "3385458",
      "postDate": "01/03/2026 10:54:32",
      "content": "<p>Even if it didn't time out, it would leave no time for post processing (in case we come up with some good post processing procedure).</p>",
      "rawMarkdown": "Even if it didn't time out, it would leave no time for post processing (in case we come up with some good post processing procedure).",
      "votes": null
    },
    {
      "id": "3385668",
      "postDate": "01/03/2026 17:18:37",
      "content": "<p>It is possible that the competition organizers are looking for methods that do not rely on post-processing.</p>",
      "rawMarkdown": "It is possible that the competition organizers are looking for methods that do not rely on post-processing.",
      "votes": null
    },
    {
      "id": "3391878",
      "postDate": "01/15/2026 21:21:15",
      "content": "<p>I am skeptical on the 5-fold ensemble, but I have been wondering about a two model ensemble (though not from two folds).</p>",
      "rawMarkdown": "I am skeptical on the 5-fold ensemble, but I have been wondering about a two model ensemble (though not from two folds).",
      "votes": null
    },
    {
      "id": "3394962",
      "postDate": "01/22/2026 01:19:53",
      "content": "<p>Why hasn't my notebook timed out after 12 hours? Isn't the limit supposed to be 9 hours?</p>",
      "rawMarkdown": "Why hasn't my notebook timed out after 12 hours? Isn't the limit supposed to be 9 hours?",
      "votes": null
    },
    {
      "id": "3394980",
      "postDate": "01/22/2026 03:28:16",
      "content": "<p>My understanding is that an extra 3–4 hours are needed for scoring; the topology score calculation is very time‑consuming</p>",
      "rawMarkdown": "My understanding is that an extra 3–4 hours are needed for scoring; the topology score calculation is very time‑consuming",
      "votes": null
    },
    {
      "id": "3394997",
      "postDate": "01/22/2026 04:09:38",
      "content": "<p>So this 12 hours notebook of mine is also feasible right?</p>",
      "rawMarkdown": "So this 12 hours notebook of mine is also feasible right?",
      "votes": null
    },
    {
      "id": "3395037",
      "postDate": "01/22/2026 06:00:00",
      "content": "<p>Inference with the P100 seems to get stuck or even impossible, while the T4 doesn't? Is it a problem with my settings? Are you all using the T4 for inference?</p>",
      "rawMarkdown": "Inference with the P100 seems to get stuck or even impossible, while the T4 doesn't? Is it a problem with my settings? Are you all using the T4 for inference?",
      "votes": null
    },
    {
      "id": "3395066",
      "postDate": "01/22/2026 07:22:59",
      "content": "<p>I'm not sure, but according to what <a href=\"https://www.kaggle.com/yuanzhe\" target=\"_blank\">@yuanzhe</a> zhou said, \"The notebook runs against the public + private dataset, which means that if you have got public score then it is ok.\" As long as you have a normal score in the public phase, the private phase will not run inference again; it only does scoring.</p>",
      "rawMarkdown": "I'm not sure, but according to what @yuanzhe zhou said, \"The notebook runs against the public + private dataset, which means that if you have got public score then it is ok.\" As long as you have a normal score in the public phase, the private phase will not run inference again; it only does scoring.",
      "votes": null
    },
    {
      "id": "3395067",
      "postDate": "01/22/2026 07:23:32",
      "content": "<p>n theory, the inference performance of 2× T4 is stronger than P100</p>",
      "rawMarkdown": "n theory, the inference performance of 2× T4 is stronger than P100",
      "votes": null
    },
    {
      "id": "3395077",
      "postDate": "01/22/2026 07:51:27",
      "content": "<p>Thank you. I got it.</p>",
      "rawMarkdown": "Thank you. I got it.",
      "votes": null
    },
    {
      "id": "3395079",
      "postDate": "01/22/2026 07:53:14",
      "content": "<p>👍I got it.</p>",
      "rawMarkdown": "👍I got it.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3385365,
      "author_name": "choudharymanas",
      "author_url": "",
      "post_date": "01/03/2026 06:02:50",
      "content": "<p>I Too have the second question, because with my current pipeline 5x the data would tle on 9 hour limit.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3385414,
          "author_name": "chengtingyi",
          "author_url": "",
          "post_date": "01/03/2026 08:58:30",
          "content": "<p>5-fold ensemble with TTA will time out.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3385458,
              "author_name": "choudharymanas",
              "author_url": "",
              "post_date": "01/03/2026 10:54:32",
              "content": "<p>Even if it didn't time out, it would leave no time for post processing (in case we come up with some good post processing procedure).</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3385668,
                  "author_name": "chengtingyi",
                  "author_url": "",
                  "post_date": "01/03/2026 17:18:37",
                  "content": "<p>It is possible that the competition organizers are looking for methods that do not rely on post-processing.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            },
            {
              "id": 3395037,
              "author_name": "tanyong666",
              "author_url": "",
              "post_date": "01/22/2026 06:00:00",
              "content": "<p>Inference with the P100 seems to get stuck or even impossible, while the T4 doesn't? Is it a problem with my settings? Are you all using the T4 for inference?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3395067,
                  "author_name": "wzyfromhust",
                  "author_url": "",
                  "post_date": "01/22/2026 07:23:32",
                  "content": "<p>n theory, the inference performance of 2× T4 is stronger than P100</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3395079,
                      "author_name": "tanyong666",
                      "author_url": "",
                      "post_date": "01/22/2026 07:53:14",
                      "content": "<p>👍I got it.</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        },
        {
          "id": 3391878,
          "author_name": "rob1080ti",
          "author_url": "",
          "post_date": "01/15/2026 21:21:15",
          "content": "<p>I am skeptical on the 5-fold ensemble, but I have been wondering about a two model ensemble (though not from two folds).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3385374,
      "author_name": "yuanzhezhou",
      "author_url": "",
      "post_date": "01/03/2026 06:27:43",
      "content": "<p>The notebook runs against the public + private dataset, which means that if you have got public score then it is ok.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3394962,
      "author_name": "tanyong666",
      "author_url": "",
      "post_date": "01/22/2026 01:19:53",
      "content": "<p>Why hasn't my notebook timed out after 12 hours? Isn't the limit supposed to be 9 hours?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3394980,
          "author_name": "wzyfromhust",
          "author_url": "",
          "post_date": "01/22/2026 03:28:16",
          "content": "<p>My understanding is that an extra 3–4 hours are needed for scoring; the topology score calculation is very time‑consuming</p>",
          "votes": null,
          "replies": [
            {
              "id": 3394997,
              "author_name": "tanyong666",
              "author_url": "",
              "post_date": "01/22/2026 04:09:38",
              "content": "<p>So this 12 hours notebook of mine is also feasible right?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3395066,
                  "author_name": "wzyfromhust",
                  "author_url": "",
                  "post_date": "01/22/2026 07:22:59",
                  "content": "<p>I'm not sure, but according to what <a href=\"https://www.kaggle.com/yuanzhe\" target=\"_blank\">@yuanzhe</a> zhou said, \"The notebook runs against the public + private dataset, which means that if you have got public score then it is ok.\" As long as you have a normal score in the public phase, the private phase will not run inference again; it only does scoring.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3395077,
                      "author_name": "tanyong666",
                      "author_url": "",
                      "post_date": "01/22/2026 07:51:27",
                      "content": "<p>Thank you. I got it.</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3385354": "Hi organizers and fellow competitors,\n\nI have two questions regarding the test set and submission evaluation:\n\n### 1. Test Set Size\nHow many samples are in the test set? Specifically:\n- Public leaderboard test set size?\n- Private leaderboard test set size?\n\nThis information would help us better plan our inference pipeline\n\n### 2. Private Leaderboard Time Limit\nThe current notebook runtime limit is **9 hours** for GPU/CPU notebooks.\nWill the same 9-hour limit apply during **private leaderboard evaluation**?\n\nThis is important for planning inference strategies - if the private set is significantly larger, we need to ensure our solution scales within the time constraint.",
    "3385365": "I Too have the second question, because with my current pipeline 5x the data would tle on 9 hour limit.",
    "3385374": "The notebook runs against the public + private dataset, which means that if you have got public score then it is ok.",
    "3385414": "5-fold ensemble with TTA will time out.",
    "3385458": "Even if it didn't time out, it would leave no time for post processing (in case we come up with some good post processing procedure).",
    "3385668": "It is possible that the competition organizers are looking for methods that do not rely on post-processing.",
    "3391878": "I am skeptical on the 5-fold ensemble, but I have been wondering about a two model ensemble (though not from two folds).",
    "3394962": "Why hasn't my notebook timed out after 12 hours? Isn't the limit supposed to be 9 hours?",
    "3394980": "My understanding is that an extra 3–4 hours are needed for scoring; the topology score calculation is very time‑consuming",
    "3394997": "So this 12 hours notebook of mine is also feasible right?",
    "3395037": "Inference with the P100 seems to get stuck or even impossible, while the T4 doesn't? Is it a problem with my settings? Are you all using the T4 for inference?",
    "3395066": "I'm not sure, but according to what @yuanzhe zhou said, \"The notebook runs against the public + private dataset, which means that if you have got public score then it is ok.\" As long as you have a normal score in the public phase, the private phase will not run inference again; it only does scoring.",
    "3395067": "n theory, the inference performance of 2× T4 is stronger than P100",
    "3395077": "Thank you. I got it.",
    "3395079": "👍I got it."
  },
  "source": "meta"
}