{
  "id": 669787,
  "title": "Too Much time on Submission",
  "url": "/competitions/vesuvius-challenge-surface-detection/discussion/669787",
  "author_name": "Kaustuk000",
  "post_date": "2026-01-24T10:37:54.117000",
  "votes": 2,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I used GPU P100 and submitted but its been more then 1 hour and submission isn't completed yet. I do know that we are expected to have around 120 volumes but if i am not wrong for the training they gave us 700+ volumes and on training my 1 epoch took arounf 16 min So, by that it shouldn't take this much time. Can Some one explain what's can be the problem as this is my first time entering in the featured competition.</p>",
  "messages": [
    {
      "id": 3396107,
      "postDate": "2026-01-24T10:37:54.117Z",
      "content": "<p>I used GPU P100 and submitted but its been more then 1 hour and submission isn't completed yet. I do know that we are expected to have around 120 volumes but if i am not wrong for the training they gave us 700+ volumes and on training my 1 epoch took arounf 16 min So, by that it shouldn't take this much time. Can Some one explain what's can be the problem as this is my first time entering in the featured competition.</p>",
      "rawMarkdown": "I used GPU P100 and submitted but its been more then 1 hour and submission isn't completed yet. I do know that we are expected to have around 120 volumes but if i am not wrong for the training they gave us 700+ volumes and on training my 1 epoch took arounf 16 min So, by that it shouldn't take this much time. Can Some one explain what's can be the problem as this is my first time entering in the featured competition.",
      "votes": 2
    },
    {
      "id": 3396121,
      "postDate": "2026-01-24T11:28:33.033Z",
      "content": "<p>Metric calculation takes about four hours.\n<a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/621455\" target=\"_blank\">https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/621455</a></p>",
      "rawMarkdown": "Metric calculation takes about four hours.\nhttps://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/621455",
      "votes": 1,
      "replies": [
        {
          "id": 3396135,
          "postDate": "2026-01-24T12:07:46Z",
          "content": "<p>thankyou for the details its helps a lot. if you don't mind can you clear a few more doubts of mine as well.</p>\n<ol>\n<li>In Completition details it's says like this CPU Notebook &lt;= 9 hours run-time\nGPU Notebook &lt;= 9 hours run-time. This is for the inference notebook we submit or its apply for training notebook as well. Since model training takes times more then that. and also the competition has hidden test set so i am little confused as to my submission is correct or not . \n I done this, like turn off the training loop when i was saving the training notebook for version.\nand for inference notebook \n I load the model weigths and used some necessary required preprocessing. \n and \n and made submission file using this\n<br>\ntest_ids = sorted([p.stem for p in TEST_IMG_DIR.glob(\"*.tif\")])<br>\nprint(\"Test fragments:\", len(test_ids))<br>\nfor fid in tqdm(test_ids):<br>\nvol = tiff.imread(TEST_IMG_DIR / f\"{fid}.tif\")  # (D,H,W)<br>\nink_prob = infer_volume(vol)  # (D,H,W) float [0,1]<br>\npred = np.zeros_like(ink_prob, dtype=np.uint8)<br>\npred[ink_prob &gt; 0.5] = 1   # ink<br>\ntiff.imwrite(<br>\n    OUT_DIR / f\"{fid}.tif\",<br>\n    pred,<br>\n    compression=\"zlib\"<br>\n)<br>\n<br></li>\n</ol>\n<p>BASE_DIR = Path(\"/kaggle/input/vesuvius-challenge-surface-detection\")<br>\nTEST_IMG_DIR = BASE_DIR / \"test_images\"<br>\nWORKDIR = Path(\"/kaggle/working\")<br>\nOUT_DIR = WORKDIR / \"predictions\"<br>\nOUT_DIR.mkdir(parents=True, exist_ok=True)<br>\nit s\nhow len = 1 and run inference on that which is i think is correct considering originally the folder had only 1 test set image.\nSo, if can and just verify that i am doing things correct and my model is running on there hidden set. IT's just confusing me.</p>",
          "rawMarkdown": "thankyou for the details its helps a lot. if you don't mind can you clear a few more doubts of mine as well.\n1. In Completition details it's says like this CPU Notebook <= 9 hours run-time\nGPU Notebook <= 9 hours run-time. This is for the inference notebook we submit or its apply for training notebook as well. Since model training takes times more then that. and also the competition has hidden test set so i am little confused as to my submission is correct or not . \n     I done this, like turn off the training loop when i was saving the training notebook for version.\nand for inference notebook \n     I load the model weigths and used some necessary required preprocessing. \n     and \n     and made submission file using this\n    <br>\ntest_ids = sorted([p.stem for p in TEST_IMG_DIR.glob(\"*.tif\")])<br>\nprint(\"Test fragments:\", len(test_ids))<br>\nfor fid in tqdm(test_ids):<br>\n    vol = tiff.imread(TEST_IMG_DIR / f\"{fid}.tif\")  # (D,H,W)<br>\n    ink_prob = infer_volume(vol)  # (D,H,W) float [0,1]<br>\n    pred = np.zeros_like(ink_prob, dtype=np.uint8)<br>\n    pred[ink_prob > 0.5] = 1   # ink<br>\n    tiff.imwrite(<br>\n        OUT_DIR / f\"{fid}.tif\",<br>\n        pred,<br>\n        compression=\"zlib\"<br>\n    )<br>\n<br>\n\nBASE_DIR = Path(\"/kaggle/input/vesuvius-challenge-surface-detection\")<br>\nTEST_IMG_DIR = BASE_DIR / \"test_images\"<br>\nWORKDIR = Path(\"/kaggle/working\")<br>\nOUT_DIR = WORKDIR / \"predictions\"<br>\nOUT_DIR.mkdir(parents=True, exist_ok=True)<br>\nit s\nhow len = 1 and run inference on that which is i think is correct considering originally the folder had only 1 test set image.\nSo, if can and just verify that i am doing things correct and my model is running on there hidden set. IT's just confusing me.",
          "replies": [
            {
              "id": 3396233,
              "postDate": "2026-01-24T15:10:14.627Z",
              "content": "<p>You are doing it right. The test folder will contain the full test set (public + private) when Kaggle reruns your notebook (you will see it run twice)</p>",
              "rawMarkdown": "You are doing it right. The test folder will contain the full test set (public + private) when Kaggle reruns your notebook (you will see it run twice)",
              "votes": 1
            }
          ]
        },
        {
          "id": 3400033,
          "postDate": "2026-01-31T18:43:27.890Z",
          "content": "<p>I tried a simple inference pipeline and submission. Inference notebook finished within 5-7 minutes. But, metric calculation ran more than 9h and exited. </p>\n<p>Couldn't see what is error as well / why submission failed when notebook was successful.</p>",
          "rawMarkdown": "I tried a simple inference pipeline and submission. Inference notebook finished within 5-7 minutes. But, metric calculation ran more than 9h and exited. \n\nCouldn't see what is error as well / why submission failed when notebook was successful.",
          "replies": [
            {
              "id": 3400057,
              "postDate": "2026-01-31T20:41:27.043Z",
              "content": "<p>Are you using the native nnU-Net pipeline with mirroring enabled? Double check you’re not also stacking extra TTA on top of that (extra flips/aug loops), because that can blow up runtime fast. Also watch what models you’re running — fold-all models (or lots of folds) will take way longer than split-fold setups (ex: 5 split folds vs just 2 fold-all models, etc.).</p>\n<p>My first few submissions timed out too — I had mirroring + extra TTA going and it was just too slow. Once I disabled the extra stuff and dialed it in, my local/in-notebook inference finishes fast. For example my log is like:</p>\n<p>225.8s ✅ Saved mask: predictions/1407735.tif\nand then it zips and writes submission.zip right after.</p>\n<p>The hidden test set run still takes a while — for me it’s a little over 8 hours to finish.</p>\n<p>Also, if you’re running 2 fold-all models, you’ll pretty much need the T4 setup and run them in parallel — on a P100 it’s likely to time out if you run them sequentially.</p>",
              "rawMarkdown": "Are you using the native nnU-Net pipeline with mirroring enabled? Double check you’re not also stacking extra TTA on top of that (extra flips/aug loops), because that can blow up runtime fast. Also watch what models you’re running — fold-all models (or lots of folds) will take way longer than split-fold setups (ex: 5 split folds vs just 2 fold-all models, etc.).\n\nMy first few submissions timed out too — I had mirroring + extra TTA going and it was just too slow. Once I disabled the extra stuff and dialed it in, my local/in-notebook inference finishes fast. For example my log is like:\n\n225.8s ✅ Saved mask: predictions/1407735.tif\nand then it zips and writes submission.zip right after.\n\nThe hidden test set run still takes a while — for me it’s a little over 8 hours to finish.\n\nAlso, if you’re running 2 fold-all models, you’ll pretty much need the T4 setup and run them in parallel — on a P100 it’s likely to time out if you run them sequentially."
            }
          ]
        }
      ]
    },
    {
      "id": 3399184,
      "postDate": "2026-01-30T11:01:04.293Z",
      "content": "<p>Do you know which GPU will be available for the testing phase?</p>",
      "rawMarkdown": "Do you know which GPU will be available for the testing phase?\n",
      "replies": [
        {
          "id": 3399190,
          "postDate": "2026-01-30T11:10:39.687Z",
          "content": "<p>For mine i was able to select it when i saved for the version. I think GPU saved on that time will be used.</p>",
          "rawMarkdown": "For mine i was able to select it when i saved for the version. I think GPU saved on that time will be used.",
          "replies": [
            {
              "id": 3399364,
              "postDate": "2026-01-30T18:09:10.930Z",
              "content": "<p>Thanks! I'm a bit confused how this is supposed to work. I have created the notebook that runs inferences to all .tif files in the input. Do you know if this is correct? I just noticed that it's doing inference for the training cases: \n<code>Doing case: /kaggle/input/vesuvius-challenge-surface-detection/train_images/2761996776.tif</code></p>",
              "rawMarkdown": "Thanks! I'm a bit confused how this is supposed to work. I have created the notebook that runs inferences to all .tif files in the input. Do you know if this is correct? I just noticed that it's doing inference for the training cases: \n`Doing case: /kaggle/input/vesuvius-challenge-surface-detection/train_images/2761996776.tif`"
            }
          ]
        }
      ]
    },
    {
      "id": 3396244,
      "postDate": "2026-01-24T15:29:22.750Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3396121,
      "author_name": "Bull",
      "author_url": "",
      "post_date": "2026-01-24T11:28:33.033000",
      "content": "<p>Metric calculation takes about four hours.\n<a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/621455\" target=\"_blank\">https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/621455</a></p>",
      "votes": 1,
      "replies": [
        {
          "id": 3396135,
          "author_name": "Kaustuk000",
          "author_url": "",
          "post_date": "2026-01-24T12:07:46",
          "content": "<p>thankyou for the details its helps a lot. if you don't mind can you clear a few more doubts of mine as well.</p>\n<ol>\n<li>In Completition details it's says like this CPU Notebook &lt;= 9 hours run-time\nGPU Notebook &lt;= 9 hours run-time. This is for the inference notebook we submit or its apply for training notebook as well. Since model training takes times more then that. and also the competition has hidden test set so i am little confused as to my submission is correct or not . \n I done this, like turn off the training loop when i was saving the training notebook for version.\nand for inference notebook \n I load the model weigths and used some necessary required preprocessing. \n and \n and made submission file using this\n<br>\ntest_ids = sorted([p.stem for p in TEST_IMG_DIR.glob(\"*.tif\")])<br>\nprint(\"Test fragments:\", len(test_ids))<br>\nfor fid in tqdm(test_ids):<br>\nvol = tiff.imread(TEST_IMG_DIR / f\"{fid}.tif\")  # (D,H,W)<br>\nink_prob = infer_volume(vol)  # (D,H,W) float [0,1]<br>\npred = np.zeros_like(ink_prob, dtype=np.uint8)<br>\npred[ink_prob &gt; 0.5] = 1   # ink<br>\ntiff.imwrite(<br>\n    OUT_DIR / f\"{fid}.tif\",<br>\n    pred,<br>\n    compression=\"zlib\"<br>\n)<br>\n<br></li>\n</ol>\n<p>BASE_DIR = Path(\"/kaggle/input/vesuvius-challenge-surface-detection\")<br>\nTEST_IMG_DIR = BASE_DIR / \"test_images\"<br>\nWORKDIR = Path(\"/kaggle/working\")<br>\nOUT_DIR = WORKDIR / \"predictions\"<br>\nOUT_DIR.mkdir(parents=True, exist_ok=True)<br>\nit s\nhow len = 1 and run inference on that which is i think is correct considering originally the folder had only 1 test set image.\nSo, if can and just verify that i am doing things correct and my model is running on there hidden set. IT's just confusing me.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3396233,
              "author_name": "Duong Nguyen",
              "author_url": "",
              "post_date": "2026-01-24T15:10:14.627000",
              "content": "<p>You are doing it right. The test folder will contain the full test set (public + private) when Kaggle reruns your notebook (you will see it run twice)</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 3400033,
          "author_name": "Maapu",
          "author_url": "",
          "post_date": "2026-01-31T18:43:27.890000",
          "content": "<p>I tried a simple inference pipeline and submission. Inference notebook finished within 5-7 minutes. But, metric calculation ran more than 9h and exited. </p>\n<p>Couldn't see what is error as well / why submission failed when notebook was successful.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3400057,
              "author_name": "TWEAK",
              "author_url": "",
              "post_date": "2026-01-31T20:41:27.043000",
              "content": "<p>Are you using the native nnU-Net pipeline with mirroring enabled? Double check you’re not also stacking extra TTA on top of that (extra flips/aug loops), because that can blow up runtime fast. Also watch what models you’re running — fold-all models (or lots of folds) will take way longer than split-fold setups (ex: 5 split folds vs just 2 fold-all models, etc.).</p>\n<p>My first few submissions timed out too — I had mirroring + extra TTA going and it was just too slow. Once I disabled the extra stuff and dialed it in, my local/in-notebook inference finishes fast. For example my log is like:</p>\n<p>225.8s ✅ Saved mask: predictions/1407735.tif\nand then it zips and writes submission.zip right after.</p>\n<p>The hidden test set run still takes a while — for me it’s a little over 8 hours to finish.</p>\n<p>Also, if you’re running 2 fold-all models, you’ll pretty much need the T4 setup and run them in parallel — on a P100 it’s likely to time out if you run them sequentially.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3399184,
      "author_name": "Andre Filipe Ferreira",
      "author_url": "",
      "post_date": "2026-01-30T11:01:04.293000",
      "content": "<p>Do you know which GPU will be available for the testing phase?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3399190,
          "author_name": "Kaustuk000",
          "author_url": "",
          "post_date": "2026-01-30T11:10:39.687000",
          "content": "<p>For mine i was able to select it when i saved for the version. I think GPU saved on that time will be used.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3399364,
              "author_name": "Andre Filipe Ferreira",
              "author_url": "",
              "post_date": "2026-01-30T18:09:10.930000",
              "content": "<p>Thanks! I'm a bit confused how this is supposed to work. I have created the notebook that runs inferences to all .tif files in the input. Do you know if this is correct? I just noticed that it's doing inference for the training cases: \n<code>Doing case: /kaggle/input/vesuvius-challenge-surface-detection/train_images/2761996776.tif</code></p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3396244,
      "author_name": "",
      "author_url": "",
      "post_date": "2026-01-24T15:29:22.750000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3396107": "I used GPU P100 and submitted but its been more then 1 hour and submission isn't completed yet. I do know that we are expected to have around 120 volumes but if i am not wrong for the training they gave us 700+ volumes and on training my 1 epoch took arounf 16 min So, by that it shouldn't take this much time. Can Some one explain what's can be the problem as this is my first time entering in the featured competition.",
    "3396121": "Metric calculation takes about four hours.\nhttps://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/621455",
    "3399184": "Do you know which GPU will be available for the testing phase?\n",
    "3396244": ""
  }
}