{
  "id": 410985,
  "title": "[Help] Save GPU quota during commit",
  "url": "/competitions/vesuvius-challenge-ink-detection/discussion/410985",
  "author_name": "",
  "post_date": "2023-05-17T08:43:54.176330400Z",
  "votes": 5,
  "comment_count": 7,
  "views": 0,
  "content": "<p>As the post says, I'm looking for a way to save GPU quota during submission time. </p>\n<p>Usually in code-competitions we use a conditional statement like <code>len(sample_test) == N</code> to check if we are in commit or submission/running hidden test phase, but here it is not applicable. </p>\n<p>I tried also with the <code>os.environ</code> variable, the logic is simple, eg </p>\n<pre><code>if os.getenv('KAGGLE_IS_COMPETITION_RERUN'):\n         &lt; run inference and submit &gt;  \nelse:\n    sub = pd.read_csv('sample_submission.csv').to_csv(\"submission.csv\", index=False)\n</code></pre>\n<p>but I get continuously Notebook \"Throw exception error\".<br>\nAny ideas on how to achieve that will be appreciated! <br>\nI guess will be helpful for many others too. <br>\nthanks</p>",
  "messages": [
    {
      "id": "2262976",
      "postDate": "05/17/2023 08:43:54",
      "content": "<p>As the post says, I'm looking for a way to save GPU quota during submission time. </p>\n<p>Usually in code-competitions we use a conditional statement like <code>len(sample_test) == N</code> to check if we are in commit or submission/running hidden test phase, but here it is not applicable. </p>\n<p>I tried also with the <code>os.environ</code> variable, the logic is simple, eg </p>\n<pre><code>if os.getenv('KAGGLE_IS_COMPETITION_RERUN'):\n         &lt; run inference and submit &gt;  \nelse:\n    sub = pd.read_csv('sample_submission.csv').to_csv(\"submission.csv\", index=False)\n</code></pre>\n<p>but I get continuously Notebook \"Throw exception error\".<br>\nAny ideas on how to achieve that will be appreciated! <br>\nI guess will be helpful for many others too. <br>\nthanks</p>",
      "rawMarkdown": "As the post says, I'm looking for a way to save GPU quota during submission time. \n\nUsually in code-competitions we use a conditional statement like `len(sample_test) == N` to check if we are in commit or submission/running hidden test phase, but here it is not applicable. \n\nI tried also with the `os.environ` variable, the logic is simple, eg \n\n```\nif os.getenv('KAGGLE_IS_COMPETITION_RERUN'):\n         < run inference and submit >  \nelse:\n    sub = pd.read_csv('sample_submission.csv').to_csv(\"submission.csv\", index=False)\n```\nbut I get continuously Notebook \"Throw exception error\".\nAny ideas on how to achieve that will be appreciated! \nI guess will be helpful for many others too. \nthanks",
      "votes": null
    },
    {
      "id": "2263079",
      "postDate": "05/17/2023 10:15:37",
      "content": "<p>I’m doing this:</p>\n<pre><code>run_inference = False\n\ntry:\n    x = cv2.imread(\"/kaggle/input/vesuvius-challenge-ink-detection/test/a/mask.png\", 0)\n    if x.shape != (2727, 6330):\n        run_inference = True\n\nexcept:\n    run_inference = True\n</code></pre>",
      "rawMarkdown": "I’m doing this:\n```\nrun_inference = False\n\ntry:\n    x = cv2.imread(\"/kaggle/input/vesuvius-challenge-ink-detection/test/a/mask.png\", 0)\n    if x.shape != (2727, 6330):\n        run_inference = True\n\nexcept:\n    run_inference = True\n```",
      "votes": null
    },
    {
      "id": "2263734",
      "postDate": "05/17/2023 20:57:22",
      "content": "<p>I am approaching it this way, and it's going well.</p>\n<pre><code>QUICK_SAVE = True\n\nsample_submission_flag = filecmp.cmp(\n    \"../input/vesuvius-challenge-ink-detection/test/a/surface_volume/00.tif\",\n    \"../input/vcid-file-check/00.tif\",\n    shallow=True\n)\n\nif sample_submission_flag and QUICK_SAVE:\n    df_sub = pd.read_csv(\"../input/vesuvius-challenge-ink-detection/sample_submission.csv\")\n    df_sub.to_csv(\"submission.csv\", index=False)\nelse:\n</code></pre>\n<p>vcid-file-check is here:<br>\n<a href=\"https://www.kaggle.com/datasets/tmyok1984/vcid-file-check\" target=\"_blank\">https://www.kaggle.com/datasets/tmyok1984/vcid-file-check</a></p>",
      "rawMarkdown": "I am approaching it this way, and it's going well.\n\n```\nQUICK_SAVE = True\n\nsample_submission_flag = filecmp.cmp(\n    \"../input/vesuvius-challenge-ink-detection/test/a/surface_volume/00.tif\",\n    \"../input/vcid-file-check/00.tif\",\n    shallow=True\n)\n\nif sample_submission_flag and QUICK_SAVE:\n    df_sub = pd.read_csv(\"../input/vesuvius-challenge-ink-detection/sample_submission.csv\")\n    df_sub.to_csv(\"submission.csv\", index=False)\nelse:\n```\n\nvcid-file-check is here:\nhttps://www.kaggle.com/datasets/tmyok1984/vcid-file-check",
      "votes": null
    },
    {
      "id": "2263829",
      "postDate": "05/18/2023 00:45:40",
      "content": "<p>Thank you! It is very convenient.</p>",
      "rawMarkdown": "Thank you! It is very convenient.",
      "votes": null
    },
    {
      "id": "2263840",
      "postDate": "05/18/2023 00:59:07",
      "content": "<p>In my case:</p>\n<pre><code>from pathlib import Path\nimport hashlib\n\ndata_root = Path(\"/kaggle/input/vesuvius-challenge-ink-detection\")\ntest_root = data_root.joinpath(\"test\")\ntest_mask = sorted(test_root.glob(\"**/*.png\"))[0]\n\nwith test_mask.open(\"rb\") as f:\n    hash_md5 = hashlib.md5(f.read()).hexdigest()\n\nfast_sub = hash_md5 == \"0b0fffdc0e88be226673846a143bb3e0\"\nfast_sub\n</code></pre>",
      "rawMarkdown": "In my case:\n\n```\nfrom pathlib import Path\nimport hashlib\n\ndata_root = Path(\"/kaggle/input/vesuvius-challenge-ink-detection\")\ntest_root = data_root.joinpath(\"test\")\ntest_mask = sorted(test_root.glob(\"**/*.png\"))[0]\n\nwith test_mask.open(\"rb\") as f:\n    hash_md5 = hashlib.md5(f.read()).hexdigest()\n    \nfast_sub = hash_md5 == \"0b0fffdc0e88be226673846a143bb3e0\"\nfast_sub\n```",
      "votes": null
    },
    {
      "id": "2263868",
      "postDate": "05/18/2023 01:41:12",
      "content": "<p>should feedback this to kaggle product manager team</p>",
      "rawMarkdown": "should feedback this to kaggle product manager team",
      "votes": null
    },
    {
      "id": "2265010",
      "postDate": "05/18/2023 23:24:45",
      "content": "<p>click cancel session after submit is the fastest (but you lose your notebook formatting)</p>",
      "rawMarkdown": "click cancel session after submit is the fastest (but you lose your notebook formatting)",
      "votes": null
    },
    {
      "id": "2265496",
      "postDate": "05/19/2023 09:29:03",
      "content": "<p>Thank you all for your time! <br>\nissue has been resolved for me, I started from scratch and build step-by-step a new inference pipeline taking into account your suggestions - I had many failed subs last days but the end result is that I managed to reduce from +1h to &lt;5min runtime during commit. </p>",
      "rawMarkdown": "Thank you all for your time! \nissue has been resolved for me, I started from scratch and build step-by-step a new inference pipeline taking into account your suggestions - I had many failed subs last days but the end result is that I managed to reduce from +1h to <5min runtime during commit.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2263079,
      "author_name": "igorkrashenyi",
      "author_url": "",
      "post_date": "05/17/2023 10:15:37",
      "content": "<p>I’m doing this:</p>\n<pre><code>run_inference = False\n\ntry:\n    x = cv2.imread(\"/kaggle/input/vesuvius-challenge-ink-detection/test/a/mask.png\", 0)\n    if x.shape != (2727, 6330):\n        run_inference = True\n\nexcept:\n    run_inference = True\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2263734,
      "author_name": "tmyok1984",
      "author_url": "",
      "post_date": "05/17/2023 20:57:22",
      "content": "<p>I am approaching it this way, and it's going well.</p>\n<pre><code>QUICK_SAVE = True\n\nsample_submission_flag = filecmp.cmp(\n    \"../input/vesuvius-challenge-ink-detection/test/a/surface_volume/00.tif\",\n    \"../input/vcid-file-check/00.tif\",\n    shallow=True\n)\n\nif sample_submission_flag and QUICK_SAVE:\n    df_sub = pd.read_csv(\"../input/vesuvius-challenge-ink-detection/sample_submission.csv\")\n    df_sub.to_csv(\"submission.csv\", index=False)\nelse:\n</code></pre>\n<p>vcid-file-check is here:<br>\n<a href=\"https://www.kaggle.com/datasets/tmyok1984/vcid-file-check\" target=\"_blank\">https://www.kaggle.com/datasets/tmyok1984/vcid-file-check</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 2263829,
          "author_name": "junseonglee11",
          "author_url": "",
          "post_date": "05/18/2023 00:45:40",
          "content": "<p>Thank you! It is very convenient.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2263840,
      "author_name": "ren4yu",
      "author_url": "",
      "post_date": "05/18/2023 00:59:07",
      "content": "<p>In my case:</p>\n<pre><code>from pathlib import Path\nimport hashlib\n\ndata_root = Path(\"/kaggle/input/vesuvius-challenge-ink-detection\")\ntest_root = data_root.joinpath(\"test\")\ntest_mask = sorted(test_root.glob(\"**/*.png\"))[0]\n\nwith test_mask.open(\"rb\") as f:\n    hash_md5 = hashlib.md5(f.read()).hexdigest()\n\nfast_sub = hash_md5 == \"0b0fffdc0e88be226673846a143bb3e0\"\nfast_sub\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2263868,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "05/18/2023 01:41:12",
      "content": "<p>should feedback this to kaggle product manager team</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2265010,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "05/18/2023 23:24:45",
      "content": "<p>click cancel session after submit is the fastest (but you lose your notebook formatting)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2265496,
      "author_name": "imeintanis",
      "author_url": "",
      "post_date": "05/19/2023 09:29:03",
      "content": "<p>Thank you all for your time! <br>\nissue has been resolved for me, I started from scratch and build step-by-step a new inference pipeline taking into account your suggestions - I had many failed subs last days but the end result is that I managed to reduce from +1h to &lt;5min runtime during commit. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2262976": "As the post says, I'm looking for a way to save GPU quota during submission time. \n\nUsually in code-competitions we use a conditional statement like `len(sample_test) == N` to check if we are in commit or submission/running hidden test phase, but here it is not applicable. \n\nI tried also with the `os.environ` variable, the logic is simple, eg \n\n```\nif os.getenv('KAGGLE_IS_COMPETITION_RERUN'):\n         < run inference and submit >  \nelse:\n    sub = pd.read_csv('sample_submission.csv').to_csv(\"submission.csv\", index=False)\n```\nbut I get continuously Notebook \"Throw exception error\".\nAny ideas on how to achieve that will be appreciated! \nI guess will be helpful for many others too. \nthanks",
    "2263079": "I’m doing this:\n```\nrun_inference = False\n\ntry:\n    x = cv2.imread(\"/kaggle/input/vesuvius-challenge-ink-detection/test/a/mask.png\", 0)\n    if x.shape != (2727, 6330):\n        run_inference = True\n\nexcept:\n    run_inference = True\n```",
    "2263734": "I am approaching it this way, and it's going well.\n\n```\nQUICK_SAVE = True\n\nsample_submission_flag = filecmp.cmp(\n    \"../input/vesuvius-challenge-ink-detection/test/a/surface_volume/00.tif\",\n    \"../input/vcid-file-check/00.tif\",\n    shallow=True\n)\n\nif sample_submission_flag and QUICK_SAVE:\n    df_sub = pd.read_csv(\"../input/vesuvius-challenge-ink-detection/sample_submission.csv\")\n    df_sub.to_csv(\"submission.csv\", index=False)\nelse:\n```\n\nvcid-file-check is here:\nhttps://www.kaggle.com/datasets/tmyok1984/vcid-file-check",
    "2263829": "Thank you! It is very convenient.",
    "2263840": "In my case:\n\n```\nfrom pathlib import Path\nimport hashlib\n\ndata_root = Path(\"/kaggle/input/vesuvius-challenge-ink-detection\")\ntest_root = data_root.joinpath(\"test\")\ntest_mask = sorted(test_root.glob(\"**/*.png\"))[0]\n\nwith test_mask.open(\"rb\") as f:\n    hash_md5 = hashlib.md5(f.read()).hexdigest()\n    \nfast_sub = hash_md5 == \"0b0fffdc0e88be226673846a143bb3e0\"\nfast_sub\n```",
    "2263868": "should feedback this to kaggle product manager team",
    "2265010": "click cancel session after submit is the fastest (but you lose your notebook formatting)",
    "2265496": "Thank you all for your time! \nissue has been resolved for me, I started from scratch and build step-by-step a new inference pipeline taking into account your suggestions - I had many failed subs last days but the end result is that I managed to reduce from +1h to <5min runtime during commit."
  },
  "source": "meta"
}