{
  "id": 497895,
  "title": "Use the exit function in python to utilize GPU resources effectively upon submission",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/497895",
  "author_name": "",
  "post_date": "2024-04-26T06:23:10.756158200Z",
  "votes": 33,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hello everyone,</p>\n<p>I am sure a lot of us face issues with the GPU quotas for the week and perhaps we could face even bigger challenges as Thursday/ Friday approaches as we run out of quotas for the week. I suggest a simple trick to perhaps palliate this problem, especially during our submission to the leaderboard-</p>\n<h2>What could be done</h2>\n<p>Python has an inbuilt <strong>exit</strong> function that can be used to terminate the program execution at any time, based on the <strong>exit status</strong>. A value of 0 here indicates a normal error-free exit, while a non-zero value indicates the presence of an error. The <strong>sys</strong> module also has an exit method that does the same task, i.e. <strong>sys.exit(0)</strong>. Other options include <strong>quit()</strong> and <strong>os._exit()</strong>. Please note that these are implemented slightly differently from each other, so use the one that works for you! In my opinion, using <strong>sys.exit()</strong> is the best recourse for general tasks outside of IPython (we usually use this well outside of Jupyter)</p>\n<p>Please find a few adjutant links herewith for this function-</p>\n<ol>\n<li><a href=\"https://www.freecodecamp.org/news/python-exit-how-to-use-an-exit-function-in-python-to-stop-a-program/\" target=\"_blank\">https://www.freecodecamp.org/news/python-exit-how-to-use-an-exit-function-in-python-to-stop-a-program/</a></li>\n<li><a href=\"https://www.freecodecamp.org/news/python-exit-how-to-use-an-exit-function-in-python-to-stop-a-program/\" target=\"_blank\">https://www.freecodecamp.org/news/python-exit-how-to-use-an-exit-function-in-python-to-stop-a-program/</a></li>\n<li><a href=\"https://superfastpython.com/exit-process/\" target=\"_blank\">https://superfastpython.com/exit-process/</a></li>\n<li><a href=\"https://www.geeksforgeeks.org/python-exit-commands-quit-exit-sys-exit-and-os-_exit/\" target=\"_blank\">https://www.geeksforgeeks.org/python-exit-commands-quit-exit-sys-exit-and-os-_exit/</a></li>\n</ol>\n<p>So how do we use it effectively in our code and save GPU resources while submitting? The trick is very simple, just add the small code cell at the start of the kernel- </p>\n<pre><code>import pandas as pd\nsubmission = pd.read_csv()\n\n len(submission) &lt;= :\n    submission[] = \n    submission = submission.set_index()\n    submission.to_csv()\n    print(f)\n    ()\n</code></pre>\n<p>We do the below herewith-</p>\n<ol>\n<li>We import the sample submission file from the input folders and check the size. </li>\n<li>If the size of the submission file is lower than 100 elements (indicating a dry run/ sample values and not the actual test data), we simply bypass any further execution and post 0.5 as the single value in the score column. This value is actually present in the sample file, so this step may be repealed as well. </li>\n<li>We set the index to the case_id column and make a successful <strong>dry submission</strong>. This effectively completes our required task to enable a submission!</li>\n<li><strong>We could ideally exit at this stage to conserve GPU quotas and we do exactly this!</strong>. Using exit(0) herewith enables us to safely leave the kernel at this stage and bypass any further execution, amounting to less than a minute of GPU usage, along with a submission to the leaderboard as well! </li>\n</ol>\n<h2>An alternative route</h2>\n<p>If one finds exit()/ sys.exit() cumbersome/ does not work well, then another option is here for you-</p>\n<ul>\n<li>Make a copy of your training process kernel</li>\n<li>Convert the kernel into a script using the file menu</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8273630%2Fdb0f98b7bc075e56793277c7de645d12%2Feditor.png?generation=1714116439033445&amp;alt=media\"></p>\n<ul>\n<li>Add the below code cell at the start of the script- </li>\n</ul>\n<pre><code> pandas  pd\nsubmission = pd.read_csv()\n\n (submission) &lt;= :\n    submission[] = \n    submission = submission.set_index()\n    submission.to_csv()\n    ()\n:\n    ..... your entire code (you need to indent it appropriately)\n</code></pre>\n<ul>\n<li>Convert the kernel editor type to <strong>kernel</strong> - this puts the entire code in 1 cell. It appears clunky but don't worry. </li>\n</ul>\n<p>This achieves the same result as the one above, without using exit/ quit, saving the GPU quota upon submission. The only problem here is that the submission code looks clunky and messy with a long cell. This should be fine as we ideally may not work actively on the kernel used primarily for submission.</p>\n<h2>What is being done instead</h2>\n<ol>\n<li>Many public kernels are using a variable DRY_RUN boolean that sets a value based on the size of the submission file</li>\n<li>This variable merely controls the size of the train data <strong>after feature creation and during model training</strong> for a smaller training pipeline</li>\n<li><strong>Please note that these kernels do not use GPU to create features, so this process of multiple feature creation is extremely CPU intensive without any GPU usage, incurring a heavy GPU opportunity cost</strong></li>\n</ol>\n<h2>Rely on CV scores</h2>\n<ol>\n<li>I may forebode that many of us wish to climb up the leaderboard and effectively use public materials as well to do so, but quite a few public kernels are not even assessing the CV scores while making a submission</li>\n<li>In a competition, CV scheme analysis and CV attribution is the most important element. If this step is bypassed, then we are no different from a random gamble.</li>\n<li>I urge participants to test models and features locally using the CV score and then decide next steps. I leave the CV scheme and other elements to the participant to determine. </li>\n</ol>\n<p>Hope this helps to make a safe submission and optimize GPU usage too! <br>\nHappy learning and best regards! </p>\n<p>All the best!</p>",
  "messages": [
    {
      "id": "2776363",
      "postDate": "04/26/2024 06:23:10",
      "content": "<p>Hello everyone,</p>\n<p>I am sure a lot of us face issues with the GPU quotas for the week and perhaps we could face even bigger challenges as Thursday/ Friday approaches as we run out of quotas for the week. I suggest a simple trick to perhaps palliate this problem, especially during our submission to the leaderboard-</p>\n<h2>What could be done</h2>\n<p>Python has an inbuilt <strong>exit</strong> function that can be used to terminate the program execution at any time, based on the <strong>exit status</strong>. A value of 0 here indicates a normal error-free exit, while a non-zero value indicates the presence of an error. The <strong>sys</strong> module also has an exit method that does the same task, i.e. <strong>sys.exit(0)</strong>. Other options include <strong>quit()</strong> and <strong>os._exit()</strong>. Please note that these are implemented slightly differently from each other, so use the one that works for you! In my opinion, using <strong>sys.exit()</strong> is the best recourse for general tasks outside of IPython (we usually use this well outside of Jupyter)</p>\n<p>Please find a few adjutant links herewith for this function-</p>\n<ol>\n<li><a href=\"https://www.freecodecamp.org/news/python-exit-how-to-use-an-exit-function-in-python-to-stop-a-program/\" target=\"_blank\">https://www.freecodecamp.org/news/python-exit-how-to-use-an-exit-function-in-python-to-stop-a-program/</a></li>\n<li><a href=\"https://www.freecodecamp.org/news/python-exit-how-to-use-an-exit-function-in-python-to-stop-a-program/\" target=\"_blank\">https://www.freecodecamp.org/news/python-exit-how-to-use-an-exit-function-in-python-to-stop-a-program/</a></li>\n<li><a href=\"https://superfastpython.com/exit-process/\" target=\"_blank\">https://superfastpython.com/exit-process/</a></li>\n<li><a href=\"https://www.geeksforgeeks.org/python-exit-commands-quit-exit-sys-exit-and-os-_exit/\" target=\"_blank\">https://www.geeksforgeeks.org/python-exit-commands-quit-exit-sys-exit-and-os-_exit/</a></li>\n</ol>\n<p>So how do we use it effectively in our code and save GPU resources while submitting? The trick is very simple, just add the small code cell at the start of the kernel- </p>\n<pre><code>import pandas as pd\nsubmission = pd.read_csv()\n\n len(submission) &lt;= :\n    submission[] = \n    submission = submission.set_index()\n    submission.to_csv()\n    print(f)\n    ()\n</code></pre>\n<p>We do the below herewith-</p>\n<ol>\n<li>We import the sample submission file from the input folders and check the size. </li>\n<li>If the size of the submission file is lower than 100 elements (indicating a dry run/ sample values and not the actual test data), we simply bypass any further execution and post 0.5 as the single value in the score column. This value is actually present in the sample file, so this step may be repealed as well. </li>\n<li>We set the index to the case_id column and make a successful <strong>dry submission</strong>. This effectively completes our required task to enable a submission!</li>\n<li><strong>We could ideally exit at this stage to conserve GPU quotas and we do exactly this!</strong>. Using exit(0) herewith enables us to safely leave the kernel at this stage and bypass any further execution, amounting to less than a minute of GPU usage, along with a submission to the leaderboard as well! </li>\n</ol>\n<h2>An alternative route</h2>\n<p>If one finds exit()/ sys.exit() cumbersome/ does not work well, then another option is here for you-</p>\n<ul>\n<li>Make a copy of your training process kernel</li>\n<li>Convert the kernel into a script using the file menu</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8273630%2Fdb0f98b7bc075e56793277c7de645d12%2Feditor.png?generation=1714116439033445&amp;alt=media\"></p>\n<ul>\n<li>Add the below code cell at the start of the script- </li>\n</ul>\n<pre><code> pandas  pd\nsubmission = pd.read_csv()\n\n (submission) &lt;= :\n    submission[] = \n    submission = submission.set_index()\n    submission.to_csv()\n    ()\n:\n    ..... your entire code (you need to indent it appropriately)\n</code></pre>\n<ul>\n<li>Convert the kernel editor type to <strong>kernel</strong> - this puts the entire code in 1 cell. It appears clunky but don't worry. </li>\n</ul>\n<p>This achieves the same result as the one above, without using exit/ quit, saving the GPU quota upon submission. The only problem here is that the submission code looks clunky and messy with a long cell. This should be fine as we ideally may not work actively on the kernel used primarily for submission.</p>\n<h2>What is being done instead</h2>\n<ol>\n<li>Many public kernels are using a variable DRY_RUN boolean that sets a value based on the size of the submission file</li>\n<li>This variable merely controls the size of the train data <strong>after feature creation and during model training</strong> for a smaller training pipeline</li>\n<li><strong>Please note that these kernels do not use GPU to create features, so this process of multiple feature creation is extremely CPU intensive without any GPU usage, incurring a heavy GPU opportunity cost</strong></li>\n</ol>\n<h2>Rely on CV scores</h2>\n<ol>\n<li>I may forebode that many of us wish to climb up the leaderboard and effectively use public materials as well to do so, but quite a few public kernels are not even assessing the CV scores while making a submission</li>\n<li>In a competition, CV scheme analysis and CV attribution is the most important element. If this step is bypassed, then we are no different from a random gamble.</li>\n<li>I urge participants to test models and features locally using the CV score and then decide next steps. I leave the CV scheme and other elements to the participant to determine. </li>\n</ol>\n<p>Hope this helps to make a safe submission and optimize GPU usage too! <br>\nHappy learning and best regards! </p>\n<p>All the best!</p>",
      "rawMarkdown": "Hello everyone,\n\nI am sure a lot of us face issues with the GPU quotas for the week and perhaps we could face even bigger challenges as Thursday/ Friday approaches as we run out of quotas for the week. I suggest a simple trick to perhaps palliate this problem, especially during our submission to the leaderboard-\n\n## What could be done\n\nPython has an inbuilt **exit** function that can be used to terminate the program execution at any time, based on the **exit status**. A value of 0 here indicates a normal error-free exit, while a non-zero value indicates the presence of an error. The **sys** module also has an exit method that does the same task, i.e. **sys.exit(0)**. Other options include **quit()** and **os._exit()**. Please note that these are implemented slightly differently from each other, so use the one that works for you! In my opinion, using **sys.exit()** is the best recourse for general tasks outside of IPython (we usually use this well outside of Jupyter)\n\nPlease find a few adjutant links herewith for this function-\n1. https://www.freecodecamp.org/news/python-exit-how-to-use-an-exit-function-in-python-to-stop-a-program/\n2. https://www.freecodecamp.org/news/python-exit-how-to-use-an-exit-function-in-python-to-stop-a-program/\n3. https://superfastpython.com/exit-process/\n4. https://www.geeksforgeeks.org/python-exit-commands-quit-exit-sys-exit-and-os-_exit/\n\nSo how do we use it effectively in our code and save GPU resources while submitting? The trick is very simple, just add the small code cell at the start of the kernel- \n\n```\nimport pandas as pd\nsubmission = pd.read_csv(\"/kaggle/input/home-credit-credit-risk-model-stability/sample_submission.csv\")\n\nif len(submission) <= 100:\n    submission[\"score\"] = 0.50\n    submission = submission.set_index(\"case_id\")\n    submission.to_csv(\"submission.csv\")\n    print(f\"Sample submission done\")\n    exit(0)\n```\n\nWe do the below herewith-\n1. We import the sample submission file from the input folders and check the size. \n2. If the size of the submission file is lower than 100 elements (indicating a dry run/ sample values and not the actual test data), we simply bypass any further execution and post 0.5 as the single value in the score column. This value is actually present in the sample file, so this step may be repealed as well. \n3. We set the index to the case_id column and make a successful **dry submission**. This effectively completes our required task to enable a submission!\n4. **We could ideally exit at this stage to conserve GPU quotas and we do exactly this!**. Using exit(0) herewith enables us to safely leave the kernel at this stage and bypass any further execution, amounting to less than a minute of GPU usage, along with a submission to the leaderboard as well! \n\n## An alternative route\nIf one finds exit()/ sys.exit() cumbersome/ does not work well, then another option is here for you-\n- Make a copy of your training process kernel\n- Convert the kernel into a script using the file menu\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8273630%2Fdb0f98b7bc075e56793277c7de645d12%2Feditor.png?generation=1714116439033445&alt=media)\n- Add the below code cell at the start of the script- \n\n```python\nimport pandas as pd\nsubmission = pd.read_csv(\"/kaggle/input/home-credit-credit-risk-model-stability/sample_submission.csv\")\n\nif len(submission) <= 100:\n    submission[\"score\"] = 0.50\n    submission = submission.set_index(\"case_id\")\n    submission.to_csv(\"submission.csv\")\n    print(f\"Sample submission done\")\nelse:\n    ..... your entire code (you need to indent it appropriately)\n```\n- Convert the kernel editor type to **kernel** - this puts the entire code in 1 cell. It appears clunky but don't worry. \n\nThis achieves the same result as the one above, without using exit/ quit, saving the GPU quota upon submission. The only problem here is that the submission code looks clunky and messy with a long cell. This should be fine as we ideally may not work actively on the kernel used primarily for submission.\n\n## What is being done instead\n1. Many public kernels are using a variable DRY_RUN boolean that sets a value based on the size of the submission file\n2. This variable merely controls the size of the train data **after feature creation and during model training** for a smaller training pipeline\n3. **Please note that these kernels do not use GPU to create features, so this process of multiple feature creation is extremely CPU intensive without any GPU usage, incurring a heavy GPU opportunity cost**\n\n## Rely on CV scores\n1. I may forebode that many of us wish to climb up the leaderboard and effectively use public materials as well to do so, but quite a few public kernels are not even assessing the CV scores while making a submission\n2. In a competition, CV scheme analysis and CV attribution is the most important element. If this step is bypassed, then we are no different from a random gamble.\n3. I urge participants to test models and features locally using the CV score and then decide next steps. I leave the CV scheme and other elements to the participant to determine. \n\nHope this helps to make a safe submission and optimize GPU usage too! \nHappy learning and best regards! \n\nAll the best!",
      "votes": null
    },
    {
      "id": "2776534",
      "postDate": "04/26/2024 08:13:53",
      "content": "<p>In addition, the feature creation can be done in the separate utility script, saved in it (training ds), and then loaded into the training/inference notebook (functions from this utility script can also be used).</p>",
      "rawMarkdown": "In addition, the feature creation can be done in the separate utility script, saved in it (training ds), and then loaded into the training/inference notebook (functions from this utility script can also be used).",
      "votes": null
    },
    {
      "id": "2776566",
      "postDate": "04/26/2024 08:27:44",
      "content": "<p>Yes, constructing features does not require a GPU, it can be done using a CPU.</p>",
      "rawMarkdown": "Yes, constructing features does not require a GPU, it can be done using a CPU.",
      "votes": null
    },
    {
      "id": "2776578",
      "postDate": "04/26/2024 08:34:32",
      "content": "<p><a href=\"https://www.kaggle.com/andreynesterov\" target=\"_blank\">@andreynesterov</a> yes of course, my public work does just that </p>",
      "rawMarkdown": "andreynesterov yes of course, my public work does just that",
      "votes": null
    },
    {
      "id": "2777030",
      "postDate": "04/26/2024 13:25:16",
      "content": "<p>Local CV scores seems unreliable, so I just cancelled the other running notebook when submit.</p>",
      "rawMarkdown": "Local CV scores seems unreliable, so I just cancelled the other running notebook when submit.",
      "votes": null
    },
    {
      "id": "2777066",
      "postDate": "04/26/2024 13:52:22",
      "content": "<p>This is a problem herewith in this competition!<br>\nAll the best <a href=\"https://www.kaggle.com/huangshibao\" target=\"_blank\">@huangshibao</a> </p>",
      "rawMarkdown": "This is a problem herewith in this competition!\nAll the best @huangshibao",
      "votes": null
    },
    {
      "id": "2777949",
      "postDate": "04/26/2024 21:44:40",
      "content": "<p>You can make it more generic by replacing </p>\n<pre><code>= :\n  ...\n</code></pre>\n<p><br>\nwith </p>\n<pre><code>  .():\n  ...\n</code></pre>\n<p>See <a href=\"https://www.kaggle.com/discussions/product-feedback/315792\" target=\"_blank\">this post</a> for more details.</p>",
      "rawMarkdown": "You can make it more generic by replacing \n```\nif len(submission) <= 100:\n  ...\n``` \nwith \n```\nif not os.getenv('KAGGLE_IS_COMPETITION_RERUN'):\n  ...\n```\n\nSee [this post](https://www.kaggle.com/discussions/product-feedback/315792) for more details.",
      "votes": null
    },
    {
      "id": "2816867",
      "postDate": "05/16/2024 15:12:51",
      "content": "<p>Thanks for your sharing~</p>",
      "rawMarkdown": "Thanks for your sharing~",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2776534,
      "author_name": "andreynesterov",
      "author_url": "",
      "post_date": "04/26/2024 08:13:53",
      "content": "<p>In addition, the feature creation can be done in the separate utility script, saved in it (training ds), and then loaded into the training/inference notebook (functions from this utility script can also be used).</p>",
      "votes": null,
      "replies": [
        {
          "id": 2776566,
          "author_name": "yunsuxiaozi",
          "author_url": "",
          "post_date": "04/26/2024 08:27:44",
          "content": "<p>Yes, constructing features does not require a GPU, it can be done using a CPU.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2776578,
          "author_name": "ravi20076",
          "author_url": "",
          "post_date": "04/26/2024 08:34:32",
          "content": "<p><a href=\"https://www.kaggle.com/andreynesterov\" target=\"_blank\">@andreynesterov</a> yes of course, my public work does just that </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2777030,
      "author_name": "huangshibao",
      "author_url": "",
      "post_date": "04/26/2024 13:25:16",
      "content": "<p>Local CV scores seems unreliable, so I just cancelled the other running notebook when submit.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2777066,
          "author_name": "ravi20076",
          "author_url": "",
          "post_date": "04/26/2024 13:52:22",
          "content": "<p>This is a problem herewith in this competition!<br>\nAll the best <a href=\"https://www.kaggle.com/huangshibao\" target=\"_blank\">@huangshibao</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2777949,
      "author_name": "kononenko",
      "author_url": "",
      "post_date": "04/26/2024 21:44:40",
      "content": "<p>You can make it more generic by replacing </p>\n<pre><code>= :\n  ...\n</code></pre>\n<p><br>\nwith </p>\n<pre><code>  .():\n  ...\n</code></pre>\n<p>See <a href=\"https://www.kaggle.com/discussions/product-feedback/315792\" target=\"_blank\">this post</a> for more details.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2816867,
      "author_name": "yumiaogao",
      "author_url": "",
      "post_date": "05/16/2024 15:12:51",
      "content": "<p>Thanks for your sharing~</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2776363": "Hello everyone,\n\nI am sure a lot of us face issues with the GPU quotas for the week and perhaps we could face even bigger challenges as Thursday/ Friday approaches as we run out of quotas for the week. I suggest a simple trick to perhaps palliate this problem, especially during our submission to the leaderboard-\n\n## What could be done\n\nPython has an inbuilt **exit** function that can be used to terminate the program execution at any time, based on the **exit status**. A value of 0 here indicates a normal error-free exit, while a non-zero value indicates the presence of an error. The **sys** module also has an exit method that does the same task, i.e. **sys.exit(0)**. Other options include **quit()** and **os._exit()**. Please note that these are implemented slightly differently from each other, so use the one that works for you! In my opinion, using **sys.exit()** is the best recourse for general tasks outside of IPython (we usually use this well outside of Jupyter)\n\nPlease find a few adjutant links herewith for this function-\n1. https://www.freecodecamp.org/news/python-exit-how-to-use-an-exit-function-in-python-to-stop-a-program/\n2. https://www.freecodecamp.org/news/python-exit-how-to-use-an-exit-function-in-python-to-stop-a-program/\n3. https://superfastpython.com/exit-process/\n4. https://www.geeksforgeeks.org/python-exit-commands-quit-exit-sys-exit-and-os-_exit/\n\nSo how do we use it effectively in our code and save GPU resources while submitting? The trick is very simple, just add the small code cell at the start of the kernel- \n\n```\nimport pandas as pd\nsubmission = pd.read_csv(\"/kaggle/input/home-credit-credit-risk-model-stability/sample_submission.csv\")\n\nif len(submission) <= 100:\n    submission[\"score\"] = 0.50\n    submission = submission.set_index(\"case_id\")\n    submission.to_csv(\"submission.csv\")\n    print(f\"Sample submission done\")\n    exit(0)\n```\n\nWe do the below herewith-\n1. We import the sample submission file from the input folders and check the size. \n2. If the size of the submission file is lower than 100 elements (indicating a dry run/ sample values and not the actual test data), we simply bypass any further execution and post 0.5 as the single value in the score column. This value is actually present in the sample file, so this step may be repealed as well. \n3. We set the index to the case_id column and make a successful **dry submission**. This effectively completes our required task to enable a submission!\n4. **We could ideally exit at this stage to conserve GPU quotas and we do exactly this!**. Using exit(0) herewith enables us to safely leave the kernel at this stage and bypass any further execution, amounting to less than a minute of GPU usage, along with a submission to the leaderboard as well! \n\n## An alternative route\nIf one finds exit()/ sys.exit() cumbersome/ does not work well, then another option is here for you-\n- Make a copy of your training process kernel\n- Convert the kernel into a script using the file menu\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8273630%2Fdb0f98b7bc075e56793277c7de645d12%2Feditor.png?generation=1714116439033445&alt=media)\n- Add the below code cell at the start of the script- \n\n```python\nimport pandas as pd\nsubmission = pd.read_csv(\"/kaggle/input/home-credit-credit-risk-model-stability/sample_submission.csv\")\n\nif len(submission) <= 100:\n    submission[\"score\"] = 0.50\n    submission = submission.set_index(\"case_id\")\n    submission.to_csv(\"submission.csv\")\n    print(f\"Sample submission done\")\nelse:\n    ..... your entire code (you need to indent it appropriately)\n```\n- Convert the kernel editor type to **kernel** - this puts the entire code in 1 cell. It appears clunky but don't worry. \n\nThis achieves the same result as the one above, without using exit/ quit, saving the GPU quota upon submission. The only problem here is that the submission code looks clunky and messy with a long cell. This should be fine as we ideally may not work actively on the kernel used primarily for submission.\n\n## What is being done instead\n1. Many public kernels are using a variable DRY_RUN boolean that sets a value based on the size of the submission file\n2. This variable merely controls the size of the train data **after feature creation and during model training** for a smaller training pipeline\n3. **Please note that these kernels do not use GPU to create features, so this process of multiple feature creation is extremely CPU intensive without any GPU usage, incurring a heavy GPU opportunity cost**\n\n## Rely on CV scores\n1. I may forebode that many of us wish to climb up the leaderboard and effectively use public materials as well to do so, but quite a few public kernels are not even assessing the CV scores while making a submission\n2. In a competition, CV scheme analysis and CV attribution is the most important element. If this step is bypassed, then we are no different from a random gamble.\n3. I urge participants to test models and features locally using the CV score and then decide next steps. I leave the CV scheme and other elements to the participant to determine. \n\nHope this helps to make a safe submission and optimize GPU usage too! \nHappy learning and best regards! \n\nAll the best!",
    "2776534": "In addition, the feature creation can be done in the separate utility script, saved in it (training ds), and then loaded into the training/inference notebook (functions from this utility script can also be used).",
    "2776566": "Yes, constructing features does not require a GPU, it can be done using a CPU.",
    "2776578": "andreynesterov yes of course, my public work does just that",
    "2777030": "Local CV scores seems unreliable, so I just cancelled the other running notebook when submit.",
    "2777066": "This is a problem herewith in this competition!\nAll the best @huangshibao",
    "2777949": "You can make it more generic by replacing \n```\nif len(submission) <= 100:\n  ...\n``` \nwith \n```\nif not os.getenv('KAGGLE_IS_COMPETITION_RERUN'):\n  ...\n```\n\nSee [this post](https://www.kaggle.com/discussions/product-feedback/315792) for more details.",
    "2816867": "Thanks for your sharing~"
  },
  "source": "meta"
}