{
  "id": 115913,
  "title": "Commits Keep on running and wasting GPU hours",
  "url": "/competitions/understanding_cloud_organization/discussion/115913",
  "author_name": "",
  "post_date": "2019-11-06T03:51:22.646537200Z",
  "votes": 18,
  "comment_count": 31,
  "views": 0,
  "content": "<p>Few times, my inference kernel keep on running and wasting GPU times. I just committed a previously successfully committed kernel by changing just pixel threshold. it was unsuccessfull and wasted all my GPU hours. Its very unfortunate for user like me who don't have personal GPU. Anybody facing same issue?</p>",
  "messages": [
    {
      "id": "666382",
      "postDate": "11/06/2019 03:51:22",
      "content": "<p>Few times, my inference kernel keep on running and wasting GPU times. I just committed a previously successfully committed kernel by changing just pixel threshold. it was unsuccessfull and wasted all my GPU hours. Its very unfortunate for user like me who don't have personal GPU. Anybody facing same issue?</p>",
      "rawMarkdown": "Few times, my inference kernel keep on running and wasting GPU times. I just committed a previously successfully committed kernel by changing just pixel threshold. it was unsuccessfull and wasted all my GPU hours. Its very unfortunate for user like me who don't have personal GPU. Anybody facing same issue?",
      "votes": null
    },
    {
      "id": "666387",
      "postDate": "11/06/2019 03:56:45",
      "content": "<p>last night i was training this model : <a href=\"https://www.kaggle.com/mobassir/deeplabv3-with-mobilenetv2\">https://www.kaggle.com/mobassir/deeplabv3-with-mobilenetv2</a>\ni was training for 1 epoch only,for 1 epoch it shouldn't even take 1 hour to complete commit but i waited for forever and my 11 hours gpu has gone,absolutely embarrassing :(</p>",
      "rawMarkdown": "last night i was training this model : https://www.kaggle.com/mobassir/deeplabv3-with-mobilenetv2\ni was training for 1 epoch only,for 1 epoch it shouldn't even take 1 hour to complete commit but i waited for forever and my 11 hours gpu has gone,absolutely embarrassing :(",
      "votes": null
    },
    {
      "id": "666389",
      "postDate": "11/06/2019 04:00:46",
      "content": "<p><a href=\"/raghaw\">@raghaw</a> Remember to turn off your kernel right after committing it, and you will not waste any GPU time 😄 </p>",
      "rawMarkdown": "raghaw Remember to turn off your kernel right after committing it, and you will not waste any GPU time 😄",
      "votes": null
    },
    {
      "id": "666393",
      "postDate": "11/06/2019 04:05:52",
      "content": "<p>Oh, you mean your submitted kernel still runs for many hours even it is unsuccessful to commit? I never face this problem before, but I think \"try-except\" could help 😉 </p>",
      "rawMarkdown": "Oh, you mean your submitted kernel still runs for many hours even it is unsuccessful to commit? I never face this problem before, but I think \"try-except\" could help 😉",
      "votes": null
    },
    {
      "id": "666400",
      "postDate": "11/06/2019 04:32:38",
      "content": "<p><a href=\"/phunghieu\">@phunghieu</a> <a href=\"/mobassir\">@mobassir</a> Yes, I think so. Its my habit to turn off GPU and interactive kernel after commit and I did the same. Before sleeping I committed a already successfully committed inference kernel by just changing pixel threshold when I woke up after 4 hours I saw that commit was still running, it should have completed within 25 minutes. At that time I cancelled the commit because it would have wasted my another 5 hours before timeout. It happened to me 3 times and I wasted my 10 GPU hours.</p>",
      "rawMarkdown": "phunghieu @mobassir Yes, I think so. Its my habit to turn off GPU and interactive kernel after commit and I did the same. Before sleeping I committed a already successfully committed inference kernel by just changing pixel threshold when I woke up after 4 hours I saw that commit was still running, it should have completed within 25 minutes. At that time I cancelled the commit because it would have wasted my another 5 hours before timeout. It happened to me 3 times and I wasted my 10 GPU hours.",
      "votes": null
    },
    {
      "id": "666402",
      "postDate": "11/06/2019 04:35:05",
      "content": "<p>same here with  me : <a href=\"https://www.kaggle.com/product-feedback/115916#latest-666401\">https://www.kaggle.com/product-feedback/115916#latest-666401</a></p>\n\n<p>:'(</p>",
      "rawMarkdown": "same here with  me : https://www.kaggle.com/product-feedback/115916#latest-666401\n\n:'(",
      "votes": null
    },
    {
      "id": "666422",
      "postDate": "11/06/2019 04:52:28",
      "content": "<p>Virtual machines are something that really hard to understand 😂 </p>",
      "rawMarkdown": "Virtual machines are something that really hard to understand 😂",
      "votes": null
    },
    {
      "id": "666429",
      "postDate": "11/06/2019 05:02:02",
      "content": "<p>Yes, the same thing happened to me today. I hit commit and it should have taken 30 minutes.. Luckily I noticed after 2 hours that it was taking too long and ended it. Then I ran it again with no changes and it finished in 30 minutes.</p>",
      "rawMarkdown": "Yes, the same thing happened to me today. I hit commit and it should have taken 30 minutes.. Luckily I noticed after 2 hours that it was taking too long and ended it. Then I ran it again with no changes and it finished in 30 minutes.",
      "votes": null
    },
    {
      "id": "666530",
      "postDate": "11/06/2019 07:28:43",
      "content": "<p><a href=\"/raghaw\">@raghaw</a> : Yeah. I also got the same issue in Lyft competition this morning. It normally takes 20 minutes for the commit to complete. But it took 3 hrs this morning and it didn't complete. I ended the session and re-running it. </p>",
      "rawMarkdown": "raghaw : Yeah. I also got the same issue in Lyft competition this morning. It normally takes 20 minutes for the commit to complete. But it took 3 hrs this morning and it didn't complete. I ended the session and re-running it.",
      "votes": null
    },
    {
      "id": "666532",
      "postDate": "11/06/2019 07:30:24",
      "content": "<p><a href=\"/phunghieu\">@phunghieu</a> : There should be a competition for that as well :). Predicting if virtual machines are in good condition or in bad condition. Whether we can use GPU to compute or not.  </p>",
      "rawMarkdown": "phunghieu : There should be a competition for that as well :). Predicting if virtual machines are in good condition or in bad condition. Whether we can use GPU to compute or not.",
      "votes": null
    },
    {
      "id": "666567",
      "postDate": "11/06/2019 08:27:34",
      "content": "<p><a href=\"/manojprabhaakr\">@manojprabhaakr</a> That sounds good 🎶 </p>",
      "rawMarkdown": "manojprabhaakr That sounds good 🎶",
      "votes": null
    },
    {
      "id": "666668",
      "postDate": "11/06/2019 11:01:34",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> <a href=\"/mobassir\">@mobassir</a> <a href=\"/phunghieu\">@phunghieu</a> <a href=\"/manojprabhaakr\">@manojprabhaakr</a>  Absolutely horrible experience!!!! In 7 hours, I run 7 inference kernels out of which 4 failed and keep on consuming GPU hours so I aborted them. So I could make just 3 submissions today and I don't have patience to try further. If it continue to do so .... I will just leave this competition!!!!</p>",
      "rawMarkdown": "cdeotte @mobassir @phunghieu @manojprabhaakr  Absolutely horrible experience!!!! In 7 hours, I run 7 inference kernels out of which 4 failed and keep on consuming GPU hours so I aborted them. So I could make just 3 submissions today and I don't have patience to try further. If it continue to do so .... I will just leave this competition!!!!",
      "votes": null
    },
    {
      "id": "666671",
      "postDate": "11/06/2019 11:06:58",
      "content": "<p>The same thing happened to me. Now, the total usage is still growing without kernel running😂 </p>",
      "rawMarkdown": "The same thing happened to me. Now, the total usage is still growing without kernel running😂",
      "votes": null
    },
    {
      "id": "666676",
      "postDate": "11/06/2019 11:13:10",
      "content": "<p>30 hours have been used up...</p>",
      "rawMarkdown": "30 hours have been used up...",
      "votes": null
    },
    {
      "id": "666687",
      "postDate": "11/06/2019 11:28:05",
      "content": "<p>The organizers should resolve this issue. When there is a constraint for GPU (30 hrs) and these kind of things happen, competitors who rely only on Kaggle GPU's will eventually leave the competition and only those who have own GPU's compete and get the medals/ranks etc.</p>",
      "rawMarkdown": "The organizers should resolve this issue. When there is a constraint for GPU (30 hrs) and these kind of things happen, competitors who rely only on Kaggle GPU's will eventually leave the competition and only those who have own GPU's compete and get the medals/ranks etc.",
      "votes": null
    },
    {
      "id": "666688",
      "postDate": "11/06/2019 11:29:00",
      "content": "<p><a href=\"/dandingclam\">@dandingclam</a> I think you might have left interactive kernel with GPU keep on running. Since Interactive kernels GPU hours are counted in 30 hrs, It might have consumed your GPU hours even though you were not running any commit.</p>",
      "rawMarkdown": "dandingclam I think you might have left interactive kernel with GPU keep on running. Since Interactive kernels GPU hours are counted in 30 hrs, It might have consumed your GPU hours even though you were not running any commit.",
      "votes": null
    },
    {
      "id": "666761",
      "postDate": "11/06/2019 13:20:53",
      "content": "<p>I committed a kernel this morning, and it should take less than 5 hours to train. Now it already spend 8 hours..I don't know should I stop it or not lol</p>",
      "rawMarkdown": "I committed a kernel this morning, and it should take less than 5 hours to train. Now it already spend 8 hours..I don't know should I stop it or not lol",
      "votes": null
    },
    {
      "id": "666859",
      "postDate": "11/06/2019 15:07:11",
      "content": "<p>Yes it is very frustrating. Furthermore, I am trying to make a public kernel to share and teach others. The errors used up all my GPU quota. Now I can't share nor experiment for the competition.</p>",
      "rawMarkdown": "Yes it is very frustrating. Furthermore, I am trying to make a public kernel to share and teach others. The errors used up all my GPU quota. Now I can't share nor experiment for the competition.",
      "votes": null
    },
    {
      "id": "666865",
      "postDate": "11/06/2019 15:12:19",
      "content": "<p>I wonder if a per challenge quota is better than per week, maybe with some top ups for submissions. No doubt this is being discussed elsewhere on Kaggle forums...</p>",
      "rawMarkdown": "I wonder if a per challenge quota is better than per week, maybe with some top ups for submissions. No doubt this is being discussed elsewhere on Kaggle forums...",
      "votes": null
    },
    {
      "id": "666868",
      "postDate": "11/06/2019 15:15:13",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> Despite 30Hrs GPU limit and these frustrating errors, you are making a public kernel to teach others. Hats off for you man!!!</p>",
      "rawMarkdown": "cdeotte Despite 30Hrs GPU limit and these frustrating errors, you are making a public kernel to teach others. Hats off for you man!!!",
      "votes": null
    },
    {
      "id": "667045",
      "postDate": "11/06/2019 18:54:29",
      "content": "<p>i wasted 18 hours this week , this exact same way..and not getting into kaggle anymore..</p>",
      "rawMarkdown": "i wasted 18 hours this week , this exact same way..and not getting into kaggle anymore..",
      "votes": null
    },
    {
      "id": "667077",
      "postDate": "11/06/2019 19:38:55",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> </p>\n\n<p>Yeah, almost everybody now thinking that no reason to make some public kernel for teaching purposes or kernel medals (which anyway teaches/helps others to have code to start on). </p>\n\n<p>Hence, the educational value of Kaggle is vastly decreasing. </p>",
      "rawMarkdown": "cdeotte \n\nYeah, almost everybody now thinking that no reason to make some public kernel for teaching purposes or kernel medals (which anyway teaches/helps others to have code to start on). \n\nHence, the educational value of Kaggle is vastly decreasing.",
      "votes": null
    },
    {
      "id": "667078",
      "postDate": "11/06/2019 19:40:03",
      "content": "<p><a href=\"/manojprabhaakr\">@manojprabhaakr</a> </p>\n\n<p>I ended up with that twice with 2-2.5h each. And I think it happens again in a new run of my kernel with another loss. So, in 30h I lost 25% for nothing. </p>",
      "rawMarkdown": "manojprabhaakr \n\nI ended up with that twice with 2-2.5h each. And I think it happens again in a new run of my kernel with another loss. So, in 30h I lost 25% for nothing.",
      "votes": null
    },
    {
      "id": "667079",
      "postDate": "11/06/2019 19:42:29",
      "content": "<p><a href=\"/mobassir\">@mobassir</a> another issue is when your code computes everything and doesn't stop, and GPU usage last f forever until reaches the limit. That's why I can not go to sleep just running kernel of 2-3 h GPU usage because when I wake up at the morning, 7-8h of GPU will be lost)))))))) <a href=\"/manojprabhaakr\">@manojprabhaakr</a> happened with you?</p>",
      "rawMarkdown": "mobassir another issue is when your code computes everything and doesn't stop, and GPU usage last f forever until reaches the limit. That's why I can not go to sleep just running kernel of 2-3 h GPU usage because when I wake up at the morning, 7-8h of GPU will be lost)))))))) @manojprabhaakr happened with you?",
      "votes": null
    },
    {
      "id": "667296",
      "postDate": "11/07/2019 04:54:26",
      "content": "<p>yes happened with me also :( </p>",
      "rawMarkdown": "yes happened with me also :(",
      "votes": null
    },
    {
      "id": "667327",
      "postDate": "11/07/2019 05:57:25",
      "content": "<p>Today's kernel seems back to normal. But I don't have gpu time to verify it. Lol</p>",
      "rawMarkdown": "Today's kernel seems back to normal. But I don't have gpu time to verify it. Lol",
      "votes": null
    },
    {
      "id": "667779",
      "postDate": "11/07/2019 17:17:54",
      "content": "<p>Plus you are double punished. If a notebook commit fails and runs extra long, then (1) you waste GPU quota and (2) your notebook's Kaggle Hottest score decreases. </p>\n\n<p>The hottest algorithm penalizes notebooks that are run without changes. My latest shared notebook <a href=\"https://www.kaggle.com/cdeotte/unsupervised-masks-cv-0-60\">here</a> is an example. Version 1 failed and I ran it a second time without changes. It currently has 14 votes in one day, but since public notebooks are sorted by \"hottest\" you don't see it when you click to see Cloud public notebooks. (It is at the bottom of the list). </p>\n\n<p>Therefore if you attempt to share something and GPU fails, you waste quota and people won't find your share anyway.</p>",
      "rawMarkdown": "Plus you are double punished. If a notebook commit fails and runs extra long, then (1) you waste GPU quota and (2) your notebook's Kaggle Hottest score decreases. \n  \nThe hottest algorithm penalizes notebooks that are run without changes. My latest shared notebook [here][1] is an example. Version 1 failed and I ran it a second time without changes. It currently has 14 votes in one day, but since public notebooks are sorted by \"hottest\" you don't see it when you click to see Cloud public notebooks. (It is at the bottom of the list). \n\nTherefore if you attempt to share something and GPU fails, you waste quota and people won't find your share anyway.\n\n[1]: https://www.kaggle.com/cdeotte/unsupervised-masks-cv-0-60",
      "votes": null
    },
    {
      "id": "667809",
      "postDate": "11/07/2019 17:53:45",
      "content": "<p>i didn't know that failed kernel version causes reduced hotness,thanks for the wisdom <a href=\"/cdeotte\">@cdeotte</a> </p>",
      "rawMarkdown": "i didn't know that failed kernel version causes reduced hotness,thanks for the wisdom @cdeotte",
      "votes": null
    },
    {
      "id": "668164",
      "postDate": "11/08/2019 04:19:32",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> <a href=\"/mobassir\">@mobassir</a> <a href=\"/manojprabhaakr\">@manojprabhaakr</a> <a href=\"/xiejialun\">@xiejialun</a> Seems back to normal. today I run 4 inference kernel, none of them failed.</p>",
      "rawMarkdown": "cdeotte @mobassir @manojprabhaakr @xiejialun Seems back to normal. today I run 4 inference kernel, none of them failed.",
      "votes": null
    },
    {
      "id": "668251",
      "postDate": "11/08/2019 07:03:36",
      "content": "<p>One of my kernel failed again... I commit the kernel after I leave the above comment and it spends rest of my GPU time. I wasted almost 15 hours of GPU time this week on this problem and can't even finish train one model.....This is very disappointed...</p>",
      "rawMarkdown": "One of my kernel failed again... I commit the kernel after I leave the above comment and it spends rest of my GPU time. I wasted almost 15 hours of GPU time this week on this problem and can't even finish train one model.....This is very disappointed...",
      "votes": null
    },
    {
      "id": "668492",
      "postDate": "11/08/2019 13:36:02",
      "content": "<p>same happened to me..😢</p>",
      "rawMarkdown": "same happened to me..😢",
      "votes": null
    },
    {
      "id": "757365",
      "postDate": "02/26/2020 17:22:41",
      "content": "<p>That's why you need to create separately training notebook and prediction/submission notebook. Try not to submit on the same training notebook as it consume a lot of time. Just create a notebook for training, get the trained weight, up load it to another notebook and make prediction and submission. It's usually take a few minute for committing the prediction notebook.</p>",
      "rawMarkdown": "That's why you need to create separately training notebook and prediction/submission notebook. Try not to submit on the same training notebook as it consume a lot of time. Just create a notebook for training, get the trained weight, up load it to another notebook and make prediction and submission. It's usually take a few minute for committing the prediction notebook.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 666387,
      "author_name": "mobassir",
      "author_url": "",
      "post_date": "11/06/2019 03:56:45",
      "content": "<p>last night i was training this model : <a href=\"https://www.kaggle.com/mobassir/deeplabv3-with-mobilenetv2\">https://www.kaggle.com/mobassir/deeplabv3-with-mobilenetv2</a>\ni was training for 1 epoch only,for 1 epoch it shouldn't even take 1 hour to complete commit but i waited for forever and my 11 hours gpu has gone,absolutely embarrassing :(</p>",
      "votes": null,
      "replies": [
        {
          "id": 667079,
          "author_name": "muhakabartay",
          "author_url": "",
          "post_date": "11/06/2019 19:42:29",
          "content": "<p><a href=\"/mobassir\">@mobassir</a> another issue is when your code computes everything and doesn't stop, and GPU usage last f forever until reaches the limit. That's why I can not go to sleep just running kernel of 2-3 h GPU usage because when I wake up at the morning, 7-8h of GPU will be lost)))))))) <a href=\"/manojprabhaakr\">@manojprabhaakr</a> happened with you?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 667296,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "11/07/2019 04:54:26",
          "content": "<p>yes happened with me also :( </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 666389,
      "author_name": "phunghieu",
      "author_url": "",
      "post_date": "11/06/2019 04:00:46",
      "content": "<p><a href=\"/raghaw\">@raghaw</a> Remember to turn off your kernel right after committing it, and you will not waste any GPU time 😄 </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 666393,
      "author_name": "phunghieu",
      "author_url": "",
      "post_date": "11/06/2019 04:05:52",
      "content": "<p>Oh, you mean your submitted kernel still runs for many hours even it is unsuccessful to commit? I never face this problem before, but I think \"try-except\" could help 😉 </p>",
      "votes": null,
      "replies": [
        {
          "id": 666400,
          "author_name": "raghaw",
          "author_url": "",
          "post_date": "11/06/2019 04:32:38",
          "content": "<p><a href=\"/phunghieu\">@phunghieu</a> <a href=\"/mobassir\">@mobassir</a> Yes, I think so. Its my habit to turn off GPU and interactive kernel after commit and I did the same. Before sleeping I committed a already successfully committed inference kernel by just changing pixel threshold when I woke up after 4 hours I saw that commit was still running, it should have completed within 25 minutes. At that time I cancelled the commit because it would have wasted my another 5 hours before timeout. It happened to me 3 times and I wasted my 10 GPU hours.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 666402,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "11/06/2019 04:35:05",
          "content": "<p>same here with  me : <a href=\"https://www.kaggle.com/product-feedback/115916#latest-666401\">https://www.kaggle.com/product-feedback/115916#latest-666401</a></p>\n\n<p>:'(</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 666422,
          "author_name": "phunghieu",
          "author_url": "",
          "post_date": "11/06/2019 04:52:28",
          "content": "<p>Virtual machines are something that really hard to understand 😂 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 666532,
          "author_name": "manojprabhaakr",
          "author_url": "",
          "post_date": "11/06/2019 07:30:24",
          "content": "<p><a href=\"/phunghieu\">@phunghieu</a> : There should be a competition for that as well :). Predicting if virtual machines are in good condition or in bad condition. Whether we can use GPU to compute or not.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 666567,
          "author_name": "phunghieu",
          "author_url": "",
          "post_date": "11/06/2019 08:27:34",
          "content": "<p><a href=\"/manojprabhaakr\">@manojprabhaakr</a> That sounds good 🎶 </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 666429,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "11/06/2019 05:02:02",
      "content": "<p>Yes, the same thing happened to me today. I hit commit and it should have taken 30 minutes.. Luckily I noticed after 2 hours that it was taking too long and ended it. Then I ran it again with no changes and it finished in 30 minutes.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 666530,
      "author_name": "manojprabhaakr",
      "author_url": "",
      "post_date": "11/06/2019 07:28:43",
      "content": "<p><a href=\"/raghaw\">@raghaw</a> : Yeah. I also got the same issue in Lyft competition this morning. It normally takes 20 minutes for the commit to complete. But it took 3 hrs this morning and it didn't complete. I ended the session and re-running it. </p>",
      "votes": null,
      "replies": [
        {
          "id": 667078,
          "author_name": "muhakabartay",
          "author_url": "",
          "post_date": "11/06/2019 19:40:03",
          "content": "<p><a href=\"/manojprabhaakr\">@manojprabhaakr</a> </p>\n\n<p>I ended up with that twice with 2-2.5h each. And I think it happens again in a new run of my kernel with another loss. So, in 30h I lost 25% for nothing. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 666668,
      "author_name": "raghaw",
      "author_url": "",
      "post_date": "11/06/2019 11:01:34",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> <a href=\"/mobassir\">@mobassir</a> <a href=\"/phunghieu\">@phunghieu</a> <a href=\"/manojprabhaakr\">@manojprabhaakr</a>  Absolutely horrible experience!!!! In 7 hours, I run 7 inference kernels out of which 4 failed and keep on consuming GPU hours so I aborted them. So I could make just 3 submissions today and I don't have patience to try further. If it continue to do so .... I will just leave this competition!!!!</p>",
      "votes": null,
      "replies": [
        {
          "id": 666687,
          "author_name": "manojprabhaakr",
          "author_url": "",
          "post_date": "11/06/2019 11:28:05",
          "content": "<p>The organizers should resolve this issue. When there is a constraint for GPU (30 hrs) and these kind of things happen, competitors who rely only on Kaggle GPU's will eventually leave the competition and only those who have own GPU's compete and get the medals/ranks etc.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 666859,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "11/06/2019 15:07:11",
          "content": "<p>Yes it is very frustrating. Furthermore, I am trying to make a public kernel to share and teach others. The errors used up all my GPU quota. Now I can't share nor experiment for the competition.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 666868,
          "author_name": "raghaw",
          "author_url": "",
          "post_date": "11/06/2019 15:15:13",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> Despite 30Hrs GPU limit and these frustrating errors, you are making a public kernel to teach others. Hats off for you man!!!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 667077,
          "author_name": "muhakabartay",
          "author_url": "",
          "post_date": "11/06/2019 19:38:55",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> </p>\n\n<p>Yeah, almost everybody now thinking that no reason to make some public kernel for teaching purposes or kernel medals (which anyway teaches/helps others to have code to start on). </p>\n\n<p>Hence, the educational value of Kaggle is vastly decreasing. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 667779,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "11/07/2019 17:17:54",
          "content": "<p>Plus you are double punished. If a notebook commit fails and runs extra long, then (1) you waste GPU quota and (2) your notebook's Kaggle Hottest score decreases. </p>\n\n<p>The hottest algorithm penalizes notebooks that are run without changes. My latest shared notebook <a href=\"https://www.kaggle.com/cdeotte/unsupervised-masks-cv-0-60\">here</a> is an example. Version 1 failed and I ran it a second time without changes. It currently has 14 votes in one day, but since public notebooks are sorted by \"hottest\" you don't see it when you click to see Cloud public notebooks. (It is at the bottom of the list). </p>\n\n<p>Therefore if you attempt to share something and GPU fails, you waste quota and people won't find your share anyway.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 667809,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "11/07/2019 17:53:45",
          "content": "<p>i didn't know that failed kernel version causes reduced hotness,thanks for the wisdom <a href=\"/cdeotte\">@cdeotte</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 666671,
      "author_name": "dandingclam",
      "author_url": "",
      "post_date": "11/06/2019 11:06:58",
      "content": "<p>The same thing happened to me. Now, the total usage is still growing without kernel running😂 </p>",
      "votes": null,
      "replies": [
        {
          "id": 666676,
          "author_name": "dandingclam",
          "author_url": "",
          "post_date": "11/06/2019 11:13:10",
          "content": "<p>30 hours have been used up...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 666688,
          "author_name": "raghaw",
          "author_url": "",
          "post_date": "11/06/2019 11:29:00",
          "content": "<p><a href=\"/dandingclam\">@dandingclam</a> I think you might have left interactive kernel with GPU keep on running. Since Interactive kernels GPU hours are counted in 30 hrs, It might have consumed your GPU hours even though you were not running any commit.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 666761,
      "author_name": "xiejialun",
      "author_url": "",
      "post_date": "11/06/2019 13:20:53",
      "content": "<p>I committed a kernel this morning, and it should take less than 5 hours to train. Now it already spend 8 hours..I don't know should I stop it or not lol</p>",
      "votes": null,
      "replies": [
        {
          "id": 667045,
          "author_name": "phoenix9032",
          "author_url": "",
          "post_date": "11/06/2019 18:54:29",
          "content": "<p>i wasted 18 hours this week , this exact same way..and not getting into kaggle anymore..</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 667327,
          "author_name": "xiejialun",
          "author_url": "",
          "post_date": "11/07/2019 05:57:25",
          "content": "<p>Today's kernel seems back to normal. But I don't have gpu time to verify it. Lol</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 668164,
          "author_name": "raghaw",
          "author_url": "",
          "post_date": "11/08/2019 04:19:32",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> <a href=\"/mobassir\">@mobassir</a> <a href=\"/manojprabhaakr\">@manojprabhaakr</a> <a href=\"/xiejialun\">@xiejialun</a> Seems back to normal. today I run 4 inference kernel, none of them failed.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 668251,
          "author_name": "xiejialun",
          "author_url": "",
          "post_date": "11/08/2019 07:03:36",
          "content": "<p>One of my kernel failed again... I commit the kernel after I leave the above comment and it spends rest of my GPU time. I wasted almost 15 hours of GPU time this week on this problem and can't even finish train one model.....This is very disappointed...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 668492,
          "author_name": "manann",
          "author_url": "",
          "post_date": "11/08/2019 13:36:02",
          "content": "<p>same happened to me..😢</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 666865,
      "author_name": "robga",
      "author_url": "",
      "post_date": "11/06/2019 15:12:19",
      "content": "<p>I wonder if a per challenge quota is better than per week, maybe with some top ups for submissions. No doubt this is being discussed elsewhere on Kaggle forums...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 757365,
      "author_name": "phamhoad",
      "author_url": "",
      "post_date": "02/26/2020 17:22:41",
      "content": "<p>That's why you need to create separately training notebook and prediction/submission notebook. Try not to submit on the same training notebook as it consume a lot of time. Just create a notebook for training, get the trained weight, up load it to another notebook and make prediction and submission. It's usually take a few minute for committing the prediction notebook.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "666382": "Few times, my inference kernel keep on running and wasting GPU times. I just committed a previously successfully committed kernel by changing just pixel threshold. it was unsuccessfull and wasted all my GPU hours. Its very unfortunate for user like me who don't have personal GPU. Anybody facing same issue?",
    "666387": "last night i was training this model : https://www.kaggle.com/mobassir/deeplabv3-with-mobilenetv2\ni was training for 1 epoch only,for 1 epoch it shouldn't even take 1 hour to complete commit but i waited for forever and my 11 hours gpu has gone,absolutely embarrassing :(",
    "666389": "raghaw Remember to turn off your kernel right after committing it, and you will not waste any GPU time 😄",
    "666393": "Oh, you mean your submitted kernel still runs for many hours even it is unsuccessful to commit? I never face this problem before, but I think \"try-except\" could help 😉",
    "666400": "phunghieu @mobassir Yes, I think so. Its my habit to turn off GPU and interactive kernel after commit and I did the same. Before sleeping I committed a already successfully committed inference kernel by just changing pixel threshold when I woke up after 4 hours I saw that commit was still running, it should have completed within 25 minutes. At that time I cancelled the commit because it would have wasted my another 5 hours before timeout. It happened to me 3 times and I wasted my 10 GPU hours.",
    "666402": "same here with  me : https://www.kaggle.com/product-feedback/115916#latest-666401\n\n:'(",
    "666422": "Virtual machines are something that really hard to understand 😂",
    "666429": "Yes, the same thing happened to me today. I hit commit and it should have taken 30 minutes.. Luckily I noticed after 2 hours that it was taking too long and ended it. Then I ran it again with no changes and it finished in 30 minutes.",
    "666530": "raghaw : Yeah. I also got the same issue in Lyft competition this morning. It normally takes 20 minutes for the commit to complete. But it took 3 hrs this morning and it didn't complete. I ended the session and re-running it.",
    "666532": "phunghieu : There should be a competition for that as well :). Predicting if virtual machines are in good condition or in bad condition. Whether we can use GPU to compute or not.",
    "666567": "manojprabhaakr That sounds good 🎶",
    "666668": "cdeotte @mobassir @phunghieu @manojprabhaakr  Absolutely horrible experience!!!! In 7 hours, I run 7 inference kernels out of which 4 failed and keep on consuming GPU hours so I aborted them. So I could make just 3 submissions today and I don't have patience to try further. If it continue to do so .... I will just leave this competition!!!!",
    "666671": "The same thing happened to me. Now, the total usage is still growing without kernel running😂",
    "666676": "30 hours have been used up...",
    "666687": "The organizers should resolve this issue. When there is a constraint for GPU (30 hrs) and these kind of things happen, competitors who rely only on Kaggle GPU's will eventually leave the competition and only those who have own GPU's compete and get the medals/ranks etc.",
    "666688": "dandingclam I think you might have left interactive kernel with GPU keep on running. Since Interactive kernels GPU hours are counted in 30 hrs, It might have consumed your GPU hours even though you were not running any commit.",
    "666761": "I committed a kernel this morning, and it should take less than 5 hours to train. Now it already spend 8 hours..I don't know should I stop it or not lol",
    "666859": "Yes it is very frustrating. Furthermore, I am trying to make a public kernel to share and teach others. The errors used up all my GPU quota. Now I can't share nor experiment for the competition.",
    "666865": "I wonder if a per challenge quota is better than per week, maybe with some top ups for submissions. No doubt this is being discussed elsewhere on Kaggle forums...",
    "666868": "cdeotte Despite 30Hrs GPU limit and these frustrating errors, you are making a public kernel to teach others. Hats off for you man!!!",
    "667045": "i wasted 18 hours this week , this exact same way..and not getting into kaggle anymore..",
    "667077": "cdeotte \n\nYeah, almost everybody now thinking that no reason to make some public kernel for teaching purposes or kernel medals (which anyway teaches/helps others to have code to start on). \n\nHence, the educational value of Kaggle is vastly decreasing.",
    "667078": "manojprabhaakr \n\nI ended up with that twice with 2-2.5h each. And I think it happens again in a new run of my kernel with another loss. So, in 30h I lost 25% for nothing.",
    "667079": "mobassir another issue is when your code computes everything and doesn't stop, and GPU usage last f forever until reaches the limit. That's why I can not go to sleep just running kernel of 2-3 h GPU usage because when I wake up at the morning, 7-8h of GPU will be lost)))))))) @manojprabhaakr happened with you?",
    "667296": "yes happened with me also :(",
    "667327": "Today's kernel seems back to normal. But I don't have gpu time to verify it. Lol",
    "667779": "Plus you are double punished. If a notebook commit fails and runs extra long, then (1) you waste GPU quota and (2) your notebook's Kaggle Hottest score decreases. \n  \nThe hottest algorithm penalizes notebooks that are run without changes. My latest shared notebook [here][1] is an example. Version 1 failed and I ran it a second time without changes. It currently has 14 votes in one day, but since public notebooks are sorted by \"hottest\" you don't see it when you click to see Cloud public notebooks. (It is at the bottom of the list). \n\nTherefore if you attempt to share something and GPU fails, you waste quota and people won't find your share anyway.\n\n[1]: https://www.kaggle.com/cdeotte/unsupervised-masks-cv-0-60",
    "667809": "i didn't know that failed kernel version causes reduced hotness,thanks for the wisdom @cdeotte",
    "668164": "cdeotte @mobassir @manojprabhaakr @xiejialun Seems back to normal. today I run 4 inference kernel, none of them failed.",
    "668251": "One of my kernel failed again... I commit the kernel after I leave the above comment and it spends rest of my GPU time. I wasted almost 15 hours of GPU time this week on this problem and can't even finish train one model.....This is very disappointed...",
    "668492": "same happened to me..😢",
    "757365": "That's why you need to create separately training notebook and prediction/submission notebook. Try not to submit on the same training notebook as it consume a lot of time. Just create a notebook for training, get the trained weight, up load it to another notebook and make prediction and submission. It's usually take a few minute for committing the prediction notebook."
  },
  "source": "meta"
}