{
  "id": 144172,
  "title": "Beware: Misleading status message with TPU notebook commit",
  "url": "/competitions/flower-classification-with-tpus/discussion/144172",
  "author_name": "",
  "post_date": "2020-04-18T01:10:46.129128Z",
  "votes": null,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Every time I click \"Save and Commit\" on my notebook, my notebook shows an error message \"Are you still there?  Your TPU notebook stops after 15 minutes of idle time. Commit your notebook for longer computations\". It also shows session as \"Disconnected\".</p>\n\n<p>I know that commit runs the whole notebook and my training within the notebook takes ~ 30-35 minutes. But after waiting for about 40 minutes, my notebook gets committed. </p>",
  "messages": [
    {
      "id": "811488",
      "postDate": "04/18/2020 01:10:46",
      "content": "<p>Every time I click \"Save and Commit\" on my notebook, my notebook shows an error message \"Are you still there?  Your TPU notebook stops after 15 minutes of idle time. Commit your notebook for longer computations\". It also shows session as \"Disconnected\".</p>\n\n<p>I know that commit runs the whole notebook and my training within the notebook takes ~ 30-35 minutes. But after waiting for about 40 minutes, my notebook gets committed. </p>",
      "rawMarkdown": "Every time I click \"Save and Commit\" on my notebook, my notebook shows an error message \"Are you still there?  Your TPU notebook stops after 15 minutes of idle time. Commit your notebook for longer computations\". It also shows session as \"Disconnected\".\n\nI know that commit runs the whole notebook and my training within the notebook takes ~ 30-35 minutes. But after waiting for about 40 minutes, my notebook gets committed.",
      "votes": null
    },
    {
      "id": "811574",
      "postDate": "04/18/2020 04:39:05",
      "content": "<p>When you save and commit your notebook, it runs in a separate batch session in the background, and that's where the current version of it is committed. The message you see is for the current interactive session that you began using and editing. The kernel that was running as you made your edits stopped after 15 minutes of idle time while the batch process continued in the background. I hope this helps.</p>",
      "rawMarkdown": "When you save and commit your notebook, it runs in a separate batch session in the background, and that's where the current version of it is committed. The message you see is for the current interactive session that you began using and editing. The kernel that was running as you made your edits stopped after 15 minutes of idle time while the batch process continued in the background. I hope this helps.",
      "votes": null
    },
    {
      "id": "811822",
      "postDate": "04/18/2020 08:58:11",
      "content": "<p>Can i stop my notebook after committing so as to save TPU quota?</p>",
      "rawMarkdown": "Can i stop my notebook after committing so as to save TPU quota?",
      "votes": null
    },
    {
      "id": "815597",
      "postDate": "04/21/2020 17:03:13",
      "content": "<p>^ This is correct. </p>\n\n<p>We're still working on explaining it effectively in the UI, but your interactive notebook is executing totally separate to 'commits/save versions'. The interactive one lets you write your notebook, executing any part of it as you go to test it and iterate on it. Once you've got a notebook that you'd like to execute  from a clean slate top-to-bottom, you can hit 'save version &gt; save &amp; run all' and it copies your interactive notebook and executes it on a background machine top-to-bottom and then saves and formats the results to be displayed in the 'viewer'.</p>\n\n<p>Unlike the interactive notebook, the background one doesn't time out if you're not using it. It only times out if you exceed the maximum runtime allowed.</p>",
      "rawMarkdown": "^ This is correct. \n\nWe're still working on explaining it effectively in the UI, but your interactive notebook is executing totally separate to 'commits/save versions'. The interactive one lets you write your notebook, executing any part of it as you go to test it and iterate on it. Once you've got a notebook that you'd like to execute  from a clean slate top-to-bottom, you can hit 'save version &gt; save &amp; run all' and it copies your interactive notebook and executes it on a background machine top-to-bottom and then saves and formats the results to be displayed in the 'viewer'.\n\nUnlike the interactive notebook, the background one doesn't time out if you're not using it. It only times out if you exceed the maximum runtime allowed.",
      "votes": null
    },
    {
      "id": "815602",
      "postDate": "04/21/2020 17:08:41",
      "content": "<p>Hi, i am trying to commit my notebook, but since last one hour same message showing GCE check failed is being displayed. I left my notebook as it is. Next day i saw commit operation was failed.\nPlease help me as i am unable to commit my version.\nIt failed with following log messages. At time 10166 error is coming but i have no idea about it. Please guide me <a href=\"/martingorner\">@martingorner</a></p>\n\n<p>Time\nLine #\nLog Message\n3.0s\n1\n[NbConvertApp] Converting notebook notebook.ipynb to notebook\n6.1s\n2\n[NbConvertApp] Executing notebook with kernel: python3\n14.3s\n3\n2020-04-20 17:18:06.289469: I tensorflow/core/platform/cpufeatureguard.cc:142] Your CPU supports instructions that this TensorFlow binary was not compiled to use: AVX512F\n14.3s\n4\n2020-04-20 17:18:06.314414: I tensorflow/core/platform/profileutils/cpuutils.cc:94] CPU Frequency: 2000179999 Hz\n14.3s\n5\n2020-04-20 17:18:06.315135: I tensorflow/compiler/xla/service/service.cc:168] XLA service 0x558c1cfbef00 initialized for platform Host (this does not guarantee that XLA will be used). Devices:\n14.3s\n6\n2020-04-20 17:18:06.315181: I tensorflow/compiler/xla/service/service.cc:176] StreamExecutor device (0): Host, Default Version\n14.3s\n7\n2020-04-20 17:18:06.344279: I tensorflow/core/distributedruntime/rpc/grpcchannel.cc:300] Initialize GrpcChannelCache for job worker -&gt; {0 -&gt; 10.0.0.2:8470}\n14.3s\n8\n2020-04-20 17:18:06.344339: I tensorflow/core/distributedruntime/rpc/grpcchannel.cc:300] Initialize GrpcChannelCache for job localhost -&gt; {0 -&gt; localhost:30013}\n14.3s\n9\n2020-04-20 17:18:06.359935: I tensorflow/core/distributedruntime/rpc/grpcchannel.cc:300] Initialize GrpcChannelCache for job worker -&gt; {0 -&gt; 10.0.0.2:8470}\n14.3s\n10\n2020-04-20 17:18:06.359995: I tensorflow/core/distributedruntime/rpc/grpcchannel.cc:300] Initialize GrpcChannelCache for job localhost -&gt; {0 -&gt; localhost:30013}\n14.3s\n11\n2020-04-20 17:18:06.366237: I tensorflow/core/distributedruntime/rpc/grpcserverlib.cc:390] Started server with target: grpc://localhost:30013 21.6s 12 2020-04-20 17:18:13.653909: W tensorflow/core/platform/cloud/googleauthprovider.cc:178] All attempts to get a Google authentication bearer token failed, returning an empty token. Retrieving token from files failed with \"Not found: Could not locate the credentials file.\". Retrieving token from GCE failed with \"Cancelled: GCE check skipped due to presence of $NOGCECHECK environment variable.\". 21.7s 13 2020-04-20 17:18:13.704236: W tensorflow/core/platform/cloud/googleauthprovider.cc:178] All attempts to get a Google authentication bearer token failed, returning an empty token. Retrieving token from files failed with \"Not found: Could not locate the credentials file.\". Retrieving token from GCE failed with \"Cancelled: GCE check skipped due to presence of $NOGCECHECK environment variable.\". 21.7s 14 2020-04-20 17:18:13.741587: W tensorflow/core/platform/cloud/googleauthprovider.cc:178] All attempts to get a Google authentication bearer token failed, returning an empty token. Retrieving token from files failed with \"Not found: Could not locate the credentials file.\". Retrieving token from GCE failed with \"Cancelled: GCE check skipped due to presence of $NOGCECHECK environment variable.\". 10166.2s 15 2020-04-20 20:07:18.218870: W tensorflow/core/distributedruntime/eager/remotetensorhandledata.cc:75] Unable to destroy remote tensor handles. If you are running a tf.function, it usually indicates some op in the graph gets an error: RecvTensor expects a different device incarnation: 14797989987728605988 vs. 1799738036951955635. Your worker job (\"/job:tpuworker/replica:0/task:0\") was probably restarted. Check your worker job for the reason why it was restarted.\n10166.2s\n16\nAdditional GRPC error information:\n10166.2s\n17\n{\"created\":\"@1587413238.218478572\",\"description\":\"Error received from peer\",\"file\":\"external/grpc/src/core/lib/surface/call.cc\",\"fileline\":1039,\"grpcmessage\":\"RecvTensor expects a different device incarnation: 14797989987728605988 vs. 1799738036951955635. Your worker job (\"/job:tpuworker/replica:0/task:0\") was probably restarted. Check your worker job for the reason why it was restarted.\",\"grpcstatus\":10}\n10169.6s\n18\n2020-04-20 20:07:21.664715: W ./tensorflow/core/distributedruntime/eager/destroytensorhandlenode.h:79] Ignoring an error encountered when deleting remote tensors handles: Invalid argument: Unable to find a contextid matching the specified one (12193600234898793608). Perhaps the worker was restarted, or the context was GC'd? 10169.6s 19 Additional GRPC error information: 10169.6s 20 {\"created\":\"@1587413241.664555067\",\"description\":\"Error received from peer\",\"file\":\"external/grpc/src/core/lib/surface/call.cc\",\"fileline\":1039,\"grpcmessage\":\"Unable to find a contextid matching the specified one (12193600234898793608). Perhaps the worker was restarted, or the context was GC'd?\",\"grpc_status\":3}\n10180.0s\n21\n[NbConvertApp] Writing 158697 bytes to notebook.ipynb\n10180.9s\n22\n[NbConvertApp] Converting notebook notebook.ipynb to html\n10182.4s\n23\n[NbConvertApp] Writing 508223 bytes to results.html\n10182.4s\n24\n10182.4s\n26\nComplete. Exited with code 0.</p>",
      "rawMarkdown": "Hi, i am trying to commit my notebook, but since last one hour same message showing GCE check failed is being displayed. I left my notebook as it is. Next day i saw commit operation was failed.\nPlease help me as i am unable to commit my version.\nIt failed with following log messages. At time 10166 error is coming but i have no idea about it. Please guide me @martingorner\n\nTime\nLine #\nLog Message\n3.0s\n1\n[NbConvertApp] Converting notebook notebook.ipynb to notebook\n6.1s\n2\n[NbConvertApp] Executing notebook with kernel: python3\n14.3s\n3\n2020-04-20 17:18:06.289469: I tensorflow/core/platform/cpufeatureguard.cc:142] Your CPU supports instructions that this TensorFlow binary was not compiled to use: AVX512F\n14.3s\n4\n2020-04-20 17:18:06.314414: I tensorflow/core/platform/profileutils/cpuutils.cc:94] CPU Frequency: 2000179999 Hz\n14.3s\n5\n2020-04-20 17:18:06.315135: I tensorflow/compiler/xla/service/service.cc:168] XLA service 0x558c1cfbef00 initialized for platform Host (this does not guarantee that XLA will be used). Devices:\n14.3s\n6\n2020-04-20 17:18:06.315181: I tensorflow/compiler/xla/service/service.cc:176] StreamExecutor device (0): Host, Default Version\n14.3s\n7\n2020-04-20 17:18:06.344279: I tensorflow/core/distributedruntime/rpc/grpcchannel.cc:300] Initialize GrpcChannelCache for job worker -&gt; {0 -&gt; 10.0.0.2:8470}\n14.3s\n8\n2020-04-20 17:18:06.344339: I tensorflow/core/distributedruntime/rpc/grpcchannel.cc:300] Initialize GrpcChannelCache for job localhost -&gt; {0 -&gt; localhost:30013}\n14.3s\n9\n2020-04-20 17:18:06.359935: I tensorflow/core/distributedruntime/rpc/grpcchannel.cc:300] Initialize GrpcChannelCache for job worker -&gt; {0 -&gt; 10.0.0.2:8470}\n14.3s\n10\n2020-04-20 17:18:06.359995: I tensorflow/core/distributedruntime/rpc/grpcchannel.cc:300] Initialize GrpcChannelCache for job localhost -&gt; {0 -&gt; localhost:30013}\n14.3s\n11\n2020-04-20 17:18:06.366237: I tensorflow/core/distributedruntime/rpc/grpcserverlib.cc:390] Started server with target: grpc://localhost:30013 21.6s 12 2020-04-20 17:18:13.653909: W tensorflow/core/platform/cloud/googleauthprovider.cc:178] All attempts to get a Google authentication bearer token failed, returning an empty token. Retrieving token from files failed with \"Not found: Could not locate the credentials file.\". Retrieving token from GCE failed with \"Cancelled: GCE check skipped due to presence of $NOGCECHECK environment variable.\". 21.7s 13 2020-04-20 17:18:13.704236: W tensorflow/core/platform/cloud/googleauthprovider.cc:178] All attempts to get a Google authentication bearer token failed, returning an empty token. Retrieving token from files failed with \"Not found: Could not locate the credentials file.\". Retrieving token from GCE failed with \"Cancelled: GCE check skipped due to presence of $NOGCECHECK environment variable.\". 21.7s 14 2020-04-20 17:18:13.741587: W tensorflow/core/platform/cloud/googleauthprovider.cc:178] All attempts to get a Google authentication bearer token failed, returning an empty token. Retrieving token from files failed with \"Not found: Could not locate the credentials file.\". Retrieving token from GCE failed with \"Cancelled: GCE check skipped due to presence of $NOGCECHECK environment variable.\". 10166.2s 15 2020-04-20 20:07:18.218870: W tensorflow/core/distributedruntime/eager/remotetensorhandledata.cc:75] Unable to destroy remote tensor handles. If you are running a tf.function, it usually indicates some op in the graph gets an error: RecvTensor expects a different device incarnation: 14797989987728605988 vs. 1799738036951955635. Your worker job (\"/job:tpuworker/replica:0/task:0\") was probably restarted. Check your worker job for the reason why it was restarted.\n10166.2s\n16\nAdditional GRPC error information:\n10166.2s\n17\n{\"created\":\"@1587413238.218478572\",\"description\":\"Error received from peer\",\"file\":\"external/grpc/src/core/lib/surface/call.cc\",\"fileline\":1039,\"grpcmessage\":\"RecvTensor expects a different device incarnation: 14797989987728605988 vs. 1799738036951955635. Your worker job (\"/job:tpuworker/replica:0/task:0\") was probably restarted. Check your worker job for the reason why it was restarted.\",\"grpcstatus\":10}\n10169.6s\n18\n2020-04-20 20:07:21.664715: W ./tensorflow/core/distributedruntime/eager/destroytensorhandlenode.h:79] Ignoring an error encountered when deleting remote tensors handles: Invalid argument: Unable to find a contextid matching the specified one (12193600234898793608). Perhaps the worker was restarted, or the context was GC'd? 10169.6s 19 Additional GRPC error information: 10169.6s 20 {\"created\":\"@1587413241.664555067\",\"description\":\"Error received from peer\",\"file\":\"external/grpc/src/core/lib/surface/call.cc\",\"fileline\":1039,\"grpcmessage\":\"Unable to find a contextid matching the specified one (12193600234898793608). Perhaps the worker was restarted, or the context was GC'd?\",\"grpc_status\":3}\n10180.0s\n21\n[NbConvertApp] Writing 158697 bytes to notebook.ipynb\n10180.9s\n22\n[NbConvertApp] Converting notebook notebook.ipynb to html\n10182.4s\n23\n[NbConvertApp] Writing 508223 bytes to results.html\n10182.4s\n24\n10182.4s\n26\nComplete. Exited with code 0.",
      "votes": null
    },
    {
      "id": "815707",
      "postDate": "04/21/2020 18:55:44",
      "content": "<p>Yes definitely, you can. Just change settings in acceleator section, so no more interactive section is not using it.</p>",
      "rawMarkdown": "Yes definitely, you can. Just change settings in acceleator section, so no more interactive section is not using it.",
      "votes": null
    },
    {
      "id": "815728",
      "postDate": "04/21/2020 19:25:26",
      "content": "<p><a href=\"/pawanverma\">@pawanverma</a> this issue seems unrelated to the current thread, might get more responses creating your own post.</p>\n\n<p>I'm not sure I follow the issue here though, I'm wondering if maybe the code is consistently causing a problem on the TPU (maybe TPU OOM or something). I'm not an expert on TPU code though.</p>",
      "rawMarkdown": "pawanverma this issue seems unrelated to the current thread, might get more responses creating your own post.\n\nI'm not sure I follow the issue here though, I'm wondering if maybe the code is consistently causing a problem on the TPU (maybe TPU OOM or something). I'm not an expert on TPU code though.",
      "votes": null
    },
    {
      "id": "816426",
      "postDate": "04/22/2020 11:00:07",
      "content": "<p>Actually my question was that during committing since there is a separate session created and our notebook is compiled there so can we stop our interactive notebook so as to save TPU quota.</p>",
      "rawMarkdown": "Actually my question was that during committing since there is a separate session created and our notebook is compiled there so can we stop our interactive notebook so as to save TPU quota.",
      "votes": null
    },
    {
      "id": "816653",
      "postDate": "04/22/2020 13:48:05",
      "content": "<p>Yes you can stop the interactive notebook without stopping the commit to save quota. The background  commit will continue to run fine.</p>",
      "rawMarkdown": "Yes you can stop the interactive notebook without stopping the commit to save quota. The background  commit will continue to run fine.",
      "votes": null
    },
    {
      "id": "889838",
      "postDate": "06/17/2020 07:19:31",
      "content": "<p><a href=\"/herbison\">@herbison</a>  &amp; Kaggle Team,</p>\n\n<p>This means if switch on GPU/TPU and then commit -&gt; It will create a background session for committing. Hence, I can switch off the GPU/TPU for interactive mode.  If we don't, then we are using 2X of the resources - this is painfully not clear in the UI.</p>\n\n<p>This should be taken on priority by Kaggle to clarify on the UI - since a lot of people are just waitinf for it to commit and wating double the resources.</p>\n\n<p>Also, a message should be sent after the commit is complete, so that we can submit the solution. </p>\n\n<p>Thanks.</p>",
      "rawMarkdown": "herbison  &amp; Kaggle Team,\n\nThis means if switch on GPU/TPU and then commit -&gt; It will create a background session for committing. Hence, I can switch off the GPU/TPU for interactive mode.  If we don't, then we are using 2X of the resources - this is painfully not clear in the UI.\n\nThis should be taken on priority by Kaggle to clarify on the UI - since a lot of people are just waitinf for it to commit and wating double the resources.\n\nAlso, a message should be sent after the commit is complete, so that we can submit the solution. \n\nThanks.",
      "votes": null
    },
    {
      "id": "891874",
      "postDate": "06/18/2020 13:57:46",
      "content": "<p>I sent along your feedback, I agree it would be nice to show it clearer in the UI that it counts as an extra session, we're working on that. Any suggestions on what would help make that clearer to you?</p>\n\n<p>I agree about notification on finish too, passed it along.</p>",
      "rawMarkdown": "I sent along your feedback, I agree it would be nice to show it clearer in the UI that it counts as an extra session, we're working on that. Any suggestions on what would help make that clearer to you?\n\nI agree about notification on finish too, passed it along.",
      "votes": null
    },
    {
      "id": "892166",
      "postDate": "06/18/2020 17:37:05",
      "content": "<p>Hi <a href=\"/herbison\">@herbison</a>,</p>\n\n<p>A couple of ideas...</p>\n\n<p>A) As soon as I run -&gt; Save and Commit All -&gt; My current interactive session should \"force\" stop GPU session and go to None. If user tries to enable GPU a warning can pop up, that a background session of the commit is still running( IF it is running). </p>\n\n<p>Another idea is to \"always\" show GPU and TPU sessions running in totality on top of a notebook. Currently i have to go to notebooks to find this.</p>\n\n<p>B) Once the background finish commit happens, a notification can arrive in the same place notifications comeup for chat messages like this one.</p>\n\n<p>Hope that helps</p>",
      "rawMarkdown": "Hi @herbison,\n\nA couple of ideas...\n\nA) As soon as I run -&gt; Save and Commit All -&gt; My current interactive session should \"force\" stop GPU session and go to None. If user tries to enable GPU a warning can pop up, that a background session of the commit is still running( IF it is running). \n\nAnother idea is to \"always\" show GPU and TPU sessions running in totality on top of a notebook. Currently i have to go to notebooks to find this.\n\n\nB) Once the background finish commit happens, a notification can arrive in the same place notifications comeup for chat messages like this one.\n\nHope that helps",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 811574,
      "author_name": "michaele919",
      "author_url": "",
      "post_date": "04/18/2020 04:39:05",
      "content": "<p>When you save and commit your notebook, it runs in a separate batch session in the background, and that's where the current version of it is committed. The message you see is for the current interactive session that you began using and editing. The kernel that was running as you made your edits stopped after 15 minutes of idle time while the batch process continued in the background. I hope this helps.</p>",
      "votes": null,
      "replies": [
        {
          "id": 815597,
          "author_name": "herbison",
          "author_url": "",
          "post_date": "04/21/2020 17:03:13",
          "content": "<p>^ This is correct. </p>\n\n<p>We're still working on explaining it effectively in the UI, but your interactive notebook is executing totally separate to 'commits/save versions'. The interactive one lets you write your notebook, executing any part of it as you go to test it and iterate on it. Once you've got a notebook that you'd like to execute  from a clean slate top-to-bottom, you can hit 'save version &gt; save &amp; run all' and it copies your interactive notebook and executes it on a background machine top-to-bottom and then saves and formats the results to be displayed in the 'viewer'.</p>\n\n<p>Unlike the interactive notebook, the background one doesn't time out if you're not using it. It only times out if you exceed the maximum runtime allowed.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 815602,
          "author_name": "pawanverma",
          "author_url": "",
          "post_date": "04/21/2020 17:08:41",
          "content": "<p>Hi, i am trying to commit my notebook, but since last one hour same message showing GCE check failed is being displayed. I left my notebook as it is. Next day i saw commit operation was failed.\nPlease help me as i am unable to commit my version.\nIt failed with following log messages. At time 10166 error is coming but i have no idea about it. Please guide me <a href=\"/martingorner\">@martingorner</a></p>\n\n<p>Time\nLine #\nLog Message\n3.0s\n1\n[NbConvertApp] Converting notebook notebook.ipynb to notebook\n6.1s\n2\n[NbConvertApp] Executing notebook with kernel: python3\n14.3s\n3\n2020-04-20 17:18:06.289469: I tensorflow/core/platform/cpufeatureguard.cc:142] Your CPU supports instructions that this TensorFlow binary was not compiled to use: AVX512F\n14.3s\n4\n2020-04-20 17:18:06.314414: I tensorflow/core/platform/profileutils/cpuutils.cc:94] CPU Frequency: 2000179999 Hz\n14.3s\n5\n2020-04-20 17:18:06.315135: I tensorflow/compiler/xla/service/service.cc:168] XLA service 0x558c1cfbef00 initialized for platform Host (this does not guarantee that XLA will be used). Devices:\n14.3s\n6\n2020-04-20 17:18:06.315181: I tensorflow/compiler/xla/service/service.cc:176] StreamExecutor device (0): Host, Default Version\n14.3s\n7\n2020-04-20 17:18:06.344279: I tensorflow/core/distributedruntime/rpc/grpcchannel.cc:300] Initialize GrpcChannelCache for job worker -&gt; {0 -&gt; 10.0.0.2:8470}\n14.3s\n8\n2020-04-20 17:18:06.344339: I tensorflow/core/distributedruntime/rpc/grpcchannel.cc:300] Initialize GrpcChannelCache for job localhost -&gt; {0 -&gt; localhost:30013}\n14.3s\n9\n2020-04-20 17:18:06.359935: I tensorflow/core/distributedruntime/rpc/grpcchannel.cc:300] Initialize GrpcChannelCache for job worker -&gt; {0 -&gt; 10.0.0.2:8470}\n14.3s\n10\n2020-04-20 17:18:06.359995: I tensorflow/core/distributedruntime/rpc/grpcchannel.cc:300] Initialize GrpcChannelCache for job localhost -&gt; {0 -&gt; localhost:30013}\n14.3s\n11\n2020-04-20 17:18:06.366237: I tensorflow/core/distributedruntime/rpc/grpcserverlib.cc:390] Started server with target: grpc://localhost:30013 21.6s 12 2020-04-20 17:18:13.653909: W tensorflow/core/platform/cloud/googleauthprovider.cc:178] All attempts to get a Google authentication bearer token failed, returning an empty token. Retrieving token from files failed with \"Not found: Could not locate the credentials file.\". Retrieving token from GCE failed with \"Cancelled: GCE check skipped due to presence of $NOGCECHECK environment variable.\". 21.7s 13 2020-04-20 17:18:13.704236: W tensorflow/core/platform/cloud/googleauthprovider.cc:178] All attempts to get a Google authentication bearer token failed, returning an empty token. Retrieving token from files failed with \"Not found: Could not locate the credentials file.\". Retrieving token from GCE failed with \"Cancelled: GCE check skipped due to presence of $NOGCECHECK environment variable.\". 21.7s 14 2020-04-20 17:18:13.741587: W tensorflow/core/platform/cloud/googleauthprovider.cc:178] All attempts to get a Google authentication bearer token failed, returning an empty token. Retrieving token from files failed with \"Not found: Could not locate the credentials file.\". Retrieving token from GCE failed with \"Cancelled: GCE check skipped due to presence of $NOGCECHECK environment variable.\". 10166.2s 15 2020-04-20 20:07:18.218870: W tensorflow/core/distributedruntime/eager/remotetensorhandledata.cc:75] Unable to destroy remote tensor handles. If you are running a tf.function, it usually indicates some op in the graph gets an error: RecvTensor expects a different device incarnation: 14797989987728605988 vs. 1799738036951955635. Your worker job (\"/job:tpuworker/replica:0/task:0\") was probably restarted. Check your worker job for the reason why it was restarted.\n10166.2s\n16\nAdditional GRPC error information:\n10166.2s\n17\n{\"created\":\"@1587413238.218478572\",\"description\":\"Error received from peer\",\"file\":\"external/grpc/src/core/lib/surface/call.cc\",\"fileline\":1039,\"grpcmessage\":\"RecvTensor expects a different device incarnation: 14797989987728605988 vs. 1799738036951955635. Your worker job (\"/job:tpuworker/replica:0/task:0\") was probably restarted. Check your worker job for the reason why it was restarted.\",\"grpcstatus\":10}\n10169.6s\n18\n2020-04-20 20:07:21.664715: W ./tensorflow/core/distributedruntime/eager/destroytensorhandlenode.h:79] Ignoring an error encountered when deleting remote tensors handles: Invalid argument: Unable to find a contextid matching the specified one (12193600234898793608). Perhaps the worker was restarted, or the context was GC'd? 10169.6s 19 Additional GRPC error information: 10169.6s 20 {\"created\":\"@1587413241.664555067\",\"description\":\"Error received from peer\",\"file\":\"external/grpc/src/core/lib/surface/call.cc\",\"fileline\":1039,\"grpcmessage\":\"Unable to find a contextid matching the specified one (12193600234898793608). Perhaps the worker was restarted, or the context was GC'd?\",\"grpc_status\":3}\n10180.0s\n21\n[NbConvertApp] Writing 158697 bytes to notebook.ipynb\n10180.9s\n22\n[NbConvertApp] Converting notebook notebook.ipynb to html\n10182.4s\n23\n[NbConvertApp] Writing 508223 bytes to results.html\n10182.4s\n24\n10182.4s\n26\nComplete. Exited with code 0.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 815728,
          "author_name": "herbison",
          "author_url": "",
          "post_date": "04/21/2020 19:25:26",
          "content": "<p><a href=\"/pawanverma\">@pawanverma</a> this issue seems unrelated to the current thread, might get more responses creating your own post.</p>\n\n<p>I'm not sure I follow the issue here though, I'm wondering if maybe the code is consistently causing a problem on the TPU (maybe TPU OOM or something). I'm not an expert on TPU code though.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 889838,
          "author_name": "watzisname",
          "author_url": "",
          "post_date": "06/17/2020 07:19:31",
          "content": "<p><a href=\"/herbison\">@herbison</a>  &amp; Kaggle Team,</p>\n\n<p>This means if switch on GPU/TPU and then commit -&gt; It will create a background session for committing. Hence, I can switch off the GPU/TPU for interactive mode.  If we don't, then we are using 2X of the resources - this is painfully not clear in the UI.</p>\n\n<p>This should be taken on priority by Kaggle to clarify on the UI - since a lot of people are just waitinf for it to commit and wating double the resources.</p>\n\n<p>Also, a message should be sent after the commit is complete, so that we can submit the solution. </p>\n\n<p>Thanks.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 891874,
          "author_name": "herbison",
          "author_url": "",
          "post_date": "06/18/2020 13:57:46",
          "content": "<p>I sent along your feedback, I agree it would be nice to show it clearer in the UI that it counts as an extra session, we're working on that. Any suggestions on what would help make that clearer to you?</p>\n\n<p>I agree about notification on finish too, passed it along.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 892166,
          "author_name": "watzisname",
          "author_url": "",
          "post_date": "06/18/2020 17:37:05",
          "content": "<p>Hi <a href=\"/herbison\">@herbison</a>,</p>\n\n<p>A couple of ideas...</p>\n\n<p>A) As soon as I run -&gt; Save and Commit All -&gt; My current interactive session should \"force\" stop GPU session and go to None. If user tries to enable GPU a warning can pop up, that a background session of the commit is still running( IF it is running). </p>\n\n<p>Another idea is to \"always\" show GPU and TPU sessions running in totality on top of a notebook. Currently i have to go to notebooks to find this.</p>\n\n<p>B) Once the background finish commit happens, a notification can arrive in the same place notifications comeup for chat messages like this one.</p>\n\n<p>Hope that helps</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 811822,
      "author_name": "pawanverma",
      "author_url": "",
      "post_date": "04/18/2020 08:58:11",
      "content": "<p>Can i stop my notebook after committing so as to save TPU quota?</p>",
      "votes": null,
      "replies": [
        {
          "id": 815707,
          "author_name": "kurianbenoy",
          "author_url": "",
          "post_date": "04/21/2020 18:55:44",
          "content": "<p>Yes definitely, you can. Just change settings in acceleator section, so no more interactive section is not using it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 816426,
          "author_name": "pawanverma",
          "author_url": "",
          "post_date": "04/22/2020 11:00:07",
          "content": "<p>Actually my question was that during committing since there is a separate session created and our notebook is compiled there so can we stop our interactive notebook so as to save TPU quota.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 816653,
          "author_name": "herbison",
          "author_url": "",
          "post_date": "04/22/2020 13:48:05",
          "content": "<p>Yes you can stop the interactive notebook without stopping the commit to save quota. The background  commit will continue to run fine.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "811488": "Every time I click \"Save and Commit\" on my notebook, my notebook shows an error message \"Are you still there?  Your TPU notebook stops after 15 minutes of idle time. Commit your notebook for longer computations\". It also shows session as \"Disconnected\".\n\nI know that commit runs the whole notebook and my training within the notebook takes ~ 30-35 minutes. But after waiting for about 40 minutes, my notebook gets committed.",
    "811574": "When you save and commit your notebook, it runs in a separate batch session in the background, and that's where the current version of it is committed. The message you see is for the current interactive session that you began using and editing. The kernel that was running as you made your edits stopped after 15 minutes of idle time while the batch process continued in the background. I hope this helps.",
    "811822": "Can i stop my notebook after committing so as to save TPU quota?",
    "815597": "^ This is correct. \n\nWe're still working on explaining it effectively in the UI, but your interactive notebook is executing totally separate to 'commits/save versions'. The interactive one lets you write your notebook, executing any part of it as you go to test it and iterate on it. Once you've got a notebook that you'd like to execute  from a clean slate top-to-bottom, you can hit 'save version &gt; save &amp; run all' and it copies your interactive notebook and executes it on a background machine top-to-bottom and then saves and formats the results to be displayed in the 'viewer'.\n\nUnlike the interactive notebook, the background one doesn't time out if you're not using it. It only times out if you exceed the maximum runtime allowed.",
    "815602": "Hi, i am trying to commit my notebook, but since last one hour same message showing GCE check failed is being displayed. I left my notebook as it is. Next day i saw commit operation was failed.\nPlease help me as i am unable to commit my version.\nIt failed with following log messages. At time 10166 error is coming but i have no idea about it. Please guide me @martingorner\n\nTime\nLine #\nLog Message\n3.0s\n1\n[NbConvertApp] Converting notebook notebook.ipynb to notebook\n6.1s\n2\n[NbConvertApp] Executing notebook with kernel: python3\n14.3s\n3\n2020-04-20 17:18:06.289469: I tensorflow/core/platform/cpufeatureguard.cc:142] Your CPU supports instructions that this TensorFlow binary was not compiled to use: AVX512F\n14.3s\n4\n2020-04-20 17:18:06.314414: I tensorflow/core/platform/profileutils/cpuutils.cc:94] CPU Frequency: 2000179999 Hz\n14.3s\n5\n2020-04-20 17:18:06.315135: I tensorflow/compiler/xla/service/service.cc:168] XLA service 0x558c1cfbef00 initialized for platform Host (this does not guarantee that XLA will be used). Devices:\n14.3s\n6\n2020-04-20 17:18:06.315181: I tensorflow/compiler/xla/service/service.cc:176] StreamExecutor device (0): Host, Default Version\n14.3s\n7\n2020-04-20 17:18:06.344279: I tensorflow/core/distributedruntime/rpc/grpcchannel.cc:300] Initialize GrpcChannelCache for job worker -&gt; {0 -&gt; 10.0.0.2:8470}\n14.3s\n8\n2020-04-20 17:18:06.344339: I tensorflow/core/distributedruntime/rpc/grpcchannel.cc:300] Initialize GrpcChannelCache for job localhost -&gt; {0 -&gt; localhost:30013}\n14.3s\n9\n2020-04-20 17:18:06.359935: I tensorflow/core/distributedruntime/rpc/grpcchannel.cc:300] Initialize GrpcChannelCache for job worker -&gt; {0 -&gt; 10.0.0.2:8470}\n14.3s\n10\n2020-04-20 17:18:06.359995: I tensorflow/core/distributedruntime/rpc/grpcchannel.cc:300] Initialize GrpcChannelCache for job localhost -&gt; {0 -&gt; localhost:30013}\n14.3s\n11\n2020-04-20 17:18:06.366237: I tensorflow/core/distributedruntime/rpc/grpcserverlib.cc:390] Started server with target: grpc://localhost:30013 21.6s 12 2020-04-20 17:18:13.653909: W tensorflow/core/platform/cloud/googleauthprovider.cc:178] All attempts to get a Google authentication bearer token failed, returning an empty token. Retrieving token from files failed with \"Not found: Could not locate the credentials file.\". Retrieving token from GCE failed with \"Cancelled: GCE check skipped due to presence of $NOGCECHECK environment variable.\". 21.7s 13 2020-04-20 17:18:13.704236: W tensorflow/core/platform/cloud/googleauthprovider.cc:178] All attempts to get a Google authentication bearer token failed, returning an empty token. Retrieving token from files failed with \"Not found: Could not locate the credentials file.\". Retrieving token from GCE failed with \"Cancelled: GCE check skipped due to presence of $NOGCECHECK environment variable.\". 21.7s 14 2020-04-20 17:18:13.741587: W tensorflow/core/platform/cloud/googleauthprovider.cc:178] All attempts to get a Google authentication bearer token failed, returning an empty token. Retrieving token from files failed with \"Not found: Could not locate the credentials file.\". Retrieving token from GCE failed with \"Cancelled: GCE check skipped due to presence of $NOGCECHECK environment variable.\". 10166.2s 15 2020-04-20 20:07:18.218870: W tensorflow/core/distributedruntime/eager/remotetensorhandledata.cc:75] Unable to destroy remote tensor handles. If you are running a tf.function, it usually indicates some op in the graph gets an error: RecvTensor expects a different device incarnation: 14797989987728605988 vs. 1799738036951955635. Your worker job (\"/job:tpuworker/replica:0/task:0\") was probably restarted. Check your worker job for the reason why it was restarted.\n10166.2s\n16\nAdditional GRPC error information:\n10166.2s\n17\n{\"created\":\"@1587413238.218478572\",\"description\":\"Error received from peer\",\"file\":\"external/grpc/src/core/lib/surface/call.cc\",\"fileline\":1039,\"grpcmessage\":\"RecvTensor expects a different device incarnation: 14797989987728605988 vs. 1799738036951955635. Your worker job (\"/job:tpuworker/replica:0/task:0\") was probably restarted. Check your worker job for the reason why it was restarted.\",\"grpcstatus\":10}\n10169.6s\n18\n2020-04-20 20:07:21.664715: W ./tensorflow/core/distributedruntime/eager/destroytensorhandlenode.h:79] Ignoring an error encountered when deleting remote tensors handles: Invalid argument: Unable to find a contextid matching the specified one (12193600234898793608). Perhaps the worker was restarted, or the context was GC'd? 10169.6s 19 Additional GRPC error information: 10169.6s 20 {\"created\":\"@1587413241.664555067\",\"description\":\"Error received from peer\",\"file\":\"external/grpc/src/core/lib/surface/call.cc\",\"fileline\":1039,\"grpcmessage\":\"Unable to find a contextid matching the specified one (12193600234898793608). Perhaps the worker was restarted, or the context was GC'd?\",\"grpc_status\":3}\n10180.0s\n21\n[NbConvertApp] Writing 158697 bytes to notebook.ipynb\n10180.9s\n22\n[NbConvertApp] Converting notebook notebook.ipynb to html\n10182.4s\n23\n[NbConvertApp] Writing 508223 bytes to results.html\n10182.4s\n24\n10182.4s\n26\nComplete. Exited with code 0.",
    "815707": "Yes definitely, you can. Just change settings in acceleator section, so no more interactive section is not using it.",
    "815728": "pawanverma this issue seems unrelated to the current thread, might get more responses creating your own post.\n\nI'm not sure I follow the issue here though, I'm wondering if maybe the code is consistently causing a problem on the TPU (maybe TPU OOM or something). I'm not an expert on TPU code though.",
    "816426": "Actually my question was that during committing since there is a separate session created and our notebook is compiled there so can we stop our interactive notebook so as to save TPU quota.",
    "816653": "Yes you can stop the interactive notebook without stopping the commit to save quota. The background  commit will continue to run fine.",
    "889838": "herbison  &amp; Kaggle Team,\n\nThis means if switch on GPU/TPU and then commit -&gt; It will create a background session for committing. Hence, I can switch off the GPU/TPU for interactive mode.  If we don't, then we are using 2X of the resources - this is painfully not clear in the UI.\n\nThis should be taken on priority by Kaggle to clarify on the UI - since a lot of people are just waitinf for it to commit and wating double the resources.\n\nAlso, a message should be sent after the commit is complete, so that we can submit the solution. \n\nThanks.",
    "891874": "I sent along your feedback, I agree it would be nice to show it clearer in the UI that it counts as an extra session, we're working on that. Any suggestions on what would help make that clearer to you?\n\nI agree about notification on finish too, passed it along.",
    "892166": "Hi @herbison,\n\nA couple of ideas...\n\nA) As soon as I run -&gt; Save and Commit All -&gt; My current interactive session should \"force\" stop GPU session and go to None. If user tries to enable GPU a warning can pop up, that a background session of the commit is still running( IF it is running). \n\nAnother idea is to \"always\" show GPU and TPU sessions running in totality on top of a notebook. Currently i have to go to notebooks to find this.\n\n\nB) Once the background finish commit happens, a notification can arrive in the same place notifications comeup for chat messages like this one.\n\nHope that helps"
  },
  "source": "meta"
}