{
  "id": 143918,
  "title": "Committing taking more than 40 mins",
  "url": "/competitions/flower-classification-with-tpus/discussion/143918",
  "author_name": "",
  "post_date": "2020-04-16T19:53:20.179556200Z",
  "votes": 1,
  "comment_count": 21,
  "views": 0,
  "content": "<p>Hi, I am new to kaggle and this will be my first submission. I am committing my notebook but from last 40 mins it is saving.\nMy notebook usually runs in 20-30 mins. Is it usual to take so much time for committing?\nOr is there any other way also</p>",
  "messages": [
    {
      "id": "810238",
      "postDate": "04/16/2020 19:53:20",
      "content": "<p>Hi, I am new to kaggle and this will be my first submission. I am committing my notebook but from last 40 mins it is saving.\nMy notebook usually runs in 20-30 mins. Is it usual to take so much time for committing?\nOr is there any other way also</p>",
      "rawMarkdown": "Hi, I am new to kaggle and this will be my first submission. I am committing my notebook but from last 40 mins it is saving.\nMy notebook usually runs in 20-30 mins. Is it usual to take so much time for committing?\nOr is there any other way also",
      "votes": null
    },
    {
      "id": "816436",
      "postDate": "04/22/2020 11:04:23",
      "content": "<p>Hi, i am trying to commit my notebook, but since last one hour same message showing GCE check failed is being displayed. I left my notebook as it is. Next day i saw commit operation was failed.\nPlease help me as i am unable to commit my version.\nIt failed with following log messages. At time 10166 error is coming but i have no idea about it. Please guide me <a href=\"/martingorner\">@martingorner</a> </p>\n\n<p>Time\nLine #\nLog Message\n3.0s\n1\n[NbConvertApp] Converting notebook notebook.ipynb to notebook\n6.1s\n2\n[NbConvertApp] Executing notebook with kernel: python3\n14.3s\n3\n2020-04-20 17:18:06.289469: I tensorflow/core/platform/cpufeatureguard.cc:142] Your CPU supports instructions that this TensorFlow binary was not compiled to use: AVX512F\n14.3s\n4\n2020-04-20 17:18:06.314414: I tensorflow/core/platform/profileutils/cpuutils.cc:94] CPU Frequency: 2000179999 Hz\n14.3s\n5\n2020-04-20 17:18:06.315135: I tensorflow/compiler/xla/service/service.cc:168] XLA service 0x558c1cfbef00 initialized for platform Host (this does not guarantee that XLA will be used). Devices:\n14.3s\n6\n2020-04-20 17:18:06.315181: I tensorflow/compiler/xla/service/service.cc:176] StreamExecutor device (0): Host, Default Version\n14.3s\n7\n2020-04-20 17:18:06.344279: I tensorflow/core/distributedruntime/rpc/grpcchannel.cc:300] Initialize GrpcChannelCache for job worker -&gt; {0 -&gt; 10.0.0.2:8470}\n14.3s\n8\n2020-04-20 17:18:06.344339: I tensorflow/core/distributedruntime/rpc/grpcchannel.cc:300] Initialize GrpcChannelCache for job localhost -&gt; {0 -&gt; localhost:30013}\n14.3s\n9\n2020-04-20 17:18:06.359935: I tensorflow/core/distributedruntime/rpc/grpcchannel.cc:300] Initialize GrpcChannelCache for job worker -&gt; {0 -&gt; 10.0.0.2:8470}\n14.3s\n10\n2020-04-20 17:18:06.359995: I tensorflow/core/distributedruntime/rpc/grpcchannel.cc:300] Initialize GrpcChannelCache for job localhost -&gt; {0 -&gt; localhost:30013}\n14.3s\n11\n2020-04-20 17:18:06.366237: I tensorflow/core/distributedruntime/rpc/grpcserverlib.cc:390] Started server with target: grpc://localhost:30013 21.6s 12 2020-04-20 17:18:13.653909: W tensorflow/core/platform/cloud/googleauthprovider.cc:178] All attempts to get a Google authentication bearer token failed, returning an empty token. Retrieving token from files failed with \"Not found: Could not locate the credentials file.\". Retrieving token from GCE failed with \"Cancelled: GCE check skipped due to presence of $NOGCECHECK environment variable.\". 21.7s 13 2020-04-20 17:18:13.704236: W tensorflow/core/platform/cloud/googleauthprovider.cc:178] All attempts to get a Google authentication bearer token failed, returning an empty token. Retrieving token from files failed with \"Not found: Could not locate the credentials file.\". Retrieving token from GCE failed with \"Cancelled: GCE check skipped due to presence of $NOGCECHECK environment variable.\". 21.7s 14 2020-04-20 17:18:13.741587: W tensorflow/core/platform/cloud/googleauthprovider.cc:178] All attempts to get a Google authentication bearer token failed, returning an empty token. Retrieving token from files failed with \"Not found: Could not locate the credentials file.\". Retrieving token from GCE failed with \"Cancelled: GCE check skipped due to presence of $NOGCECHECK environment variable.\". 10166.2s 15 2020-04-20 20:07:18.218870: W tensorflow/core/distributedruntime/eager/remotetensorhandledata.cc:75] Unable to destroy remote tensor handles. If you are running a tf.function, it usually indicates some op in the graph gets an error: RecvTensor expects a different device incarnation: 14797989987728605988 vs. 1799738036951955635. Your worker job (\"/job:tpuworker/replica:0/task:0\") was probably restarted. Check your worker job for the reason why it was restarted.\n10166.2s\n16\nAdditional GRPC error information:\n10166.2s\n17\n{\"created\":\"@1587413238.218478572\",\"description\":\"Error received from peer\",\"file\":\"external/grpc/src/core/lib/surface/call.cc\",\"fileline\":1039,\"grpcmessage\":\"RecvTensor expects a different device incarnation: 14797989987728605988 vs. 1799738036951955635. Your worker job (\"/job:tpuworker/replica:0/task:0\") was probably restarted. Check your worker job for the reason why it was restarted.\",\"grpcstatus\":10}\n10169.6s\n18\n2020-04-20 20:07:21.664715: W ./tensorflow/core/distributedruntime/eager/destroytensorhandlenode.h:79] Ignoring an error encountered when deleting remote tensors handles: Invalid argument: Unable to find a contextid matching the specified one (12193600234898793608). Perhaps the worker was restarted, or the context was GC'd? 10169.6s 19 Additional GRPC error information: 10169.6s 20 {\"created\":\"@1587413241.664555067\",\"description\":\"Error received from peer\",\"file\":\"external/grpc/src/core/lib/surface/call.cc\",\"fileline\":1039,\"grpcmessage\":\"Unable to find a contextid matching the specified one (12193600234898793608). Perhaps the worker was restarted, or the context was GC'd?\",\"grpc_status\":3}\n10180.0s\n21\n[NbConvertApp] Writing 158697 bytes to notebook.ipynb\n10180.9s\n22\n[NbConvertApp] Converting notebook notebook.ipynb to html\n10182.4s\n23\n[NbConvertApp] Writing 508223 bytes to results.html\n10182.4s\n24\n10182.4s\n26\nComplete. Exited with code 0.</p>",
      "rawMarkdown": "Hi, i am trying to commit my notebook, but since last one hour same message showing GCE check failed is being displayed. I left my notebook as it is. Next day i saw commit operation was failed.\nPlease help me as i am unable to commit my version.\nIt failed with following log messages. At time 10166 error is coming but i have no idea about it. Please guide me @martingorner \n\nTime\nLine #\nLog Message\n3.0s\n1\n[NbConvertApp] Converting notebook notebook.ipynb to notebook\n6.1s\n2\n[NbConvertApp] Executing notebook with kernel: python3\n14.3s\n3\n2020-04-20 17:18:06.289469: I tensorflow/core/platform/cpufeatureguard.cc:142] Your CPU supports instructions that this TensorFlow binary was not compiled to use: AVX512F\n14.3s\n4\n2020-04-20 17:18:06.314414: I tensorflow/core/platform/profileutils/cpuutils.cc:94] CPU Frequency: 2000179999 Hz\n14.3s\n5\n2020-04-20 17:18:06.315135: I tensorflow/compiler/xla/service/service.cc:168] XLA service 0x558c1cfbef00 initialized for platform Host (this does not guarantee that XLA will be used). Devices:\n14.3s\n6\n2020-04-20 17:18:06.315181: I tensorflow/compiler/xla/service/service.cc:176] StreamExecutor device (0): Host, Default Version\n14.3s\n7\n2020-04-20 17:18:06.344279: I tensorflow/core/distributedruntime/rpc/grpcchannel.cc:300] Initialize GrpcChannelCache for job worker -&gt; {0 -&gt; 10.0.0.2:8470}\n14.3s\n8\n2020-04-20 17:18:06.344339: I tensorflow/core/distributedruntime/rpc/grpcchannel.cc:300] Initialize GrpcChannelCache for job localhost -&gt; {0 -&gt; localhost:30013}\n14.3s\n9\n2020-04-20 17:18:06.359935: I tensorflow/core/distributedruntime/rpc/grpcchannel.cc:300] Initialize GrpcChannelCache for job worker -&gt; {0 -&gt; 10.0.0.2:8470}\n14.3s\n10\n2020-04-20 17:18:06.359995: I tensorflow/core/distributedruntime/rpc/grpcchannel.cc:300] Initialize GrpcChannelCache for job localhost -&gt; {0 -&gt; localhost:30013}\n14.3s\n11\n2020-04-20 17:18:06.366237: I tensorflow/core/distributedruntime/rpc/grpcserverlib.cc:390] Started server with target: grpc://localhost:30013 21.6s 12 2020-04-20 17:18:13.653909: W tensorflow/core/platform/cloud/googleauthprovider.cc:178] All attempts to get a Google authentication bearer token failed, returning an empty token. Retrieving token from files failed with \"Not found: Could not locate the credentials file.\". Retrieving token from GCE failed with \"Cancelled: GCE check skipped due to presence of $NOGCECHECK environment variable.\". 21.7s 13 2020-04-20 17:18:13.704236: W tensorflow/core/platform/cloud/googleauthprovider.cc:178] All attempts to get a Google authentication bearer token failed, returning an empty token. Retrieving token from files failed with \"Not found: Could not locate the credentials file.\". Retrieving token from GCE failed with \"Cancelled: GCE check skipped due to presence of $NOGCECHECK environment variable.\". 21.7s 14 2020-04-20 17:18:13.741587: W tensorflow/core/platform/cloud/googleauthprovider.cc:178] All attempts to get a Google authentication bearer token failed, returning an empty token. Retrieving token from files failed with \"Not found: Could not locate the credentials file.\". Retrieving token from GCE failed with \"Cancelled: GCE check skipped due to presence of $NOGCECHECK environment variable.\". 10166.2s 15 2020-04-20 20:07:18.218870: W tensorflow/core/distributedruntime/eager/remotetensorhandledata.cc:75] Unable to destroy remote tensor handles. If you are running a tf.function, it usually indicates some op in the graph gets an error: RecvTensor expects a different device incarnation: 14797989987728605988 vs. 1799738036951955635. Your worker job (\"/job:tpuworker/replica:0/task:0\") was probably restarted. Check your worker job for the reason why it was restarted.\n10166.2s\n16\nAdditional GRPC error information:\n10166.2s\n17\n{\"created\":\"@1587413238.218478572\",\"description\":\"Error received from peer\",\"file\":\"external/grpc/src/core/lib/surface/call.cc\",\"fileline\":1039,\"grpcmessage\":\"RecvTensor expects a different device incarnation: 14797989987728605988 vs. 1799738036951955635. Your worker job (\"/job:tpuworker/replica:0/task:0\") was probably restarted. Check your worker job for the reason why it was restarted.\",\"grpcstatus\":10}\n10169.6s\n18\n2020-04-20 20:07:21.664715: W ./tensorflow/core/distributedruntime/eager/destroytensorhandlenode.h:79] Ignoring an error encountered when deleting remote tensors handles: Invalid argument: Unable to find a contextid matching the specified one (12193600234898793608). Perhaps the worker was restarted, or the context was GC'd? 10169.6s 19 Additional GRPC error information: 10169.6s 20 {\"created\":\"@1587413241.664555067\",\"description\":\"Error received from peer\",\"file\":\"external/grpc/src/core/lib/surface/call.cc\",\"fileline\":1039,\"grpcmessage\":\"Unable to find a contextid matching the specified one (12193600234898793608). Perhaps the worker was restarted, or the context was GC'd?\",\"grpc_status\":3}\n10180.0s\n21\n[NbConvertApp] Writing 158697 bytes to notebook.ipynb\n10180.9s\n22\n[NbConvertApp] Converting notebook notebook.ipynb to html\n10182.4s\n23\n[NbConvertApp] Writing 508223 bytes to results.html\n10182.4s\n24\n10182.4s\n26\nComplete. Exited with code 0.",
      "votes": null
    },
    {
      "id": "816671",
      "postDate": "04/22/2020 13:59:52",
      "content": "<p>There can potentially be many reasons for a commit to take longer. For example, your commit may not even start executing immediately when you press the button. We have to schedule the execution on a machine, and if there is enough demand all at once, it may be queued for awhile before it starts. We do our best to avoid queuing at all times, but sometimes TPUs are so popular all at once that we have to add a little wait time.</p>\n\n<p>But that's assuming it's a scheduling problem, it could also be that your code runs into a rare condition that just takes longer (anything involving IO/network will have some level of variance in it due to to variable throughput rates).</p>",
      "rawMarkdown": "There can potentially be many reasons for a commit to take longer. For example, your commit may not even start executing immediately when you press the button. We have to schedule the execution on a machine, and if there is enough demand all at once, it may be queued for awhile before it starts. We do our best to avoid queuing at all times, but sometimes TPUs are so popular all at once that we have to add a little wait time.\n\nBut that's assuming it's a scheduling problem, it could also be that your code runs into a rare condition that just takes longer (anything involving IO/network will have some level of variance in it due to to variable throughput rates).",
      "votes": null
    },
    {
      "id": "816733",
      "postDate": "04/22/2020 14:41:53",
      "content": "<p>I tried for committing for third time in three day. But it again failed with error due to time out errot. When i run my notebook in interactive mode, all cells successfully execute so it seems there is some problem with scheduling TPU for commit operation.\nI request you to please check this issue as i am unable to submit my notebook and submission deadline is approaching.\nAttached are the latest logs for commit operation:\nTime\nLine #\nLog Message\n3.6s\n1\n[NbConvertApp] Converting notebook <strong>notebook</strong>.ipynb to notebook\n6.9s\n2\n[NbConvertApp] Executing notebook with kernel: python3\n16.0s\n3\n2020-04-22 11:08:42.138682: I tensorflow/core/platform/cpu_feature_guard.cc:142] Your CPU supports instructions that this TensorFlow binary was not compiled to use: AVX512F\n16.0s\n4\n2020-04-22 11:08:42.171854: I tensorflow/core/platform/profile_utils/cpu_utils.cc:94] CPU Frequency: 2000170000 Hz\n16.0s\n5\n2020-04-22 11:08:42.173379: I tensorflow/compiler/xla/service/service.cc:168] XLA service 0x55a1c3d9c950 initialized for platform Host (this does not guarantee that XLA will be used). Devices:\n16.0s\n6\n2020-04-22 11:08:42.173435: I tensorflow/compiler/xla/service/service.cc:176]   StreamExecutor device (0): Host, Default Version\n16.1s\n7\n2020-04-22 11:08:42.198828: I tensorflow/core/distributed_runtime/rpc/grpc_channel.cc:300] Initialize GrpcChannelCache for job worker -&gt; {0 -&gt; 10.0.0.2:8470}\n16.1s\n8\n2020-04-22 11:08:42.198889: I tensorflow/core/distributed_runtime/rpc/grpc_channel.cc:300] Initialize GrpcChannelCache for job localhost -&gt; {0 -&gt; localhost:30012}\n16.1s\n9\n2020-04-22 11:08:42.214317: I tensorflow/core/distributed_runtime/rpc/grpc_channel.cc:300] Initialize GrpcChannelCache for job worker -&gt; {0 -&gt; 10.0.0.2:8470}\n16.1s\n10\n2020-04-22 11:08:42.214390: I tensorflow/core/distributed_runtime/rpc/grpc_channel.cc:300] Initialize GrpcChannelCache for job localhost -&gt; {0 -&gt; localhost:30012}\n16.1s\n11\n2020-04-22 11:08:42.219715: I tensorflow/core/distributed_runtime/rpc/grpc_server_lib.cc:390] Started server with target: grpc://localhost:30012\n23.2s\n12\n2020-04-22 11:08:49.292527: W tensorflow/core/platform/cloud/google_auth_provider.cc:178] All attempts to get a Google authentication bearer token failed, returning an empty token. Retrieving token from files failed with \"Not found: Could not locate the credentials file.\". Retrieving token from GCE failed with \"Cancelled: GCE check skipped due to presence of $NO_GCE_CHECK environment variable.\".\n23.2s\n13\n2020-04-22 11:08:49.342698: W tensorflow/core/platform/cloud/google_auth_provider.cc:178] All attempts to get a Google authentication bearer token failed, returning an empty token. Retrieving token from files failed with \"Not found: Could not locate the credentials file.\". Retrieving token from GCE failed with \"Cancelled: GCE check skipped due to presence of $NO_GCE_CHECK environment variable.\".\n23.2s\n14\n2020-04-22 11:08:49.376955: W tensorflow/core/platform/cloud/google_auth_provider.cc:178] All attempts to get a Google authentication bearer token failed, returning an empty token. Retrieving token from files failed with \"Not found: Could not locate the credentials file.\". Retrieving token from GCE failed with \"Cancelled: GCE check skipped due to presence of $NO_GCE_CHECK environment variable.\".\n23.2s\n15\n23.2s\n17\nFailed. Exited with code 137.</p>",
      "rawMarkdown": "I tried for committing for third time in three day. But it again failed with error due to time out errot. When i run my notebook in interactive mode, all cells successfully execute so it seems there is some problem with scheduling TPU for commit operation.\nI request you to please check this issue as i am unable to submit my notebook and submission deadline is approaching.\nAttached are the latest logs for commit operation:\nTime\nLine #\nLog Message\n3.6s\n1\n[NbConvertApp] Converting notebook __notebook__.ipynb to notebook\n6.9s\n2\n[NbConvertApp] Executing notebook with kernel: python3\n16.0s\n3\n2020-04-22 11:08:42.138682: I tensorflow/core/platform/cpu_feature_guard.cc:142] Your CPU supports instructions that this TensorFlow binary was not compiled to use: AVX512F\n16.0s\n4\n2020-04-22 11:08:42.171854: I tensorflow/core/platform/profile_utils/cpu_utils.cc:94] CPU Frequency: 2000170000 Hz\n16.0s\n5\n2020-04-22 11:08:42.173379: I tensorflow/compiler/xla/service/service.cc:168] XLA service 0x55a1c3d9c950 initialized for platform Host (this does not guarantee that XLA will be used). Devices:\n16.0s\n6\n2020-04-22 11:08:42.173435: I tensorflow/compiler/xla/service/service.cc:176]   StreamExecutor device (0): Host, Default Version\n16.1s\n7\n2020-04-22 11:08:42.198828: I tensorflow/core/distributed_runtime/rpc/grpc_channel.cc:300] Initialize GrpcChannelCache for job worker -&gt; {0 -&gt; 10.0.0.2:8470}\n16.1s\n8\n2020-04-22 11:08:42.198889: I tensorflow/core/distributed_runtime/rpc/grpc_channel.cc:300] Initialize GrpcChannelCache for job localhost -&gt; {0 -&gt; localhost:30012}\n16.1s\n9\n2020-04-22 11:08:42.214317: I tensorflow/core/distributed_runtime/rpc/grpc_channel.cc:300] Initialize GrpcChannelCache for job worker -&gt; {0 -&gt; 10.0.0.2:8470}\n16.1s\n10\n2020-04-22 11:08:42.214390: I tensorflow/core/distributed_runtime/rpc/grpc_channel.cc:300] Initialize GrpcChannelCache for job localhost -&gt; {0 -&gt; localhost:30012}\n16.1s\n11\n2020-04-22 11:08:42.219715: I tensorflow/core/distributed_runtime/rpc/grpc_server_lib.cc:390] Started server with target: grpc://localhost:30012\n23.2s\n12\n2020-04-22 11:08:49.292527: W tensorflow/core/platform/cloud/google_auth_provider.cc:178] All attempts to get a Google authentication bearer token failed, returning an empty token. Retrieving token from files failed with \"Not found: Could not locate the credentials file.\". Retrieving token from GCE failed with \"Cancelled: GCE check skipped due to presence of $NO_GCE_CHECK environment variable.\".\n23.2s\n13\n2020-04-22 11:08:49.342698: W tensorflow/core/platform/cloud/google_auth_provider.cc:178] All attempts to get a Google authentication bearer token failed, returning an empty token. Retrieving token from files failed with \"Not found: Could not locate the credentials file.\". Retrieving token from GCE failed with \"Cancelled: GCE check skipped due to presence of $NO_GCE_CHECK environment variable.\".\n23.2s\n14\n2020-04-22 11:08:49.376955: W tensorflow/core/platform/cloud/google_auth_provider.cc:178] All attempts to get a Google authentication bearer token failed, returning an empty token. Retrieving token from files failed with \"Not found: Could not locate the credentials file.\". Retrieving token from GCE failed with \"Cancelled: GCE check skipped due to presence of $NO_GCE_CHECK environment variable.\".\n23.2s\n15\n23.2s\n17\nFailed. Exited with code 137.",
      "votes": null
    },
    {
      "id": "816737",
      "postDate": "04/22/2020 14:44:20",
      "content": "<p>Scheduling does not affect the timeout, only when it starts, the timer starts ticking when execution starts, not when you submit it.</p>\n\n<p>I'll pass this along to some people who know TPUs better than me to see if they know what's going on.</p>",
      "rawMarkdown": "Scheduling does not affect the timeout, only when it starts, the timer starts ticking when execution starts, not when you submit it.\n\nI'll pass this along to some people who know TPUs better than me to see if they know what's going on.",
      "votes": null
    },
    {
      "id": "816789",
      "postDate": "04/22/2020 15:39:24",
      "content": "<p>The GCE check warning messages should be harmless and  can be ignored. The next version of Tensorflow will no longer print these. The message basically says that TF will not use any credentials to access public GCS object.\nPerhaps the real error is <code>Unable to destroy remote tensor handles. If you are running a tf.function, it usually indicates some op in the graph gets an error: RecvTensor expects a different device incarnation: 14797989987728605988 vs. 1799738036951955635. Your worker job (\"/job:tpuworker/replica:0/task:0\") was probably restarted. Check your worker job for the reason why it was restarted.</code>?  </p>",
      "rawMarkdown": "The GCE check warning messages should be harmless and  can be ignored. The next version of Tensorflow will no longer print these. The message basically says that TF will not use any credentials to access public GCS object.\nPerhaps the real error is `Unable to destroy remote tensor handles. If you are running a tf.function, it usually indicates some op in the graph gets an error: RecvTensor expects a different device incarnation: 14797989987728605988 vs. 1799738036951955635. Your worker job (\"/job:tpuworker/replica:0/task:0\") was probably restarted. Check your worker job for the reason why it was restarted.`?",
      "votes": null
    },
    {
      "id": "816967",
      "postDate": "04/22/2020 18:33:09",
      "content": "<p>Thanks. That will be very helpful because this problem is recurring every time after commit. But if i run all cells without commit then there seems to be no problem.</p>",
      "rawMarkdown": "Thanks. That will be very helpful because this problem is recurring every time after commit. But if i run all cells without commit then there seems to be no problem.",
      "votes": null
    },
    {
      "id": "816971",
      "postDate": "04/22/2020 18:36:39",
      "content": "<p>There is no problem when i run all cells in interactive mode but only during commit i face this problem. Moreover for continuous 2 hrs i monitored commit operation and it kept showing the harmless GCE check warning message. And i have no idea how worker job restarts. I simply commit and wait.\nI have wasted around 10 hrs of my TPU quota for trying to commit this notebook.\nPlease check if you can find any issue.</p>",
      "rawMarkdown": "There is no problem when i run all cells in interactive mode but only during commit i face this problem. Moreover for continuous 2 hrs i monitored commit operation and it kept showing the harmless GCE check warning message. And i have no idea how worker job restarts. I simply commit and wait.\nI have wasted around 10 hrs of my TPU quota for trying to commit this notebook.\nPlease check if you can find any issue.",
      "votes": null
    },
    {
      "id": "816994",
      "postDate": "04/22/2020 18:56:28",
      "content": "<p>Yep, the Tensorflow library will probably keep outputting \"GCE check\" warnings every time the TPU accesses the data from Google Cloud Storage. Does the notebook output you shared earlier contain any more messages? Please, share if that's the case.</p>\n\n<p>Also, when you look at the commit output URL, you should see something like scriptVersionId=[numberHere]. Could you please share the id with us so that we can try to find out more info?</p>",
      "rawMarkdown": "Yep, the Tensorflow library will probably keep outputting \"GCE check\" warnings every time the TPU accesses the data from Google Cloud Storage. Does the notebook output you shared earlier contain any more messages? Please, share if that's the case.\n\nAlso, when you look at the commit output URL, you should see something like scriptVersionId=[numberHere]. Could you please share the id with us so that we can try to find out more info?",
      "votes": null
    },
    {
      "id": "817020",
      "postDate": "04/22/2020 19:22:05",
      "content": "<p><a href=\"https://www.kaggle.com/pawanverma/flowers-efficient-net-dense-xception/log?scriptVersionId=32431497\">https://www.kaggle.com/pawanverma/flowers-efficient-net-dense-xception/log?scriptVersionId=32431497</a>.\nI have pasted entire log file of my last commit attempts. Thanks in advance for all the help.</p>",
      "rawMarkdown": "https://www.kaggle.com/pawanverma/flowers-efficient-net-dense-xception/log?scriptVersionId=32431497.\nI have pasted entire log file of my last commit attempts. Thanks in advance for all the help.",
      "votes": null
    },
    {
      "id": "817122",
      "postDate": "04/22/2020 21:22:22",
      "content": "<p>Btw, <a href=\"/pawanverma\">@pawanverma</a>, when you ran your session interactively top to bottom, how long did it take?\nThe commit sessions will time out after 3 hours if I recall correctly.</p>",
      "rawMarkdown": "Btw, @pawanverma, when you ran your session interactively top to bottom, how long did it take?\nThe commit sessions will time out after 3 hours if I recall correctly.",
      "votes": null
    },
    {
      "id": "817143",
      "postDate": "04/22/2020 22:00:36",
      "content": "<p><a href=\"/martingorner\">@martingorner</a> pointed out that your notebook seem to attempt to use several large models loaded in memory at the same time. That could be the issue (though we don't see the error reported in the commit logs). He shared the following discussion with me that should be relevant here:\n<a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/131045\">https://www.kaggle.com/c/flower-classification-with-tpus/discussion/131045</a></p>\n\n<p>In the notebook referenced there, the user relies on multiple models but he only keeps the predictions from each at a time, which was enough for ensembing. The trick there seems to be release memory with either GC or by resetting the TPU via <code>initialize_tpu_system(tpu)</code> in between.</p>",
      "rawMarkdown": "martingorner pointed out that your notebook seem to attempt to use several large models loaded in memory at the same time. That could be the issue (though we don't see the error reported in the commit logs). He shared the following discussion with me that should be relevant here:\nhttps://www.kaggle.com/c/flower-classification-with-tpus/discussion/131045\n\nIn the notebook referenced there, the user relies on multiple models but he only keeps the predictions from each at a time, which was enough for ensembing. The trick there seems to be release memory with either GC or by resetting the TPU via `initialize_tpu_system(tpu)` in between.",
      "votes": null
    },
    {
      "id": "817198",
      "postDate": "04/22/2020 23:51:13",
      "content": "<p>It takes around 2 hours.</p>",
      "rawMarkdown": "It takes around 2 hours.",
      "votes": null
    },
    {
      "id": "817199",
      "postDate": "04/22/2020 23:54:14",
      "content": "<p>Earlier i was using 4 models for ensembling. But from last two commit attempts i am only using only two models and there seems to be no issue during interactive execution. So it is pretty strange and my commit operation doesn't proceed from that GCE check. Meanwhile i will also try your recommendations of employing initialize_tpu_system(tpu) in between.</p>",
      "rawMarkdown": "Earlier i was using 4 models for ensembling. But from last two commit attempts i am only using only two models and there seems to be no issue during interactive execution. So it is pretty strange and my commit operation doesn't proceed from that GCE check. Meanwhile i will also try your recommendations of employing initialize_tpu_system(tpu) in between.",
      "votes": null
    },
    {
      "id": "817203",
      "postDate": "04/23/2020 00:06:47",
      "content": "<p>Yes, please try the suggestions from Martin. I also tried the Notebook in interactive mode, and for me just training the first model was taking considerable time  - i.e. after 1  hour of training I was on epoch #20 out of 50. Then I had to leave (i.e. I could no longer keep the session active manually, and so it timed out). </p>",
      "rawMarkdown": "Yes, please try the suggestions from Martin. I also tried the Notebook in interactive mode, and for me just training the first model was taking considerable time  - i.e. after 1  hour of training I was on epoch #20 out of 50. Then I had to leave (i.e. I could no longer keep the session active manually, and so it timed out).",
      "votes": null
    },
    {
      "id": "817614",
      "postDate": "04/23/2020 09:08:58",
      "content": "<p>Actually there is nothing extra in that model. It uses all the functions as available in Getting Started Notebook. And what seems strange during commit , logs does not contain execution of any of my cells. At least starting cells should be executed.</p>",
      "rawMarkdown": "Actually there is nothing extra in that model. It uses all the functions as available in Getting Started Notebook. And what seems strange during commit , logs does not contain execution of any of my cells. At least starting cells should be executed.",
      "votes": null
    },
    {
      "id": "818371",
      "postDate": "04/23/2020 20:13:11",
      "content": "<p>You could try adding the following snippet before everything else to see the log messages:\n```\nimport logging, sys</p>\n\n<p>sys.stdout = open(\"/kaggle/working/logoutput.txt\", \"w\")</p>\n\n<p>def setup_tf_logging():\n    tf.get_logger().handlers = []\n    tf.get_logger().addHandler(logging.StreamHandler(sys.stdout))</p>\n\n<p>setup_tf_logging()\n```\nPerhaps the learning rate schedule affects the execution time? I'm not sure myself. When I run your example (just the first part) it takes about 100 minutes.</p>",
      "rawMarkdown": "You could try adding the following snippet before everything else to see the log messages:\n```\nimport logging, sys\n\nsys.stdout = open(\"/kaggle/working/logoutput.txt\", \"w\")\n\ndef setup_tf_logging():\n    tf.get_logger().handlers = []\n    tf.get_logger().addHandler(logging.StreamHandler(sys.stdout))\n\nsetup_tf_logging()\n```\nPerhaps the learning rate schedule affects the execution time? I'm not sure myself. When I run your example (just the first part) it takes about 100 minutes.",
      "votes": null
    },
    {
      "id": "818385",
      "postDate": "04/23/2020 20:30:45",
      "content": "<p>Ok. I will use this code to enable logging. Can you trying committing my notebook to find out exact issue as entire notebook finishes within 3 hrs during interactive run. </p>",
      "rawMarkdown": "Ok. I will use this code to enable logging. Can you trying committing my notebook to find out exact issue as entire notebook finishes within 3 hrs during interactive run.",
      "votes": null
    },
    {
      "id": "820105",
      "postDate": "04/25/2020 06:24:32",
      "content": "<p>Just a small clarification, if I am commiting the notebook and running TPUs on a interactive session, does both use same TPU memory?</p>",
      "rawMarkdown": "Just a small clarification, if I am commiting the notebook and running TPUs on a interactive session, does both use same TPU memory?",
      "votes": null
    },
    {
      "id": "820109",
      "postDate": "04/25/2020 06:28:06",
      "content": "<p><a href=\"/ifigotin\">@ifigotin</a> the script you provided doesn't provide any output for notebooks which have failed.</p>\n\n<p>Can you check the <a href=\"https://gist.github.com/kurianbenoy/e5dc2b3a72a31aabca2712f634bb35a6\">log</a> and say why it may have failed.</p>\n\n<p>I don't use any huge models, and I just trained for 1 epoch which finished like in 30 minutes for the interactive session</p>",
      "rawMarkdown": "ifigotin the script you provided doesn't provide any output for notebooks which have failed.\n\nCan you check the [log](https://gist.github.com/kurianbenoy/e5dc2b3a72a31aabca2712f634bb35a6) and say why it may have failed.\n\nI don't use any huge models, and I just trained for 1 epoch which finished like in 30 minutes for the interactive session",
      "votes": null
    },
    {
      "id": "820652",
      "postDate": "04/25/2020 15:51:47",
      "content": "<p>Another way to debug would be to output debug messages at various points as follows: \n<code>os.system('echo '+'your debug msg')</code>.  This <em>should</em> show up in the commit logs, at least before your session is killed. \nFor example, you can output available memory at various places, and hopefully get an indication as to at which point your notebook is killed approximately.</p>",
      "rawMarkdown": "Another way to debug would be to output debug messages at various points as follows: \n`os.system('echo '+'your debug msg') `.  This *should* show up in the commit logs, at least before your session is killed. \nFor example, you can output available memory at various places, and hopefully get an indication as to at which point your notebook is killed approximately.",
      "votes": null
    },
    {
      "id": "821003",
      "postDate": "04/25/2020 20:59:57",
      "content": "<p>I used the logging. And i checked my commit failed during training of my first model after around approximately 100 mins. As timeout is fixed to be around 3 hrs, it should not exit with \"Max time exceeded error\". </p>",
      "rawMarkdown": "I used the logging. And i checked my commit failed during training of my first model after around approximately 100 mins. As timeout is fixed to be around 3 hrs, it should not exit with \"Max time exceeded error\".",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 816436,
      "author_name": "pawanverma",
      "author_url": "",
      "post_date": "04/22/2020 11:04:23",
      "content": "<p>Hi, i am trying to commit my notebook, but since last one hour same message showing GCE check failed is being displayed. I left my notebook as it is. Next day i saw commit operation was failed.\nPlease help me as i am unable to commit my version.\nIt failed with following log messages. At time 10166 error is coming but i have no idea about it. Please guide me <a href=\"/martingorner\">@martingorner</a> </p>\n\n<p>Time\nLine #\nLog Message\n3.0s\n1\n[NbConvertApp] Converting notebook notebook.ipynb to notebook\n6.1s\n2\n[NbConvertApp] Executing notebook with kernel: python3\n14.3s\n3\n2020-04-20 17:18:06.289469: I tensorflow/core/platform/cpufeatureguard.cc:142] Your CPU supports instructions that this TensorFlow binary was not compiled to use: AVX512F\n14.3s\n4\n2020-04-20 17:18:06.314414: I tensorflow/core/platform/profileutils/cpuutils.cc:94] CPU Frequency: 2000179999 Hz\n14.3s\n5\n2020-04-20 17:18:06.315135: I tensorflow/compiler/xla/service/service.cc:168] XLA service 0x558c1cfbef00 initialized for platform Host (this does not guarantee that XLA will be used). Devices:\n14.3s\n6\n2020-04-20 17:18:06.315181: I tensorflow/compiler/xla/service/service.cc:176] StreamExecutor device (0): Host, Default Version\n14.3s\n7\n2020-04-20 17:18:06.344279: I tensorflow/core/distributedruntime/rpc/grpcchannel.cc:300] Initialize GrpcChannelCache for job worker -&gt; {0 -&gt; 10.0.0.2:8470}\n14.3s\n8\n2020-04-20 17:18:06.344339: I tensorflow/core/distributedruntime/rpc/grpcchannel.cc:300] Initialize GrpcChannelCache for job localhost -&gt; {0 -&gt; localhost:30013}\n14.3s\n9\n2020-04-20 17:18:06.359935: I tensorflow/core/distributedruntime/rpc/grpcchannel.cc:300] Initialize GrpcChannelCache for job worker -&gt; {0 -&gt; 10.0.0.2:8470}\n14.3s\n10\n2020-04-20 17:18:06.359995: I tensorflow/core/distributedruntime/rpc/grpcchannel.cc:300] Initialize GrpcChannelCache for job localhost -&gt; {0 -&gt; localhost:30013}\n14.3s\n11\n2020-04-20 17:18:06.366237: I tensorflow/core/distributedruntime/rpc/grpcserverlib.cc:390] Started server with target: grpc://localhost:30013 21.6s 12 2020-04-20 17:18:13.653909: W tensorflow/core/platform/cloud/googleauthprovider.cc:178] All attempts to get a Google authentication bearer token failed, returning an empty token. Retrieving token from files failed with \"Not found: Could not locate the credentials file.\". Retrieving token from GCE failed with \"Cancelled: GCE check skipped due to presence of $NOGCECHECK environment variable.\". 21.7s 13 2020-04-20 17:18:13.704236: W tensorflow/core/platform/cloud/googleauthprovider.cc:178] All attempts to get a Google authentication bearer token failed, returning an empty token. Retrieving token from files failed with \"Not found: Could not locate the credentials file.\". Retrieving token from GCE failed with \"Cancelled: GCE check skipped due to presence of $NOGCECHECK environment variable.\". 21.7s 14 2020-04-20 17:18:13.741587: W tensorflow/core/platform/cloud/googleauthprovider.cc:178] All attempts to get a Google authentication bearer token failed, returning an empty token. Retrieving token from files failed with \"Not found: Could not locate the credentials file.\". Retrieving token from GCE failed with \"Cancelled: GCE check skipped due to presence of $NOGCECHECK environment variable.\". 10166.2s 15 2020-04-20 20:07:18.218870: W tensorflow/core/distributedruntime/eager/remotetensorhandledata.cc:75] Unable to destroy remote tensor handles. If you are running a tf.function, it usually indicates some op in the graph gets an error: RecvTensor expects a different device incarnation: 14797989987728605988 vs. 1799738036951955635. Your worker job (\"/job:tpuworker/replica:0/task:0\") was probably restarted. Check your worker job for the reason why it was restarted.\n10166.2s\n16\nAdditional GRPC error information:\n10166.2s\n17\n{\"created\":\"@1587413238.218478572\",\"description\":\"Error received from peer\",\"file\":\"external/grpc/src/core/lib/surface/call.cc\",\"fileline\":1039,\"grpcmessage\":\"RecvTensor expects a different device incarnation: 14797989987728605988 vs. 1799738036951955635. Your worker job (\"/job:tpuworker/replica:0/task:0\") was probably restarted. Check your worker job for the reason why it was restarted.\",\"grpcstatus\":10}\n10169.6s\n18\n2020-04-20 20:07:21.664715: W ./tensorflow/core/distributedruntime/eager/destroytensorhandlenode.h:79] Ignoring an error encountered when deleting remote tensors handles: Invalid argument: Unable to find a contextid matching the specified one (12193600234898793608). Perhaps the worker was restarted, or the context was GC'd? 10169.6s 19 Additional GRPC error information: 10169.6s 20 {\"created\":\"@1587413241.664555067\",\"description\":\"Error received from peer\",\"file\":\"external/grpc/src/core/lib/surface/call.cc\",\"fileline\":1039,\"grpcmessage\":\"Unable to find a contextid matching the specified one (12193600234898793608). Perhaps the worker was restarted, or the context was GC'd?\",\"grpc_status\":3}\n10180.0s\n21\n[NbConvertApp] Writing 158697 bytes to notebook.ipynb\n10180.9s\n22\n[NbConvertApp] Converting notebook notebook.ipynb to html\n10182.4s\n23\n[NbConvertApp] Writing 508223 bytes to results.html\n10182.4s\n24\n10182.4s\n26\nComplete. Exited with code 0.</p>",
      "votes": null,
      "replies": [
        {
          "id": 816789,
          "author_name": "ifigotin",
          "author_url": "",
          "post_date": "04/22/2020 15:39:24",
          "content": "<p>The GCE check warning messages should be harmless and  can be ignored. The next version of Tensorflow will no longer print these. The message basically says that TF will not use any credentials to access public GCS object.\nPerhaps the real error is <code>Unable to destroy remote tensor handles. If you are running a tf.function, it usually indicates some op in the graph gets an error: RecvTensor expects a different device incarnation: 14797989987728605988 vs. 1799738036951955635. Your worker job (\"/job:tpuworker/replica:0/task:0\") was probably restarted. Check your worker job for the reason why it was restarted.</code>?  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 816971,
          "author_name": "pawanverma",
          "author_url": "",
          "post_date": "04/22/2020 18:36:39",
          "content": "<p>There is no problem when i run all cells in interactive mode but only during commit i face this problem. Moreover for continuous 2 hrs i monitored commit operation and it kept showing the harmless GCE check warning message. And i have no idea how worker job restarts. I simply commit and wait.\nI have wasted around 10 hrs of my TPU quota for trying to commit this notebook.\nPlease check if you can find any issue.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 816994,
          "author_name": "ifigotin",
          "author_url": "",
          "post_date": "04/22/2020 18:56:28",
          "content": "<p>Yep, the Tensorflow library will probably keep outputting \"GCE check\" warnings every time the TPU accesses the data from Google Cloud Storage. Does the notebook output you shared earlier contain any more messages? Please, share if that's the case.</p>\n\n<p>Also, when you look at the commit output URL, you should see something like scriptVersionId=[numberHere]. Could you please share the id with us so that we can try to find out more info?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 817020,
          "author_name": "pawanverma",
          "author_url": "",
          "post_date": "04/22/2020 19:22:05",
          "content": "<p><a href=\"https://www.kaggle.com/pawanverma/flowers-efficient-net-dense-xception/log?scriptVersionId=32431497\">https://www.kaggle.com/pawanverma/flowers-efficient-net-dense-xception/log?scriptVersionId=32431497</a>.\nI have pasted entire log file of my last commit attempts. Thanks in advance for all the help.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 817122,
          "author_name": "ifigotin",
          "author_url": "",
          "post_date": "04/22/2020 21:22:22",
          "content": "<p>Btw, <a href=\"/pawanverma\">@pawanverma</a>, when you ran your session interactively top to bottom, how long did it take?\nThe commit sessions will time out after 3 hours if I recall correctly.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 817143,
          "author_name": "ifigotin",
          "author_url": "",
          "post_date": "04/22/2020 22:00:36",
          "content": "<p><a href=\"/martingorner\">@martingorner</a> pointed out that your notebook seem to attempt to use several large models loaded in memory at the same time. That could be the issue (though we don't see the error reported in the commit logs). He shared the following discussion with me that should be relevant here:\n<a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/131045\">https://www.kaggle.com/c/flower-classification-with-tpus/discussion/131045</a></p>\n\n<p>In the notebook referenced there, the user relies on multiple models but he only keeps the predictions from each at a time, which was enough for ensembing. The trick there seems to be release memory with either GC or by resetting the TPU via <code>initialize_tpu_system(tpu)</code> in between.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 817198,
          "author_name": "pawanverma",
          "author_url": "",
          "post_date": "04/22/2020 23:51:13",
          "content": "<p>It takes around 2 hours.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 817199,
          "author_name": "pawanverma",
          "author_url": "",
          "post_date": "04/22/2020 23:54:14",
          "content": "<p>Earlier i was using 4 models for ensembling. But from last two commit attempts i am only using only two models and there seems to be no issue during interactive execution. So it is pretty strange and my commit operation doesn't proceed from that GCE check. Meanwhile i will also try your recommendations of employing initialize_tpu_system(tpu) in between.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 817203,
          "author_name": "ifigotin",
          "author_url": "",
          "post_date": "04/23/2020 00:06:47",
          "content": "<p>Yes, please try the suggestions from Martin. I also tried the Notebook in interactive mode, and for me just training the first model was taking considerable time  - i.e. after 1  hour of training I was on epoch #20 out of 50. Then I had to leave (i.e. I could no longer keep the session active manually, and so it timed out). </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 817614,
          "author_name": "pawanverma",
          "author_url": "",
          "post_date": "04/23/2020 09:08:58",
          "content": "<p>Actually there is nothing extra in that model. It uses all the functions as available in Getting Started Notebook. And what seems strange during commit , logs does not contain execution of any of my cells. At least starting cells should be executed.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 818371,
          "author_name": "ifigotin",
          "author_url": "",
          "post_date": "04/23/2020 20:13:11",
          "content": "<p>You could try adding the following snippet before everything else to see the log messages:\n```\nimport logging, sys</p>\n\n<p>sys.stdout = open(\"/kaggle/working/logoutput.txt\", \"w\")</p>\n\n<p>def setup_tf_logging():\n    tf.get_logger().handlers = []\n    tf.get_logger().addHandler(logging.StreamHandler(sys.stdout))</p>\n\n<p>setup_tf_logging()\n```\nPerhaps the learning rate schedule affects the execution time? I'm not sure myself. When I run your example (just the first part) it takes about 100 minutes.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 818385,
          "author_name": "pawanverma",
          "author_url": "",
          "post_date": "04/23/2020 20:30:45",
          "content": "<p>Ok. I will use this code to enable logging. Can you trying committing my notebook to find out exact issue as entire notebook finishes within 3 hrs during interactive run. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 820109,
          "author_name": "kurianbenoy",
          "author_url": "",
          "post_date": "04/25/2020 06:28:06",
          "content": "<p><a href=\"/ifigotin\">@ifigotin</a> the script you provided doesn't provide any output for notebooks which have failed.</p>\n\n<p>Can you check the <a href=\"https://gist.github.com/kurianbenoy/e5dc2b3a72a31aabca2712f634bb35a6\">log</a> and say why it may have failed.</p>\n\n<p>I don't use any huge models, and I just trained for 1 epoch which finished like in 30 minutes for the interactive session</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 820652,
          "author_name": "ifigotin",
          "author_url": "",
          "post_date": "04/25/2020 15:51:47",
          "content": "<p>Another way to debug would be to output debug messages at various points as follows: \n<code>os.system('echo '+'your debug msg')</code>.  This <em>should</em> show up in the commit logs, at least before your session is killed. \nFor example, you can output available memory at various places, and hopefully get an indication as to at which point your notebook is killed approximately.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 821003,
          "author_name": "pawanverma",
          "author_url": "",
          "post_date": "04/25/2020 20:59:57",
          "content": "<p>I used the logging. And i checked my commit failed during training of my first model after around approximately 100 mins. As timeout is fixed to be around 3 hrs, it should not exit with \"Max time exceeded error\". </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 816671,
      "author_name": "herbison",
      "author_url": "",
      "post_date": "04/22/2020 13:59:52",
      "content": "<p>There can potentially be many reasons for a commit to take longer. For example, your commit may not even start executing immediately when you press the button. We have to schedule the execution on a machine, and if there is enough demand all at once, it may be queued for awhile before it starts. We do our best to avoid queuing at all times, but sometimes TPUs are so popular all at once that we have to add a little wait time.</p>\n\n<p>But that's assuming it's a scheduling problem, it could also be that your code runs into a rare condition that just takes longer (anything involving IO/network will have some level of variance in it due to to variable throughput rates).</p>",
      "votes": null,
      "replies": [
        {
          "id": 816733,
          "author_name": "pawanverma",
          "author_url": "",
          "post_date": "04/22/2020 14:41:53",
          "content": "<p>I tried for committing for third time in three day. But it again failed with error due to time out errot. When i run my notebook in interactive mode, all cells successfully execute so it seems there is some problem with scheduling TPU for commit operation.\nI request you to please check this issue as i am unable to submit my notebook and submission deadline is approaching.\nAttached are the latest logs for commit operation:\nTime\nLine #\nLog Message\n3.6s\n1\n[NbConvertApp] Converting notebook <strong>notebook</strong>.ipynb to notebook\n6.9s\n2\n[NbConvertApp] Executing notebook with kernel: python3\n16.0s\n3\n2020-04-22 11:08:42.138682: I tensorflow/core/platform/cpu_feature_guard.cc:142] Your CPU supports instructions that this TensorFlow binary was not compiled to use: AVX512F\n16.0s\n4\n2020-04-22 11:08:42.171854: I tensorflow/core/platform/profile_utils/cpu_utils.cc:94] CPU Frequency: 2000170000 Hz\n16.0s\n5\n2020-04-22 11:08:42.173379: I tensorflow/compiler/xla/service/service.cc:168] XLA service 0x55a1c3d9c950 initialized for platform Host (this does not guarantee that XLA will be used). Devices:\n16.0s\n6\n2020-04-22 11:08:42.173435: I tensorflow/compiler/xla/service/service.cc:176]   StreamExecutor device (0): Host, Default Version\n16.1s\n7\n2020-04-22 11:08:42.198828: I tensorflow/core/distributed_runtime/rpc/grpc_channel.cc:300] Initialize GrpcChannelCache for job worker -&gt; {0 -&gt; 10.0.0.2:8470}\n16.1s\n8\n2020-04-22 11:08:42.198889: I tensorflow/core/distributed_runtime/rpc/grpc_channel.cc:300] Initialize GrpcChannelCache for job localhost -&gt; {0 -&gt; localhost:30012}\n16.1s\n9\n2020-04-22 11:08:42.214317: I tensorflow/core/distributed_runtime/rpc/grpc_channel.cc:300] Initialize GrpcChannelCache for job worker -&gt; {0 -&gt; 10.0.0.2:8470}\n16.1s\n10\n2020-04-22 11:08:42.214390: I tensorflow/core/distributed_runtime/rpc/grpc_channel.cc:300] Initialize GrpcChannelCache for job localhost -&gt; {0 -&gt; localhost:30012}\n16.1s\n11\n2020-04-22 11:08:42.219715: I tensorflow/core/distributed_runtime/rpc/grpc_server_lib.cc:390] Started server with target: grpc://localhost:30012\n23.2s\n12\n2020-04-22 11:08:49.292527: W tensorflow/core/platform/cloud/google_auth_provider.cc:178] All attempts to get a Google authentication bearer token failed, returning an empty token. Retrieving token from files failed with \"Not found: Could not locate the credentials file.\". Retrieving token from GCE failed with \"Cancelled: GCE check skipped due to presence of $NO_GCE_CHECK environment variable.\".\n23.2s\n13\n2020-04-22 11:08:49.342698: W tensorflow/core/platform/cloud/google_auth_provider.cc:178] All attempts to get a Google authentication bearer token failed, returning an empty token. Retrieving token from files failed with \"Not found: Could not locate the credentials file.\". Retrieving token from GCE failed with \"Cancelled: GCE check skipped due to presence of $NO_GCE_CHECK environment variable.\".\n23.2s\n14\n2020-04-22 11:08:49.376955: W tensorflow/core/platform/cloud/google_auth_provider.cc:178] All attempts to get a Google authentication bearer token failed, returning an empty token. Retrieving token from files failed with \"Not found: Could not locate the credentials file.\". Retrieving token from GCE failed with \"Cancelled: GCE check skipped due to presence of $NO_GCE_CHECK environment variable.\".\n23.2s\n15\n23.2s\n17\nFailed. Exited with code 137.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 816737,
          "author_name": "herbison",
          "author_url": "",
          "post_date": "04/22/2020 14:44:20",
          "content": "<p>Scheduling does not affect the timeout, only when it starts, the timer starts ticking when execution starts, not when you submit it.</p>\n\n<p>I'll pass this along to some people who know TPUs better than me to see if they know what's going on.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 816967,
          "author_name": "pawanverma",
          "author_url": "",
          "post_date": "04/22/2020 18:33:09",
          "content": "<p>Thanks. That will be very helpful because this problem is recurring every time after commit. But if i run all cells without commit then there seems to be no problem.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 820105,
          "author_name": "kurianbenoy",
          "author_url": "",
          "post_date": "04/25/2020 06:24:32",
          "content": "<p>Just a small clarification, if I am commiting the notebook and running TPUs on a interactive session, does both use same TPU memory?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "810238": "Hi, I am new to kaggle and this will be my first submission. I am committing my notebook but from last 40 mins it is saving.\nMy notebook usually runs in 20-30 mins. Is it usual to take so much time for committing?\nOr is there any other way also",
    "816436": "Hi, i am trying to commit my notebook, but since last one hour same message showing GCE check failed is being displayed. I left my notebook as it is. Next day i saw commit operation was failed.\nPlease help me as i am unable to commit my version.\nIt failed with following log messages. At time 10166 error is coming but i have no idea about it. Please guide me @martingorner \n\nTime\nLine #\nLog Message\n3.0s\n1\n[NbConvertApp] Converting notebook notebook.ipynb to notebook\n6.1s\n2\n[NbConvertApp] Executing notebook with kernel: python3\n14.3s\n3\n2020-04-20 17:18:06.289469: I tensorflow/core/platform/cpufeatureguard.cc:142] Your CPU supports instructions that this TensorFlow binary was not compiled to use: AVX512F\n14.3s\n4\n2020-04-20 17:18:06.314414: I tensorflow/core/platform/profileutils/cpuutils.cc:94] CPU Frequency: 2000179999 Hz\n14.3s\n5\n2020-04-20 17:18:06.315135: I tensorflow/compiler/xla/service/service.cc:168] XLA service 0x558c1cfbef00 initialized for platform Host (this does not guarantee that XLA will be used). Devices:\n14.3s\n6\n2020-04-20 17:18:06.315181: I tensorflow/compiler/xla/service/service.cc:176] StreamExecutor device (0): Host, Default Version\n14.3s\n7\n2020-04-20 17:18:06.344279: I tensorflow/core/distributedruntime/rpc/grpcchannel.cc:300] Initialize GrpcChannelCache for job worker -&gt; {0 -&gt; 10.0.0.2:8470}\n14.3s\n8\n2020-04-20 17:18:06.344339: I tensorflow/core/distributedruntime/rpc/grpcchannel.cc:300] Initialize GrpcChannelCache for job localhost -&gt; {0 -&gt; localhost:30013}\n14.3s\n9\n2020-04-20 17:18:06.359935: I tensorflow/core/distributedruntime/rpc/grpcchannel.cc:300] Initialize GrpcChannelCache for job worker -&gt; {0 -&gt; 10.0.0.2:8470}\n14.3s\n10\n2020-04-20 17:18:06.359995: I tensorflow/core/distributedruntime/rpc/grpcchannel.cc:300] Initialize GrpcChannelCache for job localhost -&gt; {0 -&gt; localhost:30013}\n14.3s\n11\n2020-04-20 17:18:06.366237: I tensorflow/core/distributedruntime/rpc/grpcserverlib.cc:390] Started server with target: grpc://localhost:30013 21.6s 12 2020-04-20 17:18:13.653909: W tensorflow/core/platform/cloud/googleauthprovider.cc:178] All attempts to get a Google authentication bearer token failed, returning an empty token. Retrieving token from files failed with \"Not found: Could not locate the credentials file.\". Retrieving token from GCE failed with \"Cancelled: GCE check skipped due to presence of $NOGCECHECK environment variable.\". 21.7s 13 2020-04-20 17:18:13.704236: W tensorflow/core/platform/cloud/googleauthprovider.cc:178] All attempts to get a Google authentication bearer token failed, returning an empty token. Retrieving token from files failed with \"Not found: Could not locate the credentials file.\". Retrieving token from GCE failed with \"Cancelled: GCE check skipped due to presence of $NOGCECHECK environment variable.\". 21.7s 14 2020-04-20 17:18:13.741587: W tensorflow/core/platform/cloud/googleauthprovider.cc:178] All attempts to get a Google authentication bearer token failed, returning an empty token. Retrieving token from files failed with \"Not found: Could not locate the credentials file.\". Retrieving token from GCE failed with \"Cancelled: GCE check skipped due to presence of $NOGCECHECK environment variable.\". 10166.2s 15 2020-04-20 20:07:18.218870: W tensorflow/core/distributedruntime/eager/remotetensorhandledata.cc:75] Unable to destroy remote tensor handles. If you are running a tf.function, it usually indicates some op in the graph gets an error: RecvTensor expects a different device incarnation: 14797989987728605988 vs. 1799738036951955635. Your worker job (\"/job:tpuworker/replica:0/task:0\") was probably restarted. Check your worker job for the reason why it was restarted.\n10166.2s\n16\nAdditional GRPC error information:\n10166.2s\n17\n{\"created\":\"@1587413238.218478572\",\"description\":\"Error received from peer\",\"file\":\"external/grpc/src/core/lib/surface/call.cc\",\"fileline\":1039,\"grpcmessage\":\"RecvTensor expects a different device incarnation: 14797989987728605988 vs. 1799738036951955635. Your worker job (\"/job:tpuworker/replica:0/task:0\") was probably restarted. Check your worker job for the reason why it was restarted.\",\"grpcstatus\":10}\n10169.6s\n18\n2020-04-20 20:07:21.664715: W ./tensorflow/core/distributedruntime/eager/destroytensorhandlenode.h:79] Ignoring an error encountered when deleting remote tensors handles: Invalid argument: Unable to find a contextid matching the specified one (12193600234898793608). Perhaps the worker was restarted, or the context was GC'd? 10169.6s 19 Additional GRPC error information: 10169.6s 20 {\"created\":\"@1587413241.664555067\",\"description\":\"Error received from peer\",\"file\":\"external/grpc/src/core/lib/surface/call.cc\",\"fileline\":1039,\"grpcmessage\":\"Unable to find a contextid matching the specified one (12193600234898793608). Perhaps the worker was restarted, or the context was GC'd?\",\"grpc_status\":3}\n10180.0s\n21\n[NbConvertApp] Writing 158697 bytes to notebook.ipynb\n10180.9s\n22\n[NbConvertApp] Converting notebook notebook.ipynb to html\n10182.4s\n23\n[NbConvertApp] Writing 508223 bytes to results.html\n10182.4s\n24\n10182.4s\n26\nComplete. Exited with code 0.",
    "816671": "There can potentially be many reasons for a commit to take longer. For example, your commit may not even start executing immediately when you press the button. We have to schedule the execution on a machine, and if there is enough demand all at once, it may be queued for awhile before it starts. We do our best to avoid queuing at all times, but sometimes TPUs are so popular all at once that we have to add a little wait time.\n\nBut that's assuming it's a scheduling problem, it could also be that your code runs into a rare condition that just takes longer (anything involving IO/network will have some level of variance in it due to to variable throughput rates).",
    "816733": "I tried for committing for third time in three day. But it again failed with error due to time out errot. When i run my notebook in interactive mode, all cells successfully execute so it seems there is some problem with scheduling TPU for commit operation.\nI request you to please check this issue as i am unable to submit my notebook and submission deadline is approaching.\nAttached are the latest logs for commit operation:\nTime\nLine #\nLog Message\n3.6s\n1\n[NbConvertApp] Converting notebook __notebook__.ipynb to notebook\n6.9s\n2\n[NbConvertApp] Executing notebook with kernel: python3\n16.0s\n3\n2020-04-22 11:08:42.138682: I tensorflow/core/platform/cpu_feature_guard.cc:142] Your CPU supports instructions that this TensorFlow binary was not compiled to use: AVX512F\n16.0s\n4\n2020-04-22 11:08:42.171854: I tensorflow/core/platform/profile_utils/cpu_utils.cc:94] CPU Frequency: 2000170000 Hz\n16.0s\n5\n2020-04-22 11:08:42.173379: I tensorflow/compiler/xla/service/service.cc:168] XLA service 0x55a1c3d9c950 initialized for platform Host (this does not guarantee that XLA will be used). Devices:\n16.0s\n6\n2020-04-22 11:08:42.173435: I tensorflow/compiler/xla/service/service.cc:176]   StreamExecutor device (0): Host, Default Version\n16.1s\n7\n2020-04-22 11:08:42.198828: I tensorflow/core/distributed_runtime/rpc/grpc_channel.cc:300] Initialize GrpcChannelCache for job worker -&gt; {0 -&gt; 10.0.0.2:8470}\n16.1s\n8\n2020-04-22 11:08:42.198889: I tensorflow/core/distributed_runtime/rpc/grpc_channel.cc:300] Initialize GrpcChannelCache for job localhost -&gt; {0 -&gt; localhost:30012}\n16.1s\n9\n2020-04-22 11:08:42.214317: I tensorflow/core/distributed_runtime/rpc/grpc_channel.cc:300] Initialize GrpcChannelCache for job worker -&gt; {0 -&gt; 10.0.0.2:8470}\n16.1s\n10\n2020-04-22 11:08:42.214390: I tensorflow/core/distributed_runtime/rpc/grpc_channel.cc:300] Initialize GrpcChannelCache for job localhost -&gt; {0 -&gt; localhost:30012}\n16.1s\n11\n2020-04-22 11:08:42.219715: I tensorflow/core/distributed_runtime/rpc/grpc_server_lib.cc:390] Started server with target: grpc://localhost:30012\n23.2s\n12\n2020-04-22 11:08:49.292527: W tensorflow/core/platform/cloud/google_auth_provider.cc:178] All attempts to get a Google authentication bearer token failed, returning an empty token. Retrieving token from files failed with \"Not found: Could not locate the credentials file.\". Retrieving token from GCE failed with \"Cancelled: GCE check skipped due to presence of $NO_GCE_CHECK environment variable.\".\n23.2s\n13\n2020-04-22 11:08:49.342698: W tensorflow/core/platform/cloud/google_auth_provider.cc:178] All attempts to get a Google authentication bearer token failed, returning an empty token. Retrieving token from files failed with \"Not found: Could not locate the credentials file.\". Retrieving token from GCE failed with \"Cancelled: GCE check skipped due to presence of $NO_GCE_CHECK environment variable.\".\n23.2s\n14\n2020-04-22 11:08:49.376955: W tensorflow/core/platform/cloud/google_auth_provider.cc:178] All attempts to get a Google authentication bearer token failed, returning an empty token. Retrieving token from files failed with \"Not found: Could not locate the credentials file.\". Retrieving token from GCE failed with \"Cancelled: GCE check skipped due to presence of $NO_GCE_CHECK environment variable.\".\n23.2s\n15\n23.2s\n17\nFailed. Exited with code 137.",
    "816737": "Scheduling does not affect the timeout, only when it starts, the timer starts ticking when execution starts, not when you submit it.\n\nI'll pass this along to some people who know TPUs better than me to see if they know what's going on.",
    "816789": "The GCE check warning messages should be harmless and  can be ignored. The next version of Tensorflow will no longer print these. The message basically says that TF will not use any credentials to access public GCS object.\nPerhaps the real error is `Unable to destroy remote tensor handles. If you are running a tf.function, it usually indicates some op in the graph gets an error: RecvTensor expects a different device incarnation: 14797989987728605988 vs. 1799738036951955635. Your worker job (\"/job:tpuworker/replica:0/task:0\") was probably restarted. Check your worker job for the reason why it was restarted.`?",
    "816967": "Thanks. That will be very helpful because this problem is recurring every time after commit. But if i run all cells without commit then there seems to be no problem.",
    "816971": "There is no problem when i run all cells in interactive mode but only during commit i face this problem. Moreover for continuous 2 hrs i monitored commit operation and it kept showing the harmless GCE check warning message. And i have no idea how worker job restarts. I simply commit and wait.\nI have wasted around 10 hrs of my TPU quota for trying to commit this notebook.\nPlease check if you can find any issue.",
    "816994": "Yep, the Tensorflow library will probably keep outputting \"GCE check\" warnings every time the TPU accesses the data from Google Cloud Storage. Does the notebook output you shared earlier contain any more messages? Please, share if that's the case.\n\nAlso, when you look at the commit output URL, you should see something like scriptVersionId=[numberHere]. Could you please share the id with us so that we can try to find out more info?",
    "817020": "https://www.kaggle.com/pawanverma/flowers-efficient-net-dense-xception/log?scriptVersionId=32431497.\nI have pasted entire log file of my last commit attempts. Thanks in advance for all the help.",
    "817122": "Btw, @pawanverma, when you ran your session interactively top to bottom, how long did it take?\nThe commit sessions will time out after 3 hours if I recall correctly.",
    "817143": "martingorner pointed out that your notebook seem to attempt to use several large models loaded in memory at the same time. That could be the issue (though we don't see the error reported in the commit logs). He shared the following discussion with me that should be relevant here:\nhttps://www.kaggle.com/c/flower-classification-with-tpus/discussion/131045\n\nIn the notebook referenced there, the user relies on multiple models but he only keeps the predictions from each at a time, which was enough for ensembing. The trick there seems to be release memory with either GC or by resetting the TPU via `initialize_tpu_system(tpu)` in between.",
    "817198": "It takes around 2 hours.",
    "817199": "Earlier i was using 4 models for ensembling. But from last two commit attempts i am only using only two models and there seems to be no issue during interactive execution. So it is pretty strange and my commit operation doesn't proceed from that GCE check. Meanwhile i will also try your recommendations of employing initialize_tpu_system(tpu) in between.",
    "817203": "Yes, please try the suggestions from Martin. I also tried the Notebook in interactive mode, and for me just training the first model was taking considerable time  - i.e. after 1  hour of training I was on epoch #20 out of 50. Then I had to leave (i.e. I could no longer keep the session active manually, and so it timed out).",
    "817614": "Actually there is nothing extra in that model. It uses all the functions as available in Getting Started Notebook. And what seems strange during commit , logs does not contain execution of any of my cells. At least starting cells should be executed.",
    "818371": "You could try adding the following snippet before everything else to see the log messages:\n```\nimport logging, sys\n\nsys.stdout = open(\"/kaggle/working/logoutput.txt\", \"w\")\n\ndef setup_tf_logging():\n    tf.get_logger().handlers = []\n    tf.get_logger().addHandler(logging.StreamHandler(sys.stdout))\n\nsetup_tf_logging()\n```\nPerhaps the learning rate schedule affects the execution time? I'm not sure myself. When I run your example (just the first part) it takes about 100 minutes.",
    "818385": "Ok. I will use this code to enable logging. Can you trying committing my notebook to find out exact issue as entire notebook finishes within 3 hrs during interactive run.",
    "820105": "Just a small clarification, if I am commiting the notebook and running TPUs on a interactive session, does both use same TPU memory?",
    "820109": "ifigotin the script you provided doesn't provide any output for notebooks which have failed.\n\nCan you check the [log](https://gist.github.com/kurianbenoy/e5dc2b3a72a31aabca2712f634bb35a6) and say why it may have failed.\n\nI don't use any huge models, and I just trained for 1 epoch which finished like in 30 minutes for the interactive session",
    "820652": "Another way to debug would be to output debug messages at various points as follows: \n`os.system('echo '+'your debug msg') `.  This *should* show up in the commit logs, at least before your session is killed. \nFor example, you can output available memory at various places, and hopefully get an indication as to at which point your notebook is killed approximately.",
    "821003": "I used the logging. And i checked my commit failed during training of my first model after around approximately 100 mins. As timeout is fixed to be around 3 hrs, it should not exit with \"Max time exceeded error\"."
  },
  "source": "meta"
}