{
  "id": 216408,
  "title": "curious about this error (previously working) - NOW FIXED",
  "url": "/competitions/rfcx-species-audio-detection/discussion/216408",
  "author_name": "",
  "post_date": "2021-02-02T16:23:11.704637100Z",
  "votes": 9,
  "comment_count": 15,
  "views": 0,
  "content": "<p>&lt;&lt;&lt;EDIT - This is now fixable thanks to the workaround suggested by Martin. For dataset pipelines we DONT need to explicitly specify tf graph mode with the decorator <a href=\"https://www.kaggle.com/tf.function\" target=\"_blank\">@tf.function</a>. By default dataset pipeline operates only in graph mode. Commenting out that one line fixes the issue. I also compared the prev and current times and the execution time is NOT impacted due to the commenting. Happy TPU'ing :)</p>\n<p>I would like to thank <a href=\"https://www.kaggle.com/mgornergoogle\" target=\"_blank\">@mgornergoogle</a> for helping with this issue. This has been quite unlike my prev experienice in the RiiiD competition and I am pleasantly surprised with the support extended.&gt;&gt;&gt;</p>\n<p>With over 300+ direct edits and many more 2nd and 3rd level edits, I am sure the foll kernel <a href=\"https://www.kaggle.com/yosshi999/rfcx-train-resnet50-with-tpu\" target=\"_blank\">https://www.kaggle.com/yosshi999/rfcx-train-resnet50-with-tpu</a> by <a href=\"https://www.kaggle.com/yosshi999\" target=\"_blank\">@yosshi999</a> would be the most popular one in this competition</p>\n<p>Here is an interesting observation. The below block which was working earlier fails with this \"failed to connect to all addresses…\" error (you can just do a edit and run to simulate this):</p>\n<pre><code>plt.figure(figsize=(16, 4))\nfor i, (inp, targ) in enumerate(annot_dataset.map(_preprocess).take(6)):\n    plt.subplot(2,3,i+1)\n    plt.imshow(inp.numpy()[:,:,0])\n    t = targ.numpy()\n    if t.sum() == 0:\n        plt.title(f'FP')\n    else:\n        plt.title(f'{t.nonzero()[0]}')\n    plt.colorbar()\nplt.show() \n</code></pre>\n<p>Unfortunately not too much info exists about the error on SFO. I could get it working by commenting the specaugment and the gaussian noise preprocessing in the block preceeding it. </p>\n<p>Another strange behaviour is that a simple workaround is to comment out the whole block..It is anyway a exploratory block and does not impact final processing. So subsequent dataset mappings work nicely with all the preprocessing (including specaug and gaussian noise)</p>\n<p>But interested to know why this error happens. Was some software upgraded recently @ kaggle env which is causing this. We can explore some interesting combinations also. For e..g. changing the brightness works but a simple randomfilp fails.</p>\n<p>If anybody has any pointers, do let me know</p>\n<p>Below is the full error:<br>\nUnavailableError: failed to connect to all addresses<br>\nAdditional GRPC error information from remote target /job:localhost/replica:0/task:0/device:CPU:0:<br>\n:{\"created\":\"@1612282112.116200424\",\"description\":\"Failed to pick subchannel\",\"file\":\"third_party/grpc/src/core/ext/filters/client_channel/client_channel.cc\",\"file_line\":4143,\"referenced_errors\":[{\"created\":\"@1612282112.116196618\",\"description\":\"failed to connect to all addresses\",\"file\":\"third_party/grpc/src/core/ext/filters/client_channel/lb_policy/pick_first/pick_first.cc\",\"file_line\":398,\"grpc_status\":14}]}<br>\n     [[{{node StatefulPartitionedCall}}]]</p>",
  "messages": [
    {
      "id": "1182881",
      "postDate": "02/02/2021 16:23:11",
      "content": "<p>&lt;&lt;&lt;EDIT - This is now fixable thanks to the workaround suggested by Martin. For dataset pipelines we DONT need to explicitly specify tf graph mode with the decorator <a href=\"https://www.kaggle.com/tf.function\" target=\"_blank\">@tf.function</a>. By default dataset pipeline operates only in graph mode. Commenting out that one line fixes the issue. I also compared the prev and current times and the execution time is NOT impacted due to the commenting. Happy TPU'ing :)</p>\n<p>I would like to thank <a href=\"https://www.kaggle.com/mgornergoogle\" target=\"_blank\">@mgornergoogle</a> for helping with this issue. This has been quite unlike my prev experienice in the RiiiD competition and I am pleasantly surprised with the support extended.&gt;&gt;&gt;</p>\n<p>With over 300+ direct edits and many more 2nd and 3rd level edits, I am sure the foll kernel <a href=\"https://www.kaggle.com/yosshi999/rfcx-train-resnet50-with-tpu\" target=\"_blank\">https://www.kaggle.com/yosshi999/rfcx-train-resnet50-with-tpu</a> by <a href=\"https://www.kaggle.com/yosshi999\" target=\"_blank\">@yosshi999</a> would be the most popular one in this competition</p>\n<p>Here is an interesting observation. The below block which was working earlier fails with this \"failed to connect to all addresses…\" error (you can just do a edit and run to simulate this):</p>\n<pre><code>plt.figure(figsize=(16, 4))\nfor i, (inp, targ) in enumerate(annot_dataset.map(_preprocess).take(6)):\n    plt.subplot(2,3,i+1)\n    plt.imshow(inp.numpy()[:,:,0])\n    t = targ.numpy()\n    if t.sum() == 0:\n        plt.title(f'FP')\n    else:\n        plt.title(f'{t.nonzero()[0]}')\n    plt.colorbar()\nplt.show() \n</code></pre>\n<p>Unfortunately not too much info exists about the error on SFO. I could get it working by commenting the specaugment and the gaussian noise preprocessing in the block preceeding it. </p>\n<p>Another strange behaviour is that a simple workaround is to comment out the whole block..It is anyway a exploratory block and does not impact final processing. So subsequent dataset mappings work nicely with all the preprocessing (including specaug and gaussian noise)</p>\n<p>But interested to know why this error happens. Was some software upgraded recently @ kaggle env which is causing this. We can explore some interesting combinations also. For e..g. changing the brightness works but a simple randomfilp fails.</p>\n<p>If anybody has any pointers, do let me know</p>\n<p>Below is the full error:<br>\nUnavailableError: failed to connect to all addresses<br>\nAdditional GRPC error information from remote target /job:localhost/replica:0/task:0/device:CPU:0:<br>\n:{\"created\":\"@1612282112.116200424\",\"description\":\"Failed to pick subchannel\",\"file\":\"third_party/grpc/src/core/ext/filters/client_channel/client_channel.cc\",\"file_line\":4143,\"referenced_errors\":[{\"created\":\"@1612282112.116196618\",\"description\":\"failed to connect to all addresses\",\"file\":\"third_party/grpc/src/core/ext/filters/client_channel/lb_policy/pick_first/pick_first.cc\",\"file_line\":398,\"grpc_status\":14}]}<br>\n     [[{{node StatefulPartitionedCall}}]]</p>",
      "rawMarkdown": "<<<EDIT - This is now fixable thanks to the workaround suggested by Martin. For dataset pipelines we DONT need to explicitly specify tf graph mode with the decorator @tf.function. By default dataset pipeline operates only in graph mode. Commenting out that one line fixes the issue. I also compared the prev and current times and the execution time is NOT impacted due to the commenting. Happy TPU'ing :)\n\nI would like to thank @mgornergoogle for helping with this issue. This has been quite unlike my prev experienice in the RiiiD competition and I am pleasantly surprised with the support extended.>>>\n\nWith over 300+ direct edits and many more 2nd and 3rd level edits, I am sure the foll kernel https://www.kaggle.com/yosshi999/rfcx-train-resnet50-with-tpu by @yosshi999 would be the most popular one in this competition\n\nHere is an interesting observation. The below block which was working earlier fails with this \"failed to connect to all addresses...\" error (you can just do a edit and run to simulate this):\n\n```\nplt.figure(figsize=(16, 4))\nfor i, (inp, targ) in enumerate(annot_dataset.map(_preprocess).take(6)):\n    plt.subplot(2,3,i+1)\n    plt.imshow(inp.numpy()[:,:,0])\n    t = targ.numpy()\n    if t.sum() == 0:\n        plt.title(f'FP')\n    else:\n        plt.title(f'{t.nonzero()[0]}')\n    plt.colorbar()\nplt.show() \n```\n\nUnfortunately not too much info exists about the error on SFO. I could get it working by commenting the specaugment and the gaussian noise preprocessing in the block preceeding it. \n\nAnother strange behaviour is that a simple workaround is to comment out the whole block..It is anyway a exploratory block and does not impact final processing. So subsequent dataset mappings work nicely with all the preprocessing (including specaug and gaussian noise)\n\nBut interested to know why this error happens. Was some software upgraded recently @ kaggle env which is causing this. We can explore some interesting combinations also. For e..g. changing the brightness works but a simple randomfilp fails.\n\nIf anybody has any pointers, do let me know\n\nBelow is the full error:\nUnavailableError: failed to connect to all addresses\nAdditional GRPC error information from remote target /job:localhost/replica:0/task:0/device:CPU:0:\n:{\"created\":\"@1612282112.116200424\",\"description\":\"Failed to pick subchannel\",\"file\":\"third_party/grpc/src/core/ext/filters/client_channel/client_channel.cc\",\"file_line\":4143,\"referenced_errors\":[{\"created\":\"@1612282112.116196618\",\"description\":\"failed to connect to all addresses\",\"file\":\"third_party/grpc/src/core/ext/filters/client_channel/lb_policy/pick_first/pick_first.cc\",\"file_line\":398,\"grpc_status\":14}]}\n\t [[{{node StatefulPartitionedCall}}]]",
      "votes": null
    },
    {
      "id": "1185313",
      "postDate": "02/04/2021 05:38:00",
      "content": "<p>Hi I am having a similar error did you recieve any answers?</p>",
      "rawMarkdown": "Hi I am having a similar error did you recieve any answers?",
      "votes": null
    },
    {
      "id": "1185368",
      "postDate": "02/04/2021 06:12:09",
      "content": "<p>There are multiple threads on this..I just posted the below clarifications in another thread…</p>\n<p>\"Both the old and new docker images dont work with TPU. The status 14 GPRC unavailable error is typically linked to firewall/proxy settings. gRPC is google's Remote procedure call platform which tf uses to connect the master with the slaves. Seems like some pip downloads could mess it up (for e.g. grpcio - <a href=\"https://stackoverflow.com/questions/57397723/grpc-client-failing-to-connect-to-server-with-tls-certificates\" target=\"_blank\">https://stackoverflow.com/questions/57397723/grpc-client-failing-to-connect-to-server-with-tls-certificates</a> - though I noticed it isnt used in our case..but probably some other library could have changed which does not support some service used by gRPC). This isnt the first time this is happening on Kaggle. ..see for e.g. <a href=\"https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/168316\" target=\"_blank\">https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/168316</a>. Of course, Torch users are not impacted…Most of the winning solutions (in past competitions) were Torch based..I am guessing same applies here too which explains why not too many people seem bothered about this issue\"</p>\n<p>At the moment I dont see any option other than disabling certain augmentations which dont seem to work..or change your code to use GPUs. Of course you could always request help from the Kaggle support or host but my past experience in this regard has been very poor…(<a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206719)..\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206719)..</a></p>",
      "rawMarkdown": "There are multiple threads on this..I just posted the below clarifications in another thread...\n\n\"Both the old and new docker images dont work with TPU. The status 14 GPRC unavailable error is typically linked to firewall/proxy settings. gRPC is google's Remote procedure call platform which tf uses to connect the master with the slaves. Seems like some pip downloads could mess it up (for e.g. grpcio - https://stackoverflow.com/questions/57397723/grpc-client-failing-to-connect-to-server-with-tls-certificates - though I noticed it isnt used in our case..but probably some other library could have changed which does not support some service used by gRPC). This isnt the first time this is happening on Kaggle. ..see for e.g. https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/168316. Of course, Torch users are not impacted…Most of the winning solutions (in past competitions) were Torch based..I am guessing same applies here too which explains why not too many people seem bothered about this issue\"\n\nAt the moment I dont see any option other than disabling certain augmentations which dont seem to work..or change your code to use GPUs. Of course you could always request help from the Kaggle support or host but my past experience in this regard has been very poor...(https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206719)..",
      "votes": null
    },
    {
      "id": "1185378",
      "postDate": "02/04/2021 06:20:58",
      "content": "<p>Thank you . Yes I am also running in gpu but well I would like yo use the TPU time also. Maybe if we all complain to support. Or in the organizers thread..</p>",
      "rawMarkdown": "Thank you . Yes I am also running in gpu but well I would like yo use the TPU time also. Maybe if we all complain to support. Or in the organizers thread..",
      "votes": null
    },
    {
      "id": "1186640",
      "postDate": "02/05/2021 01:00:50",
      "content": "<p>I looked into this and it looks like tfa.image.cutout is not compatible with TPUs. I wonder if it has ever been or if it is a tensorflow-addons version issue.</p>",
      "rawMarkdown": "I looked into this and it looks like tfa.image.cutout is not compatible with TPUs. I wonder if it has ever been or if it is a tensorflow-addons version issue.",
      "votes": null
    },
    {
      "id": "1186697",
      "postDate": "02/05/2021 02:18:07",
      "content": "<p>I confirm this is a regression. tfa.image.cutout used to work in TF 2.2. I filed the bug internally. Unfortunately, I cannot promise a quick resolution and I have to advise all teams to avoid tfa.image.cutout on TPU. <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> posted a TPU version of CutMix <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/132935\" target=\"_blank\">here</a> and there is a batch implementation of the same by <a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a> <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/134264\" target=\"_blank\">here</a>. I'm not sure it's a direct replacement but I hope it can help.</p>",
      "rawMarkdown": "I confirm this is a regression. tfa.image.cutout used to work in TF 2.2. I filed the bug internally. Unfortunately, I cannot promise a quick resolution and I have to advise all teams to avoid tfa.image.cutout on TPU. @cdeotte posted a TPU version of CutMix [here](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/132935) and there is a batch implementation of the same by @yihdarshieh [here](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/134264). I'm not sure it's a direct replacement but I hope it can help.",
      "votes": null
    },
    {
      "id": "1186846",
      "postDate": "02/05/2021 04:45:56",
      "content": "<p>cool. Thanks for the quick workaround. Appreciate your support</p>",
      "rawMarkdown": "cool. Thanks for the quick workaround. Appreciate your support",
      "votes": null
    },
    {
      "id": "1186965",
      "postDate": "02/05/2021 06:19:29",
      "content": "<p>i have posted in another thread that this does not seem to fix the issue. I have comment out cutout from tfa and does not seem to solve the grpc error</p>",
      "rawMarkdown": "i have posted in another thread that this does not seem to fix the issue. I have comment out cutout from tfa and does not seem to solve the grpc error",
      "votes": null
    },
    {
      "id": "1187653",
      "postDate": "02/05/2021 15:55:50",
      "content": "<p><a href=\"https://www.kaggle.com/mgornergoogle\" target=\"_blank\">@mgornergoogle</a>, In fact it does not seem to be a tfaissue. For e.g. this simple 1 liner does not work:         </p>\n<p>image = tf.image.random_flip_left_right(image)</p>\n<p>You may want to recheck the same…</p>",
      "rawMarkdown": "mgornergoogle, In fact it does not seem to be a tfaissue. For e.g. this simple 1 liner does not work:         \n\nimage = tf.image.random_flip_left_right(image)\n\nYou may want to recheck the same...",
      "votes": null
    },
    {
      "id": "1187814",
      "postDate": "02/05/2021 18:22:10",
      "content": "<p>I have been able to make your notebook work on TPU. I commented the following lines:</p>\n<p>I removed lines that apply tfa.image.cutout:</p>\n<pre><code>#for i in range(ERASE_TIME_N):\n#    image = tfa.image.cutout(image, [HEIGHT, xsize[i]], offset=[HEIGHT//2, xoff[i]])\n#for i in range(ERASE_MEL_N):\n#    image = tfa.image.cutout(image, [ysize[i], WIDTH], offset=[yoff[i], WIDTH//2])\n</code></pre>\n<p>And I also had to remove the <code>@tf.function</code> annotation on <code>def _preprocess_img(...)</code>:</p>\n<pre><code>#@tf.function\ndef _preprocess_img(x, training=False):\n</code></pre>\n<p>I don't see why the <a href=\"https://www.kaggle.com/tf.function\" target=\"_blank\">@tf.function</a> decorator poses problem there. It should not. However, I can also say that it's not needed. The image preprocessing happens in the tf.data.Dataset pipeline and will therefore be executed on the CPU. The <a href=\"https://www.kaggle.com/tf.function\" target=\"_blank\">@tf.function</a> should not add any performance in that case. </p>",
      "rawMarkdown": "I have been able to make your notebook work on TPU. I commented the following lines:\n\nI removed lines that apply tfa.image.cutout:\n\n```\n#for i in range(ERASE_TIME_N):\n#    image = tfa.image.cutout(image, [HEIGHT, xsize[i]], offset=[HEIGHT//2, xoff[i]])\n#for i in range(ERASE_MEL_N):\n#    image = tfa.image.cutout(image, [ysize[i], WIDTH], offset=[yoff[i], WIDTH//2])\n```\n\nAnd I also had to remove the `@tf.function` annotation on `def _preprocess_img(...)`:\n\n```\n#@tf.function\ndef _preprocess_img(x, training=False):\n```\n\nI don't see why the @tf.function decorator poses problem there. It should not. However, I can also say that it's not needed. The image preprocessing happens in the tf.data.Dataset pipeline and will therefore be executed on the CPU. The @tf.function should not add any performance in that case.",
      "votes": null
    },
    {
      "id": "1188196",
      "postDate": "02/06/2021 03:58:27",
      "content": "<p>aaah…i guess that may work…<br>\nI will check back and tell u if there is any issue…<br>\nInteresting error though…As you have mentioned, it is a dataset pipeline so that decorator may not be needed</p>",
      "rawMarkdown": "aaah...i guess that may work...\nI will check back and tell u if there is any issue...\nInteresting error though...As you have mentioned, it is a dataset pipeline so that decorator may not be needed",
      "votes": null
    },
    {
      "id": "1188497",
      "postDate": "02/06/2021 09:43:51",
      "content": "<p>confirmed. it works without any performance impact. Thanks!<br>\nNote - I didnt find value in specaugment so I didnt test if specaugment works with the above fix (tfa.image.cutout)</p>",
      "rawMarkdown": "confirmed. it works without any performance impact. Thanks!\nNote - I didnt find value in specaugment so I didnt test if specaugment works with the above fix (tfa.image.cutout)",
      "votes": null
    },
    {
      "id": "1189237",
      "postDate": "02/06/2021 20:12:20",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/allohvk\" target=\"_blank\">@allohvk</a>  I tried just commenting the block with the plt.figure as well as the <a href=\"https://www.kaggle.com/tf.function\" target=\"_blank\">@tf.function</a> before the preprocess_img function and it does not work. Did you do anything else to get preprocessing running with specaug and cutout?</p>",
      "rawMarkdown": "Hi @allohvk  I tried just commenting the block with the plt.figure as well as the @tf.function before the preprocess_img function and it does not work. Did you do anything else to get preprocessing running with specaug and cutout?",
      "votes": null
    },
    {
      "id": "1189518",
      "postDate": "02/07/2021 04:56:46",
      "content": "<p><a href=\"https://www.kaggle.com/felipebihaiek\" target=\"_blank\">@felipebihaiek</a> no other changes done. However, like I posted, I dont use specaugment (have turned it off since it never added any value to my model). You can try commenting this line in both training and test. This should fix the issue for you:</p>\n<p>image = tf.cond(tf.random.uniform([]) &lt; 0.5, lambda: _specaugment(image), lambda: image)</p>",
      "rawMarkdown": "felipebihaiek no other changes done. However, like I posted, I dont use specaugment (have turned it off since it never added any value to my model). You can try commenting this line in both training and test. This should fix the issue for you:\n\nimage = tf.cond(tf.random.uniform([]) < 0.5, lambda: _specaugment(image), lambda: image)",
      "votes": null
    },
    {
      "id": "1194374",
      "postDate": "02/10/2021 07:05:13",
      "content": "<p><a href=\"https://www.kaggle.com/mgornergoogle\" target=\"_blank\">@mgornergoogle</a> looks like tf kernels have stopped working since morning today. Same GPRC error. This is after commenting the <a href=\"https://www.kaggle.com/tf.function\" target=\"_blank\">@tf.function</a> decorator (which had helped fix the issue earlier)</p>",
      "rawMarkdown": "mgornergoogle looks like tf kernels have stopped working since morning today. Same GPRC error. This is after commenting the @tf.function decorator (which had helped fix the issue earlier)",
      "votes": null
    },
    {
      "id": "1194742",
      "postDate": "02/10/2021 10:38:15",
      "content": "<p><a href=\"https://www.kaggle.com/mgornergoogle\" target=\"_blank\">@mgornergoogle</a> seems to be working now after I rolled back to older versions. So good for now. Pl ignore my earlier comment</p>",
      "rawMarkdown": "mgornergoogle seems to be working now after I rolled back to older versions. So good for now. Pl ignore my earlier comment",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1185313,
      "author_name": "felipebihaiek",
      "author_url": "",
      "post_date": "02/04/2021 05:38:00",
      "content": "<p>Hi I am having a similar error did you recieve any answers?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1185368,
          "author_name": "allohvk",
          "author_url": "",
          "post_date": "02/04/2021 06:12:09",
          "content": "<p>There are multiple threads on this..I just posted the below clarifications in another thread…</p>\n<p>\"Both the old and new docker images dont work with TPU. The status 14 GPRC unavailable error is typically linked to firewall/proxy settings. gRPC is google's Remote procedure call platform which tf uses to connect the master with the slaves. Seems like some pip downloads could mess it up (for e.g. grpcio - <a href=\"https://stackoverflow.com/questions/57397723/grpc-client-failing-to-connect-to-server-with-tls-certificates\" target=\"_blank\">https://stackoverflow.com/questions/57397723/grpc-client-failing-to-connect-to-server-with-tls-certificates</a> - though I noticed it isnt used in our case..but probably some other library could have changed which does not support some service used by gRPC). This isnt the first time this is happening on Kaggle. ..see for e.g. <a href=\"https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/168316\" target=\"_blank\">https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/168316</a>. Of course, Torch users are not impacted…Most of the winning solutions (in past competitions) were Torch based..I am guessing same applies here too which explains why not too many people seem bothered about this issue\"</p>\n<p>At the moment I dont see any option other than disabling certain augmentations which dont seem to work..or change your code to use GPUs. Of course you could always request help from the Kaggle support or host but my past experience in this regard has been very poor…(<a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206719)..\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206719)..</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1185378,
          "author_name": "felipebihaiek",
          "author_url": "",
          "post_date": "02/04/2021 06:20:58",
          "content": "<p>Thank you . Yes I am also running in gpu but well I would like yo use the TPU time also. Maybe if we all complain to support. Or in the organizers thread..</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1186640,
      "author_name": "mgorner",
      "author_url": "",
      "post_date": "02/05/2021 01:00:50",
      "content": "<p>I looked into this and it looks like tfa.image.cutout is not compatible with TPUs. I wonder if it has ever been or if it is a tensorflow-addons version issue.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1186697,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "02/05/2021 02:18:07",
          "content": "<p>I confirm this is a regression. tfa.image.cutout used to work in TF 2.2. I filed the bug internally. Unfortunately, I cannot promise a quick resolution and I have to advise all teams to avoid tfa.image.cutout on TPU. <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> posted a TPU version of CutMix <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/132935\" target=\"_blank\">here</a> and there is a batch implementation of the same by <a href=\"https://www.kaggle.com/yihdarshieh\" target=\"_blank\">@yihdarshieh</a> <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/134264\" target=\"_blank\">here</a>. I'm not sure it's a direct replacement but I hope it can help.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1186846,
          "author_name": "allohvk",
          "author_url": "",
          "post_date": "02/05/2021 04:45:56",
          "content": "<p>cool. Thanks for the quick workaround. Appreciate your support</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1186965,
          "author_name": "germmie",
          "author_url": "",
          "post_date": "02/05/2021 06:19:29",
          "content": "<p>i have posted in another thread that this does not seem to fix the issue. I have comment out cutout from tfa and does not seem to solve the grpc error</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1187653,
          "author_name": "allohvk",
          "author_url": "",
          "post_date": "02/05/2021 15:55:50",
          "content": "<p><a href=\"https://www.kaggle.com/mgornergoogle\" target=\"_blank\">@mgornergoogle</a>, In fact it does not seem to be a tfaissue. For e.g. this simple 1 liner does not work:         </p>\n<p>image = tf.image.random_flip_left_right(image)</p>\n<p>You may want to recheck the same…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1187814,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "02/05/2021 18:22:10",
          "content": "<p>I have been able to make your notebook work on TPU. I commented the following lines:</p>\n<p>I removed lines that apply tfa.image.cutout:</p>\n<pre><code>#for i in range(ERASE_TIME_N):\n#    image = tfa.image.cutout(image, [HEIGHT, xsize[i]], offset=[HEIGHT//2, xoff[i]])\n#for i in range(ERASE_MEL_N):\n#    image = tfa.image.cutout(image, [ysize[i], WIDTH], offset=[yoff[i], WIDTH//2])\n</code></pre>\n<p>And I also had to remove the <code>@tf.function</code> annotation on <code>def _preprocess_img(...)</code>:</p>\n<pre><code>#@tf.function\ndef _preprocess_img(x, training=False):\n</code></pre>\n<p>I don't see why the <a href=\"https://www.kaggle.com/tf.function\" target=\"_blank\">@tf.function</a> decorator poses problem there. It should not. However, I can also say that it's not needed. The image preprocessing happens in the tf.data.Dataset pipeline and will therefore be executed on the CPU. The <a href=\"https://www.kaggle.com/tf.function\" target=\"_blank\">@tf.function</a> should not add any performance in that case. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1188196,
          "author_name": "allohvk",
          "author_url": "",
          "post_date": "02/06/2021 03:58:27",
          "content": "<p>aaah…i guess that may work…<br>\nI will check back and tell u if there is any issue…<br>\nInteresting error though…As you have mentioned, it is a dataset pipeline so that decorator may not be needed</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1188497,
          "author_name": "allohvk",
          "author_url": "",
          "post_date": "02/06/2021 09:43:51",
          "content": "<p>confirmed. it works without any performance impact. Thanks!<br>\nNote - I didnt find value in specaugment so I didnt test if specaugment works with the above fix (tfa.image.cutout)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1194374,
          "author_name": "allohvk",
          "author_url": "",
          "post_date": "02/10/2021 07:05:13",
          "content": "<p><a href=\"https://www.kaggle.com/mgornergoogle\" target=\"_blank\">@mgornergoogle</a> looks like tf kernels have stopped working since morning today. Same GPRC error. This is after commenting the <a href=\"https://www.kaggle.com/tf.function\" target=\"_blank\">@tf.function</a> decorator (which had helped fix the issue earlier)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1194742,
          "author_name": "allohvk",
          "author_url": "",
          "post_date": "02/10/2021 10:38:15",
          "content": "<p><a href=\"https://www.kaggle.com/mgornergoogle\" target=\"_blank\">@mgornergoogle</a> seems to be working now after I rolled back to older versions. So good for now. Pl ignore my earlier comment</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1189237,
      "author_name": "felipebihaiek",
      "author_url": "",
      "post_date": "02/06/2021 20:12:20",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/allohvk\" target=\"_blank\">@allohvk</a>  I tried just commenting the block with the plt.figure as well as the <a href=\"https://www.kaggle.com/tf.function\" target=\"_blank\">@tf.function</a> before the preprocess_img function and it does not work. Did you do anything else to get preprocessing running with specaug and cutout?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1189518,
          "author_name": "allohvk",
          "author_url": "",
          "post_date": "02/07/2021 04:56:46",
          "content": "<p><a href=\"https://www.kaggle.com/felipebihaiek\" target=\"_blank\">@felipebihaiek</a> no other changes done. However, like I posted, I dont use specaugment (have turned it off since it never added any value to my model). You can try commenting this line in both training and test. This should fix the issue for you:</p>\n<p>image = tf.cond(tf.random.uniform([]) &lt; 0.5, lambda: _specaugment(image), lambda: image)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1182881": "<<<EDIT - This is now fixable thanks to the workaround suggested by Martin. For dataset pipelines we DONT need to explicitly specify tf graph mode with the decorator @tf.function. By default dataset pipeline operates only in graph mode. Commenting out that one line fixes the issue. I also compared the prev and current times and the execution time is NOT impacted due to the commenting. Happy TPU'ing :)\n\nI would like to thank @mgornergoogle for helping with this issue. This has been quite unlike my prev experienice in the RiiiD competition and I am pleasantly surprised with the support extended.>>>\n\nWith over 300+ direct edits and many more 2nd and 3rd level edits, I am sure the foll kernel https://www.kaggle.com/yosshi999/rfcx-train-resnet50-with-tpu by @yosshi999 would be the most popular one in this competition\n\nHere is an interesting observation. The below block which was working earlier fails with this \"failed to connect to all addresses...\" error (you can just do a edit and run to simulate this):\n\n```\nplt.figure(figsize=(16, 4))\nfor i, (inp, targ) in enumerate(annot_dataset.map(_preprocess).take(6)):\n    plt.subplot(2,3,i+1)\n    plt.imshow(inp.numpy()[:,:,0])\n    t = targ.numpy()\n    if t.sum() == 0:\n        plt.title(f'FP')\n    else:\n        plt.title(f'{t.nonzero()[0]}')\n    plt.colorbar()\nplt.show() \n```\n\nUnfortunately not too much info exists about the error on SFO. I could get it working by commenting the specaugment and the gaussian noise preprocessing in the block preceeding it. \n\nAnother strange behaviour is that a simple workaround is to comment out the whole block..It is anyway a exploratory block and does not impact final processing. So subsequent dataset mappings work nicely with all the preprocessing (including specaug and gaussian noise)\n\nBut interested to know why this error happens. Was some software upgraded recently @ kaggle env which is causing this. We can explore some interesting combinations also. For e..g. changing the brightness works but a simple randomfilp fails.\n\nIf anybody has any pointers, do let me know\n\nBelow is the full error:\nUnavailableError: failed to connect to all addresses\nAdditional GRPC error information from remote target /job:localhost/replica:0/task:0/device:CPU:0:\n:{\"created\":\"@1612282112.116200424\",\"description\":\"Failed to pick subchannel\",\"file\":\"third_party/grpc/src/core/ext/filters/client_channel/client_channel.cc\",\"file_line\":4143,\"referenced_errors\":[{\"created\":\"@1612282112.116196618\",\"description\":\"failed to connect to all addresses\",\"file\":\"third_party/grpc/src/core/ext/filters/client_channel/lb_policy/pick_first/pick_first.cc\",\"file_line\":398,\"grpc_status\":14}]}\n\t [[{{node StatefulPartitionedCall}}]]",
    "1185313": "Hi I am having a similar error did you recieve any answers?",
    "1185368": "There are multiple threads on this..I just posted the below clarifications in another thread...\n\n\"Both the old and new docker images dont work with TPU. The status 14 GPRC unavailable error is typically linked to firewall/proxy settings. gRPC is google's Remote procedure call platform which tf uses to connect the master with the slaves. Seems like some pip downloads could mess it up (for e.g. grpcio - https://stackoverflow.com/questions/57397723/grpc-client-failing-to-connect-to-server-with-tls-certificates - though I noticed it isnt used in our case..but probably some other library could have changed which does not support some service used by gRPC). This isnt the first time this is happening on Kaggle. ..see for e.g. https://www.kaggle.com/c/alaska2-image-steganalysis/discussion/168316. Of course, Torch users are not impacted…Most of the winning solutions (in past competitions) were Torch based..I am guessing same applies here too which explains why not too many people seem bothered about this issue\"\n\nAt the moment I dont see any option other than disabling certain augmentations which dont seem to work..or change your code to use GPUs. Of course you could always request help from the Kaggle support or host but my past experience in this regard has been very poor...(https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206719)..",
    "1185378": "Thank you . Yes I am also running in gpu but well I would like yo use the TPU time also. Maybe if we all complain to support. Or in the organizers thread..",
    "1186640": "I looked into this and it looks like tfa.image.cutout is not compatible with TPUs. I wonder if it has ever been or if it is a tensorflow-addons version issue.",
    "1186697": "I confirm this is a regression. tfa.image.cutout used to work in TF 2.2. I filed the bug internally. Unfortunately, I cannot promise a quick resolution and I have to advise all teams to avoid tfa.image.cutout on TPU. @cdeotte posted a TPU version of CutMix [here](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/132935) and there is a batch implementation of the same by @yihdarshieh [here](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/134264). I'm not sure it's a direct replacement but I hope it can help.",
    "1186846": "cool. Thanks for the quick workaround. Appreciate your support",
    "1186965": "i have posted in another thread that this does not seem to fix the issue. I have comment out cutout from tfa and does not seem to solve the grpc error",
    "1187653": "mgornergoogle, In fact it does not seem to be a tfaissue. For e.g. this simple 1 liner does not work:         \n\nimage = tf.image.random_flip_left_right(image)\n\nYou may want to recheck the same...",
    "1187814": "I have been able to make your notebook work on TPU. I commented the following lines:\n\nI removed lines that apply tfa.image.cutout:\n\n```\n#for i in range(ERASE_TIME_N):\n#    image = tfa.image.cutout(image, [HEIGHT, xsize[i]], offset=[HEIGHT//2, xoff[i]])\n#for i in range(ERASE_MEL_N):\n#    image = tfa.image.cutout(image, [ysize[i], WIDTH], offset=[yoff[i], WIDTH//2])\n```\n\nAnd I also had to remove the `@tf.function` annotation on `def _preprocess_img(...)`:\n\n```\n#@tf.function\ndef _preprocess_img(x, training=False):\n```\n\nI don't see why the @tf.function decorator poses problem there. It should not. However, I can also say that it's not needed. The image preprocessing happens in the tf.data.Dataset pipeline and will therefore be executed on the CPU. The @tf.function should not add any performance in that case.",
    "1188196": "aaah...i guess that may work...\nI will check back and tell u if there is any issue...\nInteresting error though...As you have mentioned, it is a dataset pipeline so that decorator may not be needed",
    "1188497": "confirmed. it works without any performance impact. Thanks!\nNote - I didnt find value in specaugment so I didnt test if specaugment works with the above fix (tfa.image.cutout)",
    "1189237": "Hi @allohvk  I tried just commenting the block with the plt.figure as well as the @tf.function before the preprocess_img function and it does not work. Did you do anything else to get preprocessing running with specaug and cutout?",
    "1189518": "felipebihaiek no other changes done. However, like I posted, I dont use specaugment (have turned it off since it never added any value to my model). You can try commenting this line in both training and test. This should fix the issue for you:\n\nimage = tf.cond(tf.random.uniform([]) < 0.5, lambda: _specaugment(image), lambda: image)",
    "1194374": "mgornergoogle looks like tf kernels have stopped working since morning today. Same GPRC error. This is after commenting the @tf.function decorator (which had helped fix the issue earlier)",
    "1194742": "mgornergoogle seems to be working now after I rolled back to older versions. So good for now. Pl ignore my earlier comment"
  },
  "source": "meta"
}