{
  "id": 394207,
  "title": "Why is the GRU network not supported when I use Torch to convert to the TF2 model? How can I bypass this error?",
  "url": "/competitions/asl-signs/discussion/394207",
  "author_name": "",
  "post_date": "2023-03-12T15:33:44.673288500Z",
  "votes": 2,
  "comment_count": 5,
  "views": 0,
  "content": "<p>`RuntimeError: in user code:</p>\n<pre><code>File \"/opt/conda/lib/python3.7/site-packages/onnx_tf/backend_tf_module.py\", line 98, in __call__  *\n    output_ops = self.backend._onnx_node_to_tensorflow_op(onnx_node,\nFile \"/opt/conda/lib/python3.7/site-packages/onnx_tf/backend.py\", line 289, in _onnx_node_to_tensorflow_op  *\n    return handler.handle(node, tensor_dict=tensor_dict, strict=strict)\nFile \"/opt/conda/lib/python3.7/site-packages/onnx_tf/handlers/handler.py\", line 58, in handle  *\n    cls.args_check(node, **kwargs)\nFile \"/opt/conda/lib/python3.7/site-packages/onnx_tf/handlers/backend/gru.py\", line 68, in args_check  *\n    exception.OP_UNSUPPORTED_EXCEPT(\"GRU with linear_before_reset\",\nFile \"/opt/conda/lib/python3.7/site-packages/onnx_tf/common/exception.py\", line 50, in __call__  *\n    raise self._func(self.get_message(op, framework))\n\nRuntimeError: GRU with linear_before_reset is not supported in Tensorflow.`\n</code></pre>",
  "messages": [
    {
      "id": "2178648",
      "postDate": "03/12/2023 15:33:44",
      "content": "<p>`RuntimeError: in user code:</p>\n<pre><code>File \"/opt/conda/lib/python3.7/site-packages/onnx_tf/backend_tf_module.py\", line 98, in __call__  *\n    output_ops = self.backend._onnx_node_to_tensorflow_op(onnx_node,\nFile \"/opt/conda/lib/python3.7/site-packages/onnx_tf/backend.py\", line 289, in _onnx_node_to_tensorflow_op  *\n    return handler.handle(node, tensor_dict=tensor_dict, strict=strict)\nFile \"/opt/conda/lib/python3.7/site-packages/onnx_tf/handlers/handler.py\", line 58, in handle  *\n    cls.args_check(node, **kwargs)\nFile \"/opt/conda/lib/python3.7/site-packages/onnx_tf/handlers/backend/gru.py\", line 68, in args_check  *\n    exception.OP_UNSUPPORTED_EXCEPT(\"GRU with linear_before_reset\",\nFile \"/opt/conda/lib/python3.7/site-packages/onnx_tf/common/exception.py\", line 50, in __call__  *\n    raise self._func(self.get_message(op, framework))\n\nRuntimeError: GRU with linear_before_reset is not supported in Tensorflow.`\n</code></pre>",
      "rawMarkdown": "`RuntimeError: in user code:\n\n    File \"/opt/conda/lib/python3.7/site-packages/onnx_tf/backend_tf_module.py\", line 98, in __call__  *\n        output_ops = self.backend._onnx_node_to_tensorflow_op(onnx_node,\n    File \"/opt/conda/lib/python3.7/site-packages/onnx_tf/backend.py\", line 289, in _onnx_node_to_tensorflow_op  *\n        return handler.handle(node, tensor_dict=tensor_dict, strict=strict)\n    File \"/opt/conda/lib/python3.7/site-packages/onnx_tf/handlers/handler.py\", line 58, in handle  *\n        cls.args_check(node, **kwargs)\n    File \"/opt/conda/lib/python3.7/site-packages/onnx_tf/handlers/backend/gru.py\", line 68, in args_check  *\n        exception.OP_UNSUPPORTED_EXCEPT(\"GRU with linear_before_reset\",\n    File \"/opt/conda/lib/python3.7/site-packages/onnx_tf/common/exception.py\", line 50, in __call__  *\n        raise self._func(self.get_message(op, framework))\n\n    RuntimeError: GRU with linear_before_reset is not supported in Tensorflow.`",
      "votes": null
    },
    {
      "id": "2180402",
      "postDate": "03/13/2023 19:21:09",
      "content": "<p>I ran into this same issue and unfortunately I don't think it is bypassable (or at least I couldn't spending a decent amount of time on it). The Tensorflow GRU that is used in the conversion process literally has a different sequence of operations from the Pytorch definition, which is the reason for the issue. There is a relevant pull request that unfortunately never got merged here on the onnx-tensorflow repo: <a href=\"https://github.com/onnx/onnx-tensorflow/pull/452\" target=\"_blank\">https://github.com/onnx/onnx-tensorflow/pull/452</a></p>",
      "rawMarkdown": "I ran into this same issue and unfortunately I don't think it is bypassable (or at least I couldn't spending a decent amount of time on it). The Tensorflow GRU that is used in the conversion process literally has a different sequence of operations from the Pytorch definition, which is the reason for the issue. There is a relevant pull request that unfortunately never got merged here on the onnx-tensorflow repo: https://github.com/onnx/onnx-tensorflow/pull/452",
      "votes": null
    },
    {
      "id": "2180751",
      "postDate": "03/14/2023 04:29:56",
      "content": "<p>this is my suggestion and is waht i am doing.</p>\n<ol>\n<li>the network in your solution is going to be shallow. it is not going to be complex becuase:</li>\n</ol>\n<ul>\n<li>you have computation complexity limit within 100ms</li>\n<li>the datasei size is small. it is easily to overfit</li>\n</ul>\n<ol>\n<li><p>it s easy to write GRU, attention, LSTM, even graph GCN using basic ops like linear, matmul, etc. <br>\nhence i understand keras layer implementation equations from keras doc. I verify by rewriting keras layers in basic keras ops to confirm results are the same.</p></li>\n<li><p>i build keras model using keras layer. i build an equivalent pytorch model using pytorch basic ops. (it can be using pytorch layers and rewriting the forward functions if pytorch equation and keras are different) </p></li>\n<li><p>i train in pytorch.</p></li>\n<li><p>i manually copied the weights from pytorch to keras after training.</p></li>\n<li><p>tflite is converted from keras.  </p></li>\n</ol>\n<hr>\n<p>this is a good exercise to refresh your understanding and knowledge of fundament equations of the basic layers in modeling today.<br>\nbecuase you know the equation well, it is easy to optmize model if required.</p>",
      "rawMarkdown": "this is my suggestion and is waht i am doing.\n\n1. the network in your solution is going to be shallow. it is not going to be complex becuase:\n- you have computation complexity limit within 100ms\n- the datasei size is small. it is easily to overfit\n\n2. it s easy to write GRU, attention, LSTM, even graph GCN using basic ops like linear, matmul, etc. \nhence i understand keras layer implementation equations from keras doc. I verify by rewriting keras layers in basic keras ops to confirm results are the same.\n\n3. i build keras model using keras layer. i build an equivalent pytorch model using pytorch basic ops. (it can be using pytorch layers and rewriting the forward functions if pytorch equation and keras are different) \n\n4. i train in pytorch.\n\n5. i manually copied the weights from pytorch to keras after training.\n\n6. tflite is converted from keras.  \n\n---\n\nthis is a good exercise to refresh your understanding and knowledge of fundament equations of the basic layers in modeling today.\nbecuase you know the equation well, it is easy to optmize model if required.",
      "votes": null
    },
    {
      "id": "2180754",
      "postDate": "03/14/2023 04:32:54",
      "content": "<p>if you google around there are sucessful code and example to rewrite keras GRU using pytorch operations. i<br>\nthe issue is really if we break up pytorch GRU, whether onnx call graph will be optimzed. if onnx call graph is not optimized correctly, then tf/tflite call graph will be a mess.</p>",
      "rawMarkdown": "if you google around there are sucessful code and example to rewrite keras GRU using pytorch operations. i\nthe issue is really if we break up pytorch GRU, whether onnx call graph will be optimzed. if onnx call graph is not optimized correctly, then tf/tflite call graph will be a mess.",
      "votes": null
    },
    {
      "id": "2181447",
      "postDate": "03/14/2023 14:23:57",
      "content": "<p>thx! It certainly is very good exercise to understanding the knowledge!<br>\nin the end, I will try to optimize it with your model, hope it can be improved! Thank you very much for your contribution!👍👍🙏</p>",
      "rawMarkdown": "thx! It certainly is very good exercise to understanding the knowledge!\nin the end, I will try to optimize it with your model, hope it can be improved! Thank you very much for your contribution!👍👍🙏",
      "votes": null
    },
    {
      "id": "2184100",
      "postDate": "03/16/2023 06:23:19",
      "content": "<p><a href=\"https://www.kaggle.com/lau01b\" target=\"_blank\">@lau01b</a> </p>\n<p><a href=\"https://www.kaggle.com/code/hengck23/example-from-pytorch-to-keras-gru-by-hand\" target=\"_blank\">https://www.kaggle.com/code/hengck23/example-from-pytorch-to-keras-gru-by-hand</a></p>\n<p>this is still in development. you can check regularly for update</p>",
      "rawMarkdown": "lau01b \n\nhttps://www.kaggle.com/code/hengck23/example-from-pytorch-to-keras-gru-by-hand\n\nthis is still in development. you can check regularly for update",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2180402,
      "author_name": "mitchelldehaven",
      "author_url": "",
      "post_date": "03/13/2023 19:21:09",
      "content": "<p>I ran into this same issue and unfortunately I don't think it is bypassable (or at least I couldn't spending a decent amount of time on it). The Tensorflow GRU that is used in the conversion process literally has a different sequence of operations from the Pytorch definition, which is the reason for the issue. There is a relevant pull request that unfortunately never got merged here on the onnx-tensorflow repo: <a href=\"https://github.com/onnx/onnx-tensorflow/pull/452\" target=\"_blank\">https://github.com/onnx/onnx-tensorflow/pull/452</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 2180754,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "03/14/2023 04:32:54",
          "content": "<p>if you google around there are sucessful code and example to rewrite keras GRU using pytorch operations. i<br>\nthe issue is really if we break up pytorch GRU, whether onnx call graph will be optimzed. if onnx call graph is not optimized correctly, then tf/tflite call graph will be a mess.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2180751,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/14/2023 04:29:56",
      "content": "<p>this is my suggestion and is waht i am doing.</p>\n<ol>\n<li>the network in your solution is going to be shallow. it is not going to be complex becuase:</li>\n</ol>\n<ul>\n<li>you have computation complexity limit within 100ms</li>\n<li>the datasei size is small. it is easily to overfit</li>\n</ul>\n<ol>\n<li><p>it s easy to write GRU, attention, LSTM, even graph GCN using basic ops like linear, matmul, etc. <br>\nhence i understand keras layer implementation equations from keras doc. I verify by rewriting keras layers in basic keras ops to confirm results are the same.</p></li>\n<li><p>i build keras model using keras layer. i build an equivalent pytorch model using pytorch basic ops. (it can be using pytorch layers and rewriting the forward functions if pytorch equation and keras are different) </p></li>\n<li><p>i train in pytorch.</p></li>\n<li><p>i manually copied the weights from pytorch to keras after training.</p></li>\n<li><p>tflite is converted from keras.  </p></li>\n</ol>\n<hr>\n<p>this is a good exercise to refresh your understanding and knowledge of fundament equations of the basic layers in modeling today.<br>\nbecuase you know the equation well, it is easy to optmize model if required.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2181447,
          "author_name": "lau01b",
          "author_url": "",
          "post_date": "03/14/2023 14:23:57",
          "content": "<p>thx! It certainly is very good exercise to understanding the knowledge!<br>\nin the end, I will try to optimize it with your model, hope it can be improved! Thank you very much for your contribution!👍👍🙏</p>",
          "votes": null,
          "replies": [
            {
              "id": 2184100,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "03/16/2023 06:23:19",
              "content": "<p><a href=\"https://www.kaggle.com/lau01b\" target=\"_blank\">@lau01b</a> </p>\n<p><a href=\"https://www.kaggle.com/code/hengck23/example-from-pytorch-to-keras-gru-by-hand\" target=\"_blank\">https://www.kaggle.com/code/hengck23/example-from-pytorch-to-keras-gru-by-hand</a></p>\n<p>this is still in development. you can check regularly for update</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2178648": "`RuntimeError: in user code:\n\n    File \"/opt/conda/lib/python3.7/site-packages/onnx_tf/backend_tf_module.py\", line 98, in __call__  *\n        output_ops = self.backend._onnx_node_to_tensorflow_op(onnx_node,\n    File \"/opt/conda/lib/python3.7/site-packages/onnx_tf/backend.py\", line 289, in _onnx_node_to_tensorflow_op  *\n        return handler.handle(node, tensor_dict=tensor_dict, strict=strict)\n    File \"/opt/conda/lib/python3.7/site-packages/onnx_tf/handlers/handler.py\", line 58, in handle  *\n        cls.args_check(node, **kwargs)\n    File \"/opt/conda/lib/python3.7/site-packages/onnx_tf/handlers/backend/gru.py\", line 68, in args_check  *\n        exception.OP_UNSUPPORTED_EXCEPT(\"GRU with linear_before_reset\",\n    File \"/opt/conda/lib/python3.7/site-packages/onnx_tf/common/exception.py\", line 50, in __call__  *\n        raise self._func(self.get_message(op, framework))\n\n    RuntimeError: GRU with linear_before_reset is not supported in Tensorflow.`",
    "2180402": "I ran into this same issue and unfortunately I don't think it is bypassable (or at least I couldn't spending a decent amount of time on it). The Tensorflow GRU that is used in the conversion process literally has a different sequence of operations from the Pytorch definition, which is the reason for the issue. There is a relevant pull request that unfortunately never got merged here on the onnx-tensorflow repo: https://github.com/onnx/onnx-tensorflow/pull/452",
    "2180751": "this is my suggestion and is waht i am doing.\n\n1. the network in your solution is going to be shallow. it is not going to be complex becuase:\n- you have computation complexity limit within 100ms\n- the datasei size is small. it is easily to overfit\n\n2. it s easy to write GRU, attention, LSTM, even graph GCN using basic ops like linear, matmul, etc. \nhence i understand keras layer implementation equations from keras doc. I verify by rewriting keras layers in basic keras ops to confirm results are the same.\n\n3. i build keras model using keras layer. i build an equivalent pytorch model using pytorch basic ops. (it can be using pytorch layers and rewriting the forward functions if pytorch equation and keras are different) \n\n4. i train in pytorch.\n\n5. i manually copied the weights from pytorch to keras after training.\n\n6. tflite is converted from keras.  \n\n---\n\nthis is a good exercise to refresh your understanding and knowledge of fundament equations of the basic layers in modeling today.\nbecuase you know the equation well, it is easy to optmize model if required.",
    "2180754": "if you google around there are sucessful code and example to rewrite keras GRU using pytorch operations. i\nthe issue is really if we break up pytorch GRU, whether onnx call graph will be optimzed. if onnx call graph is not optimized correctly, then tf/tflite call graph will be a mess.",
    "2181447": "thx! It certainly is very good exercise to understanding the knowledge!\nin the end, I will try to optimize it with your model, hope it can be improved! Thank you very much for your contribution!👍👍🙏",
    "2184100": "lau01b \n\nhttps://www.kaggle.com/code/hengck23/example-from-pytorch-to-keras-gru-by-hand\n\nthis is still in development. you can check regularly for update"
  },
  "source": "meta"
}