{
  "id": 175129,
  "title": "Some reverse engineering attempts",
  "url": "/competitions/landmark-retrieval-2020/discussion/175129",
  "author_name": "",
  "post_date": "2020-08-17T08:48:57.826568700Z",
  "votes": 10,
  "comment_count": 2,
  "views": 0,
  "content": "<p>When the organizing team released their DELG kernel for the recognition challenge, I was curious about its performance in this challenge, but the signatures for the released model (which allowed multiple scales and also gives local features and boxes) is compatible for the scoring script. So I've decided to do a little bit of reverse engineering to recover the original model from the tf SavedModel format.</p>\n<p>Here's a brief summary of what I did:</p>\n<ol>\n<li>Using tf-onnx to convert the tensorflow SavedModel to onnx</li>\n<li>Using onnx-runtime to verify correctness</li>\n<li>Inspect network structure using netron</li>\n<li>Essentially building a onnx-to-keras converter from scratch to convert onnx model to tf.keras</li>\n<li>During conversion, record specific nodes that are not needed (Loop, NMS, and basically the entire local branch) and cut them off from the original graph</li>\n<li>Implement multi-scale inference in another Keras model wrapper</li>\n<li>Re-do 2-6 to generate a nicer-looking graph (because step 5 gives out really ugly results)</li>\n</ol>\n<p>I've verified single scale Keras output with the released version, and the average absolute difference is 1e-9. Due to differences in resize implementation I couldn't replicate multi-scale outputs, but I don't think that makes a lot of difference.</p>\n<p>And I managed to submit both DELG models with global features to the competition, with results ranging from 0.22 to 0.24. And that should align with the official results reported in the GLD-v2 paper, where the retrained ResNet101 should outperform DELG global features.</p>\n<p>To sum up, I think we should make it a community initiative to maintain a stable onnx to keras convertor, so that more models trained in PyTorch or simply came from SavedModel format can be more easily adapted and deployed. Let me know what you think of my approach, and what is your result in trying to submit DELG models to this competition!</p>",
  "messages": [
    {
      "id": "973330",
      "postDate": "08/17/2020 08:48:57",
      "content": "<p>When the organizing team released their DELG kernel for the recognition challenge, I was curious about its performance in this challenge, but the signatures for the released model (which allowed multiple scales and also gives local features and boxes) is compatible for the scoring script. So I've decided to do a little bit of reverse engineering to recover the original model from the tf SavedModel format.</p>\n<p>Here's a brief summary of what I did:</p>\n<ol>\n<li>Using tf-onnx to convert the tensorflow SavedModel to onnx</li>\n<li>Using onnx-runtime to verify correctness</li>\n<li>Inspect network structure using netron</li>\n<li>Essentially building a onnx-to-keras converter from scratch to convert onnx model to tf.keras</li>\n<li>During conversion, record specific nodes that are not needed (Loop, NMS, and basically the entire local branch) and cut them off from the original graph</li>\n<li>Implement multi-scale inference in another Keras model wrapper</li>\n<li>Re-do 2-6 to generate a nicer-looking graph (because step 5 gives out really ugly results)</li>\n</ol>\n<p>I've verified single scale Keras output with the released version, and the average absolute difference is 1e-9. Due to differences in resize implementation I couldn't replicate multi-scale outputs, but I don't think that makes a lot of difference.</p>\n<p>And I managed to submit both DELG models with global features to the competition, with results ranging from 0.22 to 0.24. And that should align with the official results reported in the GLD-v2 paper, where the retrained ResNet101 should outperform DELG global features.</p>\n<p>To sum up, I think we should make it a community initiative to maintain a stable onnx to keras convertor, so that more models trained in PyTorch or simply came from SavedModel format can be more easily adapted and deployed. Let me know what you think of my approach, and what is your result in trying to submit DELG models to this competition!</p>",
      "rawMarkdown": "When the organizing team released their DELG kernel for the recognition challenge, I was curious about its performance in this challenge, but the signatures for the released model (which allowed multiple scales and also gives local features and boxes) is compatible for the scoring script. So I've decided to do a little bit of reverse engineering to recover the original model from the tf SavedModel format.\n\nHere's a brief summary of what I did:\n\n1. Using tf-onnx to convert the tensorflow SavedModel to onnx\n2. Using onnx-runtime to verify correctness\n3. Inspect network structure using netron\n4. Essentially building a onnx-to-keras converter from scratch to convert onnx model to tf.keras\n5. During conversion, record specific nodes that are not needed (Loop, NMS, and basically the entire local branch) and cut them off from the original graph\n6. Implement multi-scale inference in another Keras model wrapper\n7. Re-do 2-6 to generate a nicer-looking graph (because step 5 gives out really ugly results)\n\nI've verified single scale Keras output with the released version, and the average absolute difference is 1e-9. Due to differences in resize implementation I couldn't replicate multi-scale outputs, but I don't think that makes a lot of difference.\n\nAnd I managed to submit both DELG models with global features to the competition, with results ranging from 0.22 to 0.24. And that should align with the official results reported in the GLD-v2 paper, where the retrained ResNet101 should outperform DELG global features.\n\nTo sum up, I think we should make it a community initiative to maintain a stable onnx to keras convertor, so that more models trained in PyTorch or simply came from SavedModel format can be more easily adapted and deployed. Let me know what you think of my approach, and what is your result in trying to submit DELG models to this competition!",
      "votes": null
    },
    {
      "id": "973774",
      "postDate": "08/17/2020 14:05:28",
      "content": "<p>I think it would be really hard to maintain community converter without support of Google. And, it seems, due to reasons, Google doesn't want to support ONNX data format [ <a href=\"https://github.com/keras-team/keras/issues/8638\" target=\"_blank\">https://github.com/keras-team/keras/issues/8638</a> ].</p>\n<p>Firstly, Tensorflow's internals are hairy and change very fast. I mean, you have the concrete example in the form of published baseline model - it was trained using TFv1 API and it's really difficult to work with it in TFv2 [ <a href=\"https://www.kaggle.com/c/landmark-retrieval-2020/discussion/166149\" target=\"_blank\">https://www.kaggle.com/c/landmark-retrieval-2020/discussion/166149</a> ]. It's interesting to notice that, even though DELG paper was published Jan 2020, the training code is still not yet released - I can only suspect that this is somehow connected to rewriting everything into purely TFv2 API, which seems to be a hard task even for Google itself. DELG would not be the only such project - the Object Detection API [ <a href=\"https://github.com/tensorflow/models/blob/master/research/object_detection/README.md\" target=\"_blank\">https://github.com/tensorflow/models/blob/master/research/object_detection/README.md</a> ] was also unsupported in TFv2 for a long time…</p>\n<p>Moreover, I don't think ONNX is a stable data format as well. From my experience with working with it using NVIDIA's TensorRT and PyTorch - you often run into strange problems when trying to export/import something slightly more complex than usual CNN's. To be frank though, last time I worked with that format was about a year ago, so maybe the situation is different now. </p>\n<p>All of those reasons lead me to thinking that it would be very difficult to make that converter stable and something more than a hacky project. But maybe I'm just a pessimist ;)</p>",
      "rawMarkdown": "I think it would be really hard to maintain community converter without support of Google. And, it seems, due to reasons, Google doesn't want to support ONNX data format [ https://github.com/keras-team/keras/issues/8638 ].\n\nFirstly, Tensorflow's internals are hairy and change very fast. I mean, you have the concrete example in the form of published baseline model - it was trained using TFv1 API and it's really difficult to work with it in TFv2 [ https://www.kaggle.com/c/landmark-retrieval-2020/discussion/166149 ]. It's interesting to notice that, even though DELG paper was published Jan 2020, the training code is still not yet released - I can only suspect that this is somehow connected to rewriting everything into purely TFv2 API, which seems to be a hard task even for Google itself. DELG would not be the only such project - the Object Detection API [ https://github.com/tensorflow/models/blob/master/research/object_detection/README.md ] was also unsupported in TFv2 for a long time...\n\nMoreover, I don't think ONNX is a stable data format as well. From my experience with working with it using NVIDIA's TensorRT and PyTorch - you often run into strange problems when trying to export/import something slightly more complex than usual CNN's. To be frank though, last time I worked with that format was about a year ago, so maybe the situation is different now. \n\nAll of those reasons lead me to thinking that it would be very difficult to make that converter stable and something more than a hacky project. But maybe I'm just a pessimist ;)",
      "votes": null
    },
    {
      "id": "974337",
      "postDate": "08/17/2020 22:49:58",
      "content": "<p>There is an easier conversion method if you are not trying to re-train or fine-tune the model:</p>\n<ol>\n<li>Load the saved model using <code>tf.saved_model.load</code>, and get the inference function <code>infer</code> under signature <code>serving_default</code></li>\n<li>Freeze the graph of that function using this tutorial: <a href=\"https://leimao.github.io/blog/Save-Load-Inference-From-TF2-Frozen-Graph/\" target=\"_blank\">https://leimao.github.io/blog/Save-Load-Inference-From-TF2-Frozen-Graph/</a>. This should get rid of any <code>tf.Variable</code> in the graph (this tutorial is for <code>tf.function</code>, they extract <code>ConcreteFunction</code> from it; and guess what? <code>infer</code> is a <code>WrappedFunction</code>, a class inherited from <code>ConcreteFunction</code>.</li>\n<li>Write your own custom <code>tf.Module</code> wrapper as in the official TF 2.3 documentation <a href=\"https://www.tensorflow.org/guide/saved_model#saving_a_custom_model\" target=\"_blank\">https://www.tensorflow.org/guide/saved_model#saving_a_custom_model</a></li>\n<li>Save it as SavedModel, as in the tutorial. That's all.</li>\n</ol>",
      "rawMarkdown": "There is an easier conversion method if you are not trying to re-train or fine-tune the model:\n\n1. Load the saved model using `tf.saved_model.load`, and get the inference function `infer` under signature `serving_default`\n2. Freeze the graph of that function using this tutorial: https://leimao.github.io/blog/Save-Load-Inference-From-TF2-Frozen-Graph/. This should get rid of any `tf.Variable` in the graph (this tutorial is for `tf.function`, they extract `ConcreteFunction` from it; and guess what? `infer` is a `WrappedFunction`, a class inherited from `ConcreteFunction`.\n3. Write your own custom `tf.Module` wrapper as in the official TF 2.3 documentation https://www.tensorflow.org/guide/saved_model#saving_a_custom_model\n4. Save it as SavedModel, as in the tutorial. That's all.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 973774,
      "author_name": "qiubit",
      "author_url": "",
      "post_date": "08/17/2020 14:05:28",
      "content": "<p>I think it would be really hard to maintain community converter without support of Google. And, it seems, due to reasons, Google doesn't want to support ONNX data format [ <a href=\"https://github.com/keras-team/keras/issues/8638\" target=\"_blank\">https://github.com/keras-team/keras/issues/8638</a> ].</p>\n<p>Firstly, Tensorflow's internals are hairy and change very fast. I mean, you have the concrete example in the form of published baseline model - it was trained using TFv1 API and it's really difficult to work with it in TFv2 [ <a href=\"https://www.kaggle.com/c/landmark-retrieval-2020/discussion/166149\" target=\"_blank\">https://www.kaggle.com/c/landmark-retrieval-2020/discussion/166149</a> ]. It's interesting to notice that, even though DELG paper was published Jan 2020, the training code is still not yet released - I can only suspect that this is somehow connected to rewriting everything into purely TFv2 API, which seems to be a hard task even for Google itself. DELG would not be the only such project - the Object Detection API [ <a href=\"https://github.com/tensorflow/models/blob/master/research/object_detection/README.md\" target=\"_blank\">https://github.com/tensorflow/models/blob/master/research/object_detection/README.md</a> ] was also unsupported in TFv2 for a long time…</p>\n<p>Moreover, I don't think ONNX is a stable data format as well. From my experience with working with it using NVIDIA's TensorRT and PyTorch - you often run into strange problems when trying to export/import something slightly more complex than usual CNN's. To be frank though, last time I worked with that format was about a year ago, so maybe the situation is different now. </p>\n<p>All of those reasons lead me to thinking that it would be very difficult to make that converter stable and something more than a hacky project. But maybe I'm just a pessimist ;)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 974337,
      "author_name": "chankhavu",
      "author_url": "",
      "post_date": "08/17/2020 22:49:58",
      "content": "<p>There is an easier conversion method if you are not trying to re-train or fine-tune the model:</p>\n<ol>\n<li>Load the saved model using <code>tf.saved_model.load</code>, and get the inference function <code>infer</code> under signature <code>serving_default</code></li>\n<li>Freeze the graph of that function using this tutorial: <a href=\"https://leimao.github.io/blog/Save-Load-Inference-From-TF2-Frozen-Graph/\" target=\"_blank\">https://leimao.github.io/blog/Save-Load-Inference-From-TF2-Frozen-Graph/</a>. This should get rid of any <code>tf.Variable</code> in the graph (this tutorial is for <code>tf.function</code>, they extract <code>ConcreteFunction</code> from it; and guess what? <code>infer</code> is a <code>WrappedFunction</code>, a class inherited from <code>ConcreteFunction</code>.</li>\n<li>Write your own custom <code>tf.Module</code> wrapper as in the official TF 2.3 documentation <a href=\"https://www.tensorflow.org/guide/saved_model#saving_a_custom_model\" target=\"_blank\">https://www.tensorflow.org/guide/saved_model#saving_a_custom_model</a></li>\n<li>Save it as SavedModel, as in the tutorial. That's all.</li>\n</ol>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "973330": "When the organizing team released their DELG kernel for the recognition challenge, I was curious about its performance in this challenge, but the signatures for the released model (which allowed multiple scales and also gives local features and boxes) is compatible for the scoring script. So I've decided to do a little bit of reverse engineering to recover the original model from the tf SavedModel format.\n\nHere's a brief summary of what I did:\n\n1. Using tf-onnx to convert the tensorflow SavedModel to onnx\n2. Using onnx-runtime to verify correctness\n3. Inspect network structure using netron\n4. Essentially building a onnx-to-keras converter from scratch to convert onnx model to tf.keras\n5. During conversion, record specific nodes that are not needed (Loop, NMS, and basically the entire local branch) and cut them off from the original graph\n6. Implement multi-scale inference in another Keras model wrapper\n7. Re-do 2-6 to generate a nicer-looking graph (because step 5 gives out really ugly results)\n\nI've verified single scale Keras output with the released version, and the average absolute difference is 1e-9. Due to differences in resize implementation I couldn't replicate multi-scale outputs, but I don't think that makes a lot of difference.\n\nAnd I managed to submit both DELG models with global features to the competition, with results ranging from 0.22 to 0.24. And that should align with the official results reported in the GLD-v2 paper, where the retrained ResNet101 should outperform DELG global features.\n\nTo sum up, I think we should make it a community initiative to maintain a stable onnx to keras convertor, so that more models trained in PyTorch or simply came from SavedModel format can be more easily adapted and deployed. Let me know what you think of my approach, and what is your result in trying to submit DELG models to this competition!",
    "973774": "I think it would be really hard to maintain community converter without support of Google. And, it seems, due to reasons, Google doesn't want to support ONNX data format [ https://github.com/keras-team/keras/issues/8638 ].\n\nFirstly, Tensorflow's internals are hairy and change very fast. I mean, you have the concrete example in the form of published baseline model - it was trained using TFv1 API and it's really difficult to work with it in TFv2 [ https://www.kaggle.com/c/landmark-retrieval-2020/discussion/166149 ]. It's interesting to notice that, even though DELG paper was published Jan 2020, the training code is still not yet released - I can only suspect that this is somehow connected to rewriting everything into purely TFv2 API, which seems to be a hard task even for Google itself. DELG would not be the only such project - the Object Detection API [ https://github.com/tensorflow/models/blob/master/research/object_detection/README.md ] was also unsupported in TFv2 for a long time...\n\nMoreover, I don't think ONNX is a stable data format as well. From my experience with working with it using NVIDIA's TensorRT and PyTorch - you often run into strange problems when trying to export/import something slightly more complex than usual CNN's. To be frank though, last time I worked with that format was about a year ago, so maybe the situation is different now. \n\nAll of those reasons lead me to thinking that it would be very difficult to make that converter stable and something more than a hacky project. But maybe I'm just a pessimist ;)",
    "974337": "There is an easier conversion method if you are not trying to re-train or fine-tune the model:\n\n1. Load the saved model using `tf.saved_model.load`, and get the inference function `infer` under signature `serving_default`\n2. Freeze the graph of that function using this tutorial: https://leimao.github.io/blog/Save-Load-Inference-From-TF2-Frozen-Graph/. This should get rid of any `tf.Variable` in the graph (this tutorial is for `tf.function`, they extract `ConcreteFunction` from it; and guess what? `infer` is a `WrappedFunction`, a class inherited from `ConcreteFunction`.\n3. Write your own custom `tf.Module` wrapper as in the official TF 2.3 documentation https://www.tensorflow.org/guide/saved_model#saving_a_custom_model\n4. Save it as SavedModel, as in the tutorial. That's all."
  },
  "source": "meta"
}