{
  "id": 163589,
  "title": "How to do data preprocessing while inference?",
  "url": "/competitions/landmark-retrieval-2020/discussion/163589",
  "author_name": "",
  "post_date": "2020-07-02T16:21:18.613269300Z",
  "votes": 11,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Hi Team,\n     i have a question related to data preprocessing during inference. the nature of the competition is to submit the trained model alone. in this case, how to specify or carryout data preprocessing, normalization if any, applied while training. </p>\n\n<p>thanks\nyuvaram </p>",
  "messages": [
    {
      "id": "912657",
      "postDate": "07/02/2020 16:21:18",
      "content": "<p>Hi Team,\n     i have a question related to data preprocessing during inference. the nature of the competition is to submit the trained model alone. in this case, how to specify or carryout data preprocessing, normalization if any, applied while training. </p>\n\n<p>thanks\nyuvaram </p>",
      "rawMarkdown": "Hi Team,\n     i have a question related to data preprocessing during inference. the nature of the competition is to submit the trained model alone. in this case, how to specify or carryout data preprocessing, normalization if any, applied while training. \n\nthanks\nyuvaram",
      "votes": null
    },
    {
      "id": "914073",
      "postDate": "07/03/2020 15:58:03",
      "content": "<p>If your model needs any type of data preprocessing, you should export the preprocessing operations within the submitted SavedModel, together with the rest of your model.</p>\n\n<p>The SavedModel should take a [H,W,3] uint8 tensor as input (see <a href=\"https://github.com/tensorflow/models/blob/master/research/delf/delf/python/training/model/export_global_model.py#L112\">here</a> for an input example), and the output should be a dict containing key 'global_descriptor' mapped to a [D] float tensor (see <a href=\"https://github.com/tensorflow/models/blob/master/research/delf/delf/python/training/model/export_global_model.py#L106\">here</a> for an output example).</p>\n\n<p>Another concrete example is the <a href=\"https://www.kaggle.com/camaskew/baseline-landmark-retrieval-model\">baseline model</a> we released. You can also try out the <a href=\"https://www.tensorflow.org/guide/saved_model#details_of_the_savedmodel_command_line_interface\">SavedModel command-line interface</a> to inspect this model, which can show you the above-mentioned input and output structure.</p>",
      "rawMarkdown": "If your model needs any type of data preprocessing, you should export the preprocessing operations within the submitted SavedModel, together with the rest of your model.\n\nThe SavedModel should take a [H,W,3] uint8 tensor as input (see [here](https://github.com/tensorflow/models/blob/master/research/delf/delf/python/training/model/export_global_model.py#L112) for an input example), and the output should be a dict containing key 'global_descriptor' mapped to a [D] float tensor (see [here](https://github.com/tensorflow/models/blob/master/research/delf/delf/python/training/model/export_global_model.py#L106) for an output example).\n\nAnother concrete example is the [baseline model](https://www.kaggle.com/camaskew/baseline-landmark-retrieval-model) we released. You can also try out the [SavedModel command-line interface](https://www.tensorflow.org/guide/saved_model#details_of_the_savedmodel_command_line_interface) to inspect this model, which can show you the above-mentioned input and output structure.",
      "votes": null
    },
    {
      "id": "914188",
      "postDate": "07/03/2020 17:05:43",
      "content": "<p>hi <a href=\"/andrefaraujo\">@andrefaraujo</a> . can you provide a simpler code snippet on how to map a model input and output into the required format . the github code is too huge to follow :P</p>",
      "rawMarkdown": "hi @andrefaraujo . can you provide a simpler code snippet on how to map a model input and output into the required format . the github code is too huge to follow :P",
      "votes": null
    },
    {
      "id": "914399",
      "postDate": "07/03/2020 20:44:16",
      "content": "<p><a href=\"/andrefaraujo\">@andrefaraujo</a> I think this should be put up in a place where everyone can see it. I have spent 6-7 hours trying to work out this. No one is going to adopt TF2.0 if it isn't made user friendly and if there aren't any documentations available.</p>",
      "rawMarkdown": "andrefaraujo I think this should be put up in a place where everyone can see it. I have spent 6-7 hours trying to work out this. No one is going to adopt TF2.0 if it isn't made user friendly and if there aren't any documentations available.",
      "votes": null
    },
    {
      "id": "914421",
      "postDate": "07/03/2020 21:07:55",
      "content": "<p><a href=\"/mayukh18\">@mayukh18</a>  do you have a solution for mapping the layers . is it possible to share code snippet. </p>",
      "rawMarkdown": "mayukh18  do you have a solution for mapping the layers . is it possible to share code snippet.",
      "votes": null
    },
    {
      "id": "914429",
      "postDate": "07/03/2020 21:25:25",
      "content": "<p>Currently working on it. Haven't got anything concrete yet.</p>",
      "rawMarkdown": "Currently working on it. Haven't got anything concrete yet.",
      "votes": null
    },
    {
      "id": "916081",
      "postDate": "07/05/2020 10:39:05",
      "content": "<p><a href=\"/andrefaraujo\">@andrefaraujo</a>  can we have a batch dimension in the input (None,None,None,3) . is this acceptable?</p>",
      "rawMarkdown": "andrefaraujo  can we have a batch dimension in the input (None,None,None,3) . is this acceptable?",
      "votes": null
    },
    {
      "id": "916489",
      "postDate": "07/05/2020 17:42:45",
      "content": "<p>No, the input must be of shape [None,None,3].</p>\n\n<p>To export a model that is trained with [None,None,None,3] shapes, you can follow a recipe very similar to the one used in the code linked in the <a href=\"https://www.kaggle.com/c/landmark-retrieval-2020/discussion/163350\">\"second submission\" post</a>. Essentially, just define the exported function to have an input [None,None,3] then expand the dimensions. Something like this:</p>\n\n<p><code>python\n@tf.function(input_signature=[\n      tf.TensorSpec(shape=[None, None, 3], dtype=tf.uint8, name='input_image')\n  ])\ndef YourModelExportFunction(self, input_image):\n  image_tensor = tf.expand_dims(image_tensor, 0, name='image/expand_dims')\n  extracted_feature = your_function_that_computes_embedding(image_tensor)\n  # Note: you may need to reshape \"extracted_feature\" in order to make it have \n  # shape (D,).\n  named_output_tensors[] = {}\n  named_output_tensors['global_descriptor'] = tf.identity(\n          extracted_feature, name='global_descriptor')\n  return named_output_tensors\n</code></p>\n\n<p>For the codebase I mentioned, you can see that this is done <a href=\"https://github.com/tensorflow/models/blob/master/research/delf/delf/python/training/model/export_model_utils.py#L213\">here</a>.</p>",
      "rawMarkdown": "No, the input must be of shape [None,None,3].\n\nTo export a model that is trained with [None,None,None,3] shapes, you can follow a recipe very similar to the one used in the code linked in the [\"second submission\" post](https://www.kaggle.com/c/landmark-retrieval-2020/discussion/163350). Essentially, just define the exported function to have an input [None,None,3] then expand the dimensions. Something like this:\n\n```python\n@tf.function(input_signature=[\n      tf.TensorSpec(shape=[None, None, 3], dtype=tf.uint8, name='input_image')\n  ])\ndef YourModelExportFunction(self, input_image):\n  image_tensor = tf.expand_dims(image_tensor, 0, name='image/expand_dims')\n  extracted_feature = your_function_that_computes_embedding(image_tensor)\n  # Note: you may need to reshape \"extracted_feature\" in order to make it have \n  # shape (D,).\n  named_output_tensors[] = {}\n  named_output_tensors['global_descriptor'] = tf.identity(\n          extracted_feature, name='global_descriptor')\n  return named_output_tensors\n```\n\nFor the codebase I mentioned, you can see that this is done [here](https://github.com/tensorflow/models/blob/master/research/delf/delf/python/training/model/export_model_utils.py#L213).",
      "votes": null
    },
    {
      "id": "916979",
      "postDate": "07/06/2020 06:51:02",
      "content": "<p>soo.. there is no requirement of \"named_output_tensors\" dict now, we can return the output directly..</p>",
      "rawMarkdown": "soo.. there is no requirement of \"named_output_tensors\" dict now, we can return the output directly..",
      "votes": null
    },
    {
      "id": "917548",
      "postDate": "07/06/2020 15:53:26",
      "content": "<p>Sorry, I forgot to include that part -- I was emphasizing the input handling. Just updated it.</p>",
      "rawMarkdown": "Sorry, I forgot to include that part -- I was emphasizing the input handling. Just updated it.",
      "votes": null
    },
    {
      "id": "918209",
      "postDate": "07/07/2020 04:44:58",
      "content": "<p>why even the failed submissions takes a long time. Could you please look into it?</p>",
      "rawMarkdown": "why even the failed submissions takes a long time. Could you please look into it?",
      "votes": null
    },
    {
      "id": "918667",
      "postDate": "07/07/2020 11:45:08",
      "content": "<p>I am trying to implement this, but I run into a problem:\nYou can't save a tf2 model with subclassed keras model.</p>\n\n<p>Source:\n<a href=\"https://stackoverflow.com/questions/51806852/cant-save-custom-subclassed-model\">https://stackoverflow.com/questions/51806852/cant-save-custom-subclassed-model</a></p>\n\n<p>Is there any way to use Keras for this competition?</p>",
      "rawMarkdown": "I am trying to implement this, but I run into a problem:\nYou can't save a tf2 model with subclassed keras model.\n\nSource:\nhttps://stackoverflow.com/questions/51806852/cant-save-custom-subclassed-model\n\nIs there any way to use Keras for this competition?",
      "votes": null
    },
    {
      "id": "919028",
      "postDate": "07/07/2020 17:04:48",
      "content": "<p>Yes, absolutely, Keras can be directly used. As a matter of fact, the example in our <a href=\"https://www.kaggle.com/c/landmark-retrieval-2020/discussion/163350\">second submission post</a> is using a Keras model.</p>",
      "rawMarkdown": "Yes, absolutely, Keras can be directly used. As a matter of fact, the example in our [second submission post](https://www.kaggle.com/c/landmark-retrieval-2020/discussion/163350) is using a Keras model.",
      "votes": null
    },
    {
      "id": "919067",
      "postDate": "07/07/2020 17:31:08",
      "content": "<p>Sorry, you are right. Managed to get it to work. I'll make a notebook for others.</p>",
      "rawMarkdown": "Sorry, you are right. Managed to get it to work. I'll make a notebook for others.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 914073,
      "author_name": "andrefaraujo",
      "author_url": "",
      "post_date": "07/03/2020 15:58:03",
      "content": "<p>If your model needs any type of data preprocessing, you should export the preprocessing operations within the submitted SavedModel, together with the rest of your model.</p>\n\n<p>The SavedModel should take a [H,W,3] uint8 tensor as input (see <a href=\"https://github.com/tensorflow/models/blob/master/research/delf/delf/python/training/model/export_global_model.py#L112\">here</a> for an input example), and the output should be a dict containing key 'global_descriptor' mapped to a [D] float tensor (see <a href=\"https://github.com/tensorflow/models/blob/master/research/delf/delf/python/training/model/export_global_model.py#L106\">here</a> for an output example).</p>\n\n<p>Another concrete example is the <a href=\"https://www.kaggle.com/camaskew/baseline-landmark-retrieval-model\">baseline model</a> we released. You can also try out the <a href=\"https://www.tensorflow.org/guide/saved_model#details_of_the_savedmodel_command_line_interface\">SavedModel command-line interface</a> to inspect this model, which can show you the above-mentioned input and output structure.</p>",
      "votes": null,
      "replies": [
        {
          "id": 914188,
          "author_name": "yuvaramsingh",
          "author_url": "",
          "post_date": "07/03/2020 17:05:43",
          "content": "<p>hi <a href=\"/andrefaraujo\">@andrefaraujo</a> . can you provide a simpler code snippet on how to map a model input and output into the required format . the github code is too huge to follow :P</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 914399,
          "author_name": "mayukh18",
          "author_url": "",
          "post_date": "07/03/2020 20:44:16",
          "content": "<p><a href=\"/andrefaraujo\">@andrefaraujo</a> I think this should be put up in a place where everyone can see it. I have spent 6-7 hours trying to work out this. No one is going to adopt TF2.0 if it isn't made user friendly and if there aren't any documentations available.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 914421,
          "author_name": "yuvaramsingh",
          "author_url": "",
          "post_date": "07/03/2020 21:07:55",
          "content": "<p><a href=\"/mayukh18\">@mayukh18</a>  do you have a solution for mapping the layers . is it possible to share code snippet. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 914429,
          "author_name": "mayukh18",
          "author_url": "",
          "post_date": "07/03/2020 21:25:25",
          "content": "<p>Currently working on it. Haven't got anything concrete yet.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 916081,
          "author_name": "yuvaramsingh",
          "author_url": "",
          "post_date": "07/05/2020 10:39:05",
          "content": "<p><a href=\"/andrefaraujo\">@andrefaraujo</a>  can we have a batch dimension in the input (None,None,None,3) . is this acceptable?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 916489,
          "author_name": "andrefaraujo",
          "author_url": "",
          "post_date": "07/05/2020 17:42:45",
          "content": "<p>No, the input must be of shape [None,None,3].</p>\n\n<p>To export a model that is trained with [None,None,None,3] shapes, you can follow a recipe very similar to the one used in the code linked in the <a href=\"https://www.kaggle.com/c/landmark-retrieval-2020/discussion/163350\">\"second submission\" post</a>. Essentially, just define the exported function to have an input [None,None,3] then expand the dimensions. Something like this:</p>\n\n<p><code>python\n@tf.function(input_signature=[\n      tf.TensorSpec(shape=[None, None, 3], dtype=tf.uint8, name='input_image')\n  ])\ndef YourModelExportFunction(self, input_image):\n  image_tensor = tf.expand_dims(image_tensor, 0, name='image/expand_dims')\n  extracted_feature = your_function_that_computes_embedding(image_tensor)\n  # Note: you may need to reshape \"extracted_feature\" in order to make it have \n  # shape (D,).\n  named_output_tensors[] = {}\n  named_output_tensors['global_descriptor'] = tf.identity(\n          extracted_feature, name='global_descriptor')\n  return named_output_tensors\n</code></p>\n\n<p>For the codebase I mentioned, you can see that this is done <a href=\"https://github.com/tensorflow/models/blob/master/research/delf/delf/python/training/model/export_model_utils.py#L213\">here</a>.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 916979,
          "author_name": "udaygurugubelli",
          "author_url": "",
          "post_date": "07/06/2020 06:51:02",
          "content": "<p>soo.. there is no requirement of \"named_output_tensors\" dict now, we can return the output directly..</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 917548,
          "author_name": "andrefaraujo",
          "author_url": "",
          "post_date": "07/06/2020 15:53:26",
          "content": "<p>Sorry, I forgot to include that part -- I was emphasizing the input handling. Just updated it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 918209,
          "author_name": "udaygurugubelli",
          "author_url": "",
          "post_date": "07/07/2020 04:44:58",
          "content": "<p>why even the failed submissions takes a long time. Could you please look into it?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 918667,
          "author_name": "reflexion",
          "author_url": "",
          "post_date": "07/07/2020 11:45:08",
          "content": "<p>I am trying to implement this, but I run into a problem:\nYou can't save a tf2 model with subclassed keras model.</p>\n\n<p>Source:\n<a href=\"https://stackoverflow.com/questions/51806852/cant-save-custom-subclassed-model\">https://stackoverflow.com/questions/51806852/cant-save-custom-subclassed-model</a></p>\n\n<p>Is there any way to use Keras for this competition?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 919028,
          "author_name": "andrefaraujo",
          "author_url": "",
          "post_date": "07/07/2020 17:04:48",
          "content": "<p>Yes, absolutely, Keras can be directly used. As a matter of fact, the example in our <a href=\"https://www.kaggle.com/c/landmark-retrieval-2020/discussion/163350\">second submission post</a> is using a Keras model.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 919067,
          "author_name": "reflexion",
          "author_url": "",
          "post_date": "07/07/2020 17:31:08",
          "content": "<p>Sorry, you are right. Managed to get it to work. I'll make a notebook for others.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "912657": "Hi Team,\n     i have a question related to data preprocessing during inference. the nature of the competition is to submit the trained model alone. in this case, how to specify or carryout data preprocessing, normalization if any, applied while training. \n\nthanks\nyuvaram",
    "914073": "If your model needs any type of data preprocessing, you should export the preprocessing operations within the submitted SavedModel, together with the rest of your model.\n\nThe SavedModel should take a [H,W,3] uint8 tensor as input (see [here](https://github.com/tensorflow/models/blob/master/research/delf/delf/python/training/model/export_global_model.py#L112) for an input example), and the output should be a dict containing key 'global_descriptor' mapped to a [D] float tensor (see [here](https://github.com/tensorflow/models/blob/master/research/delf/delf/python/training/model/export_global_model.py#L106) for an output example).\n\nAnother concrete example is the [baseline model](https://www.kaggle.com/camaskew/baseline-landmark-retrieval-model) we released. You can also try out the [SavedModel command-line interface](https://www.tensorflow.org/guide/saved_model#details_of_the_savedmodel_command_line_interface) to inspect this model, which can show you the above-mentioned input and output structure.",
    "914188": "hi @andrefaraujo . can you provide a simpler code snippet on how to map a model input and output into the required format . the github code is too huge to follow :P",
    "914399": "andrefaraujo I think this should be put up in a place where everyone can see it. I have spent 6-7 hours trying to work out this. No one is going to adopt TF2.0 if it isn't made user friendly and if there aren't any documentations available.",
    "914421": "mayukh18  do you have a solution for mapping the layers . is it possible to share code snippet.",
    "914429": "Currently working on it. Haven't got anything concrete yet.",
    "916081": "andrefaraujo  can we have a batch dimension in the input (None,None,None,3) . is this acceptable?",
    "916489": "No, the input must be of shape [None,None,3].\n\nTo export a model that is trained with [None,None,None,3] shapes, you can follow a recipe very similar to the one used in the code linked in the [\"second submission\" post](https://www.kaggle.com/c/landmark-retrieval-2020/discussion/163350). Essentially, just define the exported function to have an input [None,None,3] then expand the dimensions. Something like this:\n\n```python\n@tf.function(input_signature=[\n      tf.TensorSpec(shape=[None, None, 3], dtype=tf.uint8, name='input_image')\n  ])\ndef YourModelExportFunction(self, input_image):\n  image_tensor = tf.expand_dims(image_tensor, 0, name='image/expand_dims')\n  extracted_feature = your_function_that_computes_embedding(image_tensor)\n  # Note: you may need to reshape \"extracted_feature\" in order to make it have \n  # shape (D,).\n  named_output_tensors[] = {}\n  named_output_tensors['global_descriptor'] = tf.identity(\n          extracted_feature, name='global_descriptor')\n  return named_output_tensors\n```\n\nFor the codebase I mentioned, you can see that this is done [here](https://github.com/tensorflow/models/blob/master/research/delf/delf/python/training/model/export_model_utils.py#L213).",
    "916979": "soo.. there is no requirement of \"named_output_tensors\" dict now, we can return the output directly..",
    "917548": "Sorry, I forgot to include that part -- I was emphasizing the input handling. Just updated it.",
    "918209": "why even the failed submissions takes a long time. Could you please look into it?",
    "918667": "I am trying to implement this, but I run into a problem:\nYou can't save a tf2 model with subclassed keras model.\n\nSource:\nhttps://stackoverflow.com/questions/51806852/cant-save-custom-subclassed-model\n\nIs there any way to use Keras for this competition?",
    "919028": "Yes, absolutely, Keras can be directly used. As a matter of fact, the example in our [second submission post](https://www.kaggle.com/c/landmark-retrieval-2020/discussion/163350) is using a Keras model.",
    "919067": "Sorry, you are right. Managed to get it to work. I'll make a notebook for others."
  },
  "source": "meta"
}