{
  "id": 391301,
  "title": "Handling Pre-processing in PyTorch",
  "url": "/competitions/asl-signs/discussion/391301",
  "author_name": "Mayukh Bhattacharyya",
  "post_date": "2023-03-01T02:05:23.577000",
  "votes": 25,
  "comment_count": 0,
  "views": 0,
  "content": "<p>So far in the competition, the biggest obstacle seemed to be the pre-processing of the incoming data during inference. I think I have found a good modular way to serve that purpose without having to convert your pytorch code inside the tensorflow model during inference. This <a href=\"https://www.kaggle.com/code/mayukh18/end-to-end-pytorch-training-submission\" target=\"_blank\">notebook</a> shows the whole solution.</p>\n<p><strong>The Problem:</strong> Inference passes a tensor of (n_frames, 543, 3) dimensions. Well we can trainour pytorch model have an input of (n_frames, 543, 3) right? Yes, but then we have to do a lot of tensor gymnastics to train that model in batches.</p>\n<p><strong>The Solution:</strong> Having a separate pre-processing or feature generation model as I like to call it. It goes something like:</p>\n<pre><code>class FeatureGen(nn.Module):\n    def __init__(self):\n        super(FeatureGen, self).__init__()\n        pass\n\n    def forward(self, x):\n        face_x = x[:,:468,:].contiguous().view(-1, 468*3)\n        lefth_x = x[:,468:489,:].contiguous().view(-1, 21*3)\n        pose_x = x[:,489:522,:].contiguous().view(-1, 33*3)\n        righth_x = x[:,522:,:].contiguous().view(-1, 21*3)\n\n        lefth_x = lefth_x[~torch.any(torch.isnan(lefth_x), dim=1),:]\n        righth_x = righth_x[~torch.any(torch.isnan(righth_x), dim=1),:]\n\n        x1m = torch.mean(face_x, 0)\n        x2m = torch.mean(lefth_x, 0)\n        x3m = torch.mean(pose_x, 0)\n        x4m = torch.mean(righth_x, 0)\n\n        x1s = torch.std(face_x, 0)\n        x2s = torch.std(lefth_x, 0)\n        x3s = torch.std(pose_x, 0)\n        x4s = torch.std(righth_x, 0)\n\n        xfeat = torch.cat([x1m,x2m,x3m,x4m, x1s,x2s,x3s,x4s], axis=0)\n        xfeat = torch.where(torch.isnan(xfeat), torch.tensor(0.0, dtype=torch.float32), xfeat)\n\n        return xfeat\n</code></pre>\n<p>It converts a (n_frames, 543, 3) array to (n_features,) array which we can use both in training and inference. Do note how it does not have any trainable parameters. Because we don't want it to have any. What this does is that it bundles all the torch computations within a torch Model which ONNX knows very well how to convert. </p>\n<p><strong>Inference:</strong> During inference we need to call this model to convert the data before we call our actual model. Something like:</p>\n<pre><code>class TFInference(tf.Module):\n    def __init__(self):\n        super(TFInference, self).__init__()\n        self.feature_gen = tf.saved_model.load(tf_feat_gen_path)\n        self.model = tf.saved_model.load(tf_model_path)\n        self.feature_gen.trainable = False\n        self.model.trainable = False\n\n    @tf.function(input_signature=[\n      tf.TensorSpec(shape=[None, 543, 3], dtype=tf.float32, name='inputs')\n    ])\n    def call(self, input):\n        output_tensors = {}\n        features = self.feature_gen(**{'input': input})['output']\n        output_tensors['outputs'] = self.model(**{'input': tf.expand_dims(features, 0)})['output'][0,:]\n        return output_tensors\n</code></pre>\n<p>Don't forget to expand your dimensions to make it (1,n_features) because that's the shape your model will remember from training.</p>",
  "messages": [
    {
      "id": 2163686,
      "postDate": "2023-03-01T02:05:23.577Z",
      "content": "<p>So far in the competition, the biggest obstacle seemed to be the pre-processing of the incoming data during inference. I think I have found a good modular way to serve that purpose without having to convert your pytorch code inside the tensorflow model during inference. This <a href=\"https://www.kaggle.com/code/mayukh18/end-to-end-pytorch-training-submission\" target=\"_blank\">notebook</a> shows the whole solution.</p>\n<p><strong>The Problem:</strong> Inference passes a tensor of (n_frames, 543, 3) dimensions. Well we can trainour pytorch model have an input of (n_frames, 543, 3) right? Yes, but then we have to do a lot of tensor gymnastics to train that model in batches.</p>\n<p><strong>The Solution:</strong> Having a separate pre-processing or feature generation model as I like to call it. It goes something like:</p>\n<pre><code>class FeatureGen(nn.Module):\n    def __init__(self):\n        super(FeatureGen, self).__init__()\n        pass\n\n    def forward(self, x):\n        face_x = x[:,:468,:].contiguous().view(-1, 468*3)\n        lefth_x = x[:,468:489,:].contiguous().view(-1, 21*3)\n        pose_x = x[:,489:522,:].contiguous().view(-1, 33*3)\n        righth_x = x[:,522:,:].contiguous().view(-1, 21*3)\n\n        lefth_x = lefth_x[~torch.any(torch.isnan(lefth_x), dim=1),:]\n        righth_x = righth_x[~torch.any(torch.isnan(righth_x), dim=1),:]\n\n        x1m = torch.mean(face_x, 0)\n        x2m = torch.mean(lefth_x, 0)\n        x3m = torch.mean(pose_x, 0)\n        x4m = torch.mean(righth_x, 0)\n\n        x1s = torch.std(face_x, 0)\n        x2s = torch.std(lefth_x, 0)\n        x3s = torch.std(pose_x, 0)\n        x4s = torch.std(righth_x, 0)\n\n        xfeat = torch.cat([x1m,x2m,x3m,x4m, x1s,x2s,x3s,x4s], axis=0)\n        xfeat = torch.where(torch.isnan(xfeat), torch.tensor(0.0, dtype=torch.float32), xfeat)\n\n        return xfeat\n</code></pre>\n<p>It converts a (n_frames, 543, 3) array to (n_features,) array which we can use both in training and inference. Do note how it does not have any trainable parameters. Because we don't want it to have any. What this does is that it bundles all the torch computations within a torch Model which ONNX knows very well how to convert. </p>\n<p><strong>Inference:</strong> During inference we need to call this model to convert the data before we call our actual model. Something like:</p>\n<pre><code>class TFInference(tf.Module):\n    def __init__(self):\n        super(TFInference, self).__init__()\n        self.feature_gen = tf.saved_model.load(tf_feat_gen_path)\n        self.model = tf.saved_model.load(tf_model_path)\n        self.feature_gen.trainable = False\n        self.model.trainable = False\n\n    @tf.function(input_signature=[\n      tf.TensorSpec(shape=[None, 543, 3], dtype=tf.float32, name='inputs')\n    ])\n    def call(self, input):\n        output_tensors = {}\n        features = self.feature_gen(**{'input': input})['output']\n        output_tensors['outputs'] = self.model(**{'input': tf.expand_dims(features, 0)})['output'][0,:]\n        return output_tensors\n</code></pre>\n<p>Don't forget to expand your dimensions to make it (1,n_features) because that's the shape your model will remember from training.</p>",
      "rawMarkdown": "So far in the competition, the biggest obstacle seemed to be the pre-processing of the incoming data during inference. I think I have found a good modular way to serve that purpose without having to convert your pytorch code inside the tensorflow model during inference. This [notebook](https://www.kaggle.com/code/mayukh18/end-to-end-pytorch-training-submission) shows the whole solution.\n\n**The Problem:** Inference passes a tensor of (n_frames, 543, 3) dimensions. Well we can trainour pytorch model have an input of (n_frames, 543, 3) right? Yes, but then we have to do a lot of tensor gymnastics to train that model in batches.\n\n**The Solution:** Having a separate pre-processing or feature generation model as I like to call it. It goes something like:\n\n```\nclass FeatureGen(nn.Module):\n    def __init__(self):\n        super(FeatureGen, self).__init__()\n        pass\n    \n    def forward(self, x):\n        face_x = x[:,:468,:].contiguous().view(-1, 468*3)\n        lefth_x = x[:,468:489,:].contiguous().view(-1, 21*3)\n        pose_x = x[:,489:522,:].contiguous().view(-1, 33*3)\n        righth_x = x[:,522:,:].contiguous().view(-1, 21*3)\n        \n        lefth_x = lefth_x[~torch.any(torch.isnan(lefth_x), dim=1),:]\n        righth_x = righth_x[~torch.any(torch.isnan(righth_x), dim=1),:]\n        \n        x1m = torch.mean(face_x, 0)\n        x2m = torch.mean(lefth_x, 0)\n        x3m = torch.mean(pose_x, 0)\n        x4m = torch.mean(righth_x, 0)\n        \n        x1s = torch.std(face_x, 0)\n        x2s = torch.std(lefth_x, 0)\n        x3s = torch.std(pose_x, 0)\n        x4s = torch.std(righth_x, 0)\n        \n        xfeat = torch.cat([x1m,x2m,x3m,x4m, x1s,x2s,x3s,x4s], axis=0)\n        xfeat = torch.where(torch.isnan(xfeat), torch.tensor(0.0, dtype=torch.float32), xfeat)\n        \n        return xfeat\n```\nIt converts a (n_frames, 543, 3) array to (n_features,) array which we can use both in training and inference. Do note how it does not have any trainable parameters. Because we don't want it to have any. What this does is that it bundles all the torch computations within a torch Model which ONNX knows very well how to convert. \n\n**Inference:** During inference we need to call this model to convert the data before we call our actual model. Something like:\n\n```\nclass TFInference(tf.Module):\n    def __init__(self):\n        super(TFInference, self).__init__()\n        self.feature_gen = tf.saved_model.load(tf_feat_gen_path)\n        self.model = tf.saved_model.load(tf_model_path)\n        self.feature_gen.trainable = False\n        self.model.trainable = False\n    \n    @tf.function(input_signature=[\n      tf.TensorSpec(shape=[None, 543, 3], dtype=tf.float32, name='inputs')\n    ])\n    def call(self, input):\n        output_tensors = {}\n        features = self.feature_gen(**{'input': input})['output']\n        output_tensors['outputs'] = self.model(**{'input': tf.expand_dims(features, 0)})['output'][0,:]\n        return output_tensors\n```\n\nDon't forget to expand your dimensions to make it (1,n_features) because that's the shape your model will remember from training.",
      "votes": 24
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2163686": "So far in the competition, the biggest obstacle seemed to be the pre-processing of the incoming data during inference. I think I have found a good modular way to serve that purpose without having to convert your pytorch code inside the tensorflow model during inference. This [notebook](https://www.kaggle.com/code/mayukh18/end-to-end-pytorch-training-submission) shows the whole solution.\n\n**The Problem:** Inference passes a tensor of (n_frames, 543, 3) dimensions. Well we can trainour pytorch model have an input of (n_frames, 543, 3) right? Yes, but then we have to do a lot of tensor gymnastics to train that model in batches.\n\n**The Solution:** Having a separate pre-processing or feature generation model as I like to call it. It goes something like:\n\n```\nclass FeatureGen(nn.Module):\n    def __init__(self):\n        super(FeatureGen, self).__init__()\n        pass\n    \n    def forward(self, x):\n        face_x = x[:,:468,:].contiguous().view(-1, 468*3)\n        lefth_x = x[:,468:489,:].contiguous().view(-1, 21*3)\n        pose_x = x[:,489:522,:].contiguous().view(-1, 33*3)\n        righth_x = x[:,522:,:].contiguous().view(-1, 21*3)\n        \n        lefth_x = lefth_x[~torch.any(torch.isnan(lefth_x), dim=1),:]\n        righth_x = righth_x[~torch.any(torch.isnan(righth_x), dim=1),:]\n        \n        x1m = torch.mean(face_x, 0)\n        x2m = torch.mean(lefth_x, 0)\n        x3m = torch.mean(pose_x, 0)\n        x4m = torch.mean(righth_x, 0)\n        \n        x1s = torch.std(face_x, 0)\n        x2s = torch.std(lefth_x, 0)\n        x3s = torch.std(pose_x, 0)\n        x4s = torch.std(righth_x, 0)\n        \n        xfeat = torch.cat([x1m,x2m,x3m,x4m, x1s,x2s,x3s,x4s], axis=0)\n        xfeat = torch.where(torch.isnan(xfeat), torch.tensor(0.0, dtype=torch.float32), xfeat)\n        \n        return xfeat\n```\nIt converts a (n_frames, 543, 3) array to (n_features,) array which we can use both in training and inference. Do note how it does not have any trainable parameters. Because we don't want it to have any. What this does is that it bundles all the torch computations within a torch Model which ONNX knows very well how to convert. \n\n**Inference:** During inference we need to call this model to convert the data before we call our actual model. Something like:\n\n```\nclass TFInference(tf.Module):\n    def __init__(self):\n        super(TFInference, self).__init__()\n        self.feature_gen = tf.saved_model.load(tf_feat_gen_path)\n        self.model = tf.saved_model.load(tf_model_path)\n        self.feature_gen.trainable = False\n        self.model.trainable = False\n    \n    @tf.function(input_signature=[\n      tf.TensorSpec(shape=[None, 543, 3], dtype=tf.float32, name='inputs')\n    ])\n    def call(self, input):\n        output_tensors = {}\n        features = self.feature_gen(**{'input': input})['output']\n        output_tensors['outputs'] = self.model(**{'input': tf.expand_dims(features, 0)})['output'][0,:]\n        return output_tensors\n```\n\nDon't forget to expand your dimensions to make it (1,n_features) because that's the shape your model will remember from training."
  }
}