{
  "id": 406354,
  "title": "Top 8% Bronze Medal Solution",
  "url": "/competitions/asl-signs/writeups/abhinand-top-8-bronze-medal-solution",
  "author_name": "",
  "post_date": "2023-05-02T04:01:15.050587600Z",
  "votes": 7,
  "comment_count": 1,
  "views": 0,
  "content": "<blockquote>\n  <p><strong>Many congratulations to all the winners in this competition. Huge thanks to all the discussion topics, public notebooks and datasets - most (if not all of) of my learning came from them.</strong></p>\n</blockquote>\n<p>First I would like to thank the organizers of this wonderful competition. </p>\n<p>To say that I learned a lot by competing in this competition would really be an understatement. I wanted to take on the challenge of using Tensorflow in this competition and boy did it test me. If there was one thing that gave me hope in this competition it was <a href=\"https://www.kaggle.com/markwijkhuizen\" target=\"_blank\">@markwijkhuizen</a>'s amazing public <a href=\"https://www.kaggle.com/code/markwijkhuizen/gislr-tf-data-processing-transformer-training\" target=\"_blank\">notebook</a>. It was the foundation to my entire work in this competition.</p>\n<p>Initially I was using my own Transformer architectures in this competition which were not bad themselves but as <a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> pointed it, the public notebooks had really good Transformer architectures, with a few tweaks and code cleanups I was able to extract good results both on CV and LB, this gave me a significant increase in CV and LB right on my first experiment. </p>\n<blockquote>\n  <p>I have open sourced the code base on GitHub -&gt; <a href=\"https://github.com/abhinand5/isolated-sign-language-recognition\" target=\"_blank\">https://github.com/abhinand5/isolated-sign-language-recognition</a>. </p>\n  <p>Leave a ⭐️ if you like my work.</p>\n</blockquote>\n<h3>What worked for me?</h3>\n<ul>\n<li>Flip augmentation (<code>aug_prob = 0.4</code>)</li>\n<li>10 Fold Cross Validation</li>\n<li><code>label_smoothing = 0.3</code></li>\n<li><code>embed_dim = 256</code></li>\n<li><code>input_size = 32</code></li>\n<li><code>num_transformer_blocks = 2</code> (Using 1 was under-fitting and 3 was over fitting even after several experiments of trying to tune)</li>\n<li>Attention heads: <code>MHA_HEADS = 4</code> and <code>MHA_HEADS = 8</code></li>\n<li>Train on the entire dataset before submission</li>\n<li>FP16 Quantization</li>\n<li>GeLU activation</li>\n</ul>\n<h3>What did NOT work for me?</h3>\n<ul>\n<li><code>input_size = 64</code></li>\n<li>Knowledge Distillation (wasted a lot of time in this)</li>\n<li>Rotate augmentations (degraded CV and LB)</li>\n<li>Adding Gaussian noise to the inputs (degraded CV and LB)</li>\n<li>Removing <code>LayerNorm</code> from Transformer blocks (although the arch is shallow it did not work for me)</li>\n<li>Strip Pruning </li>\n<li>Dynamic Range Quantization (increased the inference time)</li>\n<li>2D Affine transforms</li>\n</ul>\n<h3>Final Solution:</h3>\n<p>There were two different model architectures in the final solution. </p>\n<p><strong>Dataset:</strong> Only Competition Data</p>\n<p><strong>Model 1:</strong></p>\n<ul>\n<li>Training <code>epochs = 50</code></li>\n<li>Add LayerNorm</li>\n<li>CV Strategy: <code>StratifiedGroupKFold(k=10)</code></li>\n<li>Optuna tuning -&gt; Best OOF on all folds</li>\n<li><code>SEED = 555</code></li>\n<li>Ensemble of Best Folds: <code>[3,6,7,9]</code> (based on OOF score)</li>\n<li><code>QUANTIZE_MODEL = True</code></li>\n<li><code>QUANT_METHOD = \"float16\"</code></li>\n<li>No Augmentations</li>\n</ul>\n<p>Below is the complete model conf:</p>\n<pre><code>\n  \n   \n   \n   \n  \n   \n  \n   \n\n  \n   \n   \n   \n   \n  \n   \n   \n   \n   \n   \n</code></pre>\n<p><strong>Model 2:</strong></p>\n<ul>\n<li>Training <code>epochs = 100</code></li>\n<li>Add LayerNorm</li>\n<li>Single Split -&gt; Optuna Tuning -&gt; Use best hyper params to train on entire train data</li>\n<li><code>SEED = 555</code></li>\n<li><code>QUANTIZE_MODEL = True</code></li>\n<li><code>QUANT_METHOD = \"float16\"</code></li>\n<li>Flip Augmentations (<code>aug_prob=0.4</code>)</li>\n</ul>\n<p>Below is the complete model conf:</p>\n<pre><code>\n  \n   \n   \n   \n  \n   \n  \n   \n\n  \n   \n   \n   \n   \n  \n   \n   \n   \n   \n   \n</code></pre>\n<p><strong>Submission:</strong><br>\nThe submission is a simple ensemble of 5 models </p>\n<pre><code>Model 1 (folds=[3,6,7,9]) + Model 2\n\nPrivate Score: 0.82815\n\nPublic Score: 0.73896\n</code></pre>\n<p>Ensembling code can be found below:</p>\n<pre><code> (tf.Module):\n     ():\n        (TFLiteEnsembleModel, self).__init__()\n\n        \n        self.preprocess_layer = preprocess_layer\n        self.models = models\n\n        \n        self.weights = weights  [/(models)] * (models)\n        self.weights = tf.reshape(self.weights, (-, ))\n\n\n     ():\n        \n        x, non_empty_frame_idxs = self.proprocess_inputs(inputs)\n\n        outputs = []\n        \n         _model  self.models:\n            output = _model({ : x, : non_empty_frame_idxs })\n            outputs.append(output)\n\n        \n        outputs = tf.concat(outputs, axis=)\n        weighted_outputs = outputs * self.weights\n        weighted_sum = tf.reduce_sum(weighted_outputs, axis=, keepdims=)\n\n        \n        outputs = tf.squeeze(weighted_sum, axis=)\n\n        \n         {: outputs}\n\n     ():\n        \n        x, non_empty_frame_idxs = self.preprocess_layer(inputs)\n        \n        x = tf.expand_dims(x, axis=)\n        non_empty_frame_idxs = tf.expand_dims(non_empty_frame_idxs, axis=)\n\n         x, non_empty_frame_idxs\n</code></pre>",
  "messages": [
    {
      "id": "2242131",
      "postDate": "05/02/2023 04:01:15",
      "content": "<blockquote>\n  <p><strong>Many congratulations to all the winners in this competition. Huge thanks to all the discussion topics, public notebooks and datasets - most (if not all of) of my learning came from them.</strong></p>\n</blockquote>\n<p>First I would like to thank the organizers of this wonderful competition. </p>\n<p>To say that I learned a lot by competing in this competition would really be an understatement. I wanted to take on the challenge of using Tensorflow in this competition and boy did it test me. If there was one thing that gave me hope in this competition it was <a href=\"https://www.kaggle.com/markwijkhuizen\" target=\"_blank\">@markwijkhuizen</a>'s amazing public <a href=\"https://www.kaggle.com/code/markwijkhuizen/gislr-tf-data-processing-transformer-training\" target=\"_blank\">notebook</a>. It was the foundation to my entire work in this competition.</p>\n<p>Initially I was using my own Transformer architectures in this competition which were not bad themselves but as <a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> pointed it, the public notebooks had really good Transformer architectures, with a few tweaks and code cleanups I was able to extract good results both on CV and LB, this gave me a significant increase in CV and LB right on my first experiment. </p>\n<blockquote>\n  <p>I have open sourced the code base on GitHub -&gt; <a href=\"https://github.com/abhinand5/isolated-sign-language-recognition\" target=\"_blank\">https://github.com/abhinand5/isolated-sign-language-recognition</a>. </p>\n  <p>Leave a ⭐️ if you like my work.</p>\n</blockquote>\n<h3>What worked for me?</h3>\n<ul>\n<li>Flip augmentation (<code>aug_prob = 0.4</code>)</li>\n<li>10 Fold Cross Validation</li>\n<li><code>label_smoothing = 0.3</code></li>\n<li><code>embed_dim = 256</code></li>\n<li><code>input_size = 32</code></li>\n<li><code>num_transformer_blocks = 2</code> (Using 1 was under-fitting and 3 was over fitting even after several experiments of trying to tune)</li>\n<li>Attention heads: <code>MHA_HEADS = 4</code> and <code>MHA_HEADS = 8</code></li>\n<li>Train on the entire dataset before submission</li>\n<li>FP16 Quantization</li>\n<li>GeLU activation</li>\n</ul>\n<h3>What did NOT work for me?</h3>\n<ul>\n<li><code>input_size = 64</code></li>\n<li>Knowledge Distillation (wasted a lot of time in this)</li>\n<li>Rotate augmentations (degraded CV and LB)</li>\n<li>Adding Gaussian noise to the inputs (degraded CV and LB)</li>\n<li>Removing <code>LayerNorm</code> from Transformer blocks (although the arch is shallow it did not work for me)</li>\n<li>Strip Pruning </li>\n<li>Dynamic Range Quantization (increased the inference time)</li>\n<li>2D Affine transforms</li>\n</ul>\n<h3>Final Solution:</h3>\n<p>There were two different model architectures in the final solution. </p>\n<p><strong>Dataset:</strong> Only Competition Data</p>\n<p><strong>Model 1:</strong></p>\n<ul>\n<li>Training <code>epochs = 50</code></li>\n<li>Add LayerNorm</li>\n<li>CV Strategy: <code>StratifiedGroupKFold(k=10)</code></li>\n<li>Optuna tuning -&gt; Best OOF on all folds</li>\n<li><code>SEED = 555</code></li>\n<li>Ensemble of Best Folds: <code>[3,6,7,9]</code> (based on OOF score)</li>\n<li><code>QUANTIZE_MODEL = True</code></li>\n<li><code>QUANT_METHOD = \"float16\"</code></li>\n<li>No Augmentations</li>\n</ul>\n<p>Below is the complete model conf:</p>\n<pre><code>\n  \n   \n   \n   \n  \n   \n  \n   \n\n  \n   \n   \n   \n   \n  \n   \n   \n   \n   \n   \n</code></pre>\n<p><strong>Model 2:</strong></p>\n<ul>\n<li>Training <code>epochs = 100</code></li>\n<li>Add LayerNorm</li>\n<li>Single Split -&gt; Optuna Tuning -&gt; Use best hyper params to train on entire train data</li>\n<li><code>SEED = 555</code></li>\n<li><code>QUANTIZE_MODEL = True</code></li>\n<li><code>QUANT_METHOD = \"float16\"</code></li>\n<li>Flip Augmentations (<code>aug_prob=0.4</code>)</li>\n</ul>\n<p>Below is the complete model conf:</p>\n<pre><code>\n  \n   \n   \n   \n  \n   \n  \n   \n\n  \n   \n   \n   \n   \n  \n   \n   \n   \n   \n   \n</code></pre>\n<p><strong>Submission:</strong><br>\nThe submission is a simple ensemble of 5 models </p>\n<pre><code>Model 1 (folds=[3,6,7,9]) + Model 2\n\nPrivate Score: 0.82815\n\nPublic Score: 0.73896\n</code></pre>\n<p>Ensembling code can be found below:</p>\n<pre><code> (tf.Module):\n     ():\n        (TFLiteEnsembleModel, self).__init__()\n\n        \n        self.preprocess_layer = preprocess_layer\n        self.models = models\n\n        \n        self.weights = weights  [/(models)] * (models)\n        self.weights = tf.reshape(self.weights, (-, ))\n\n\n     ():\n        \n        x, non_empty_frame_idxs = self.proprocess_inputs(inputs)\n\n        outputs = []\n        \n         _model  self.models:\n            output = _model({ : x, : non_empty_frame_idxs })\n            outputs.append(output)\n\n        \n        outputs = tf.concat(outputs, axis=)\n        weighted_outputs = outputs * self.weights\n        weighted_sum = tf.reduce_sum(weighted_outputs, axis=, keepdims=)\n\n        \n        outputs = tf.squeeze(weighted_sum, axis=)\n\n        \n         {: outputs}\n\n     ():\n        \n        x, non_empty_frame_idxs = self.preprocess_layer(inputs)\n        \n        x = tf.expand_dims(x, axis=)\n        non_empty_frame_idxs = tf.expand_dims(non_empty_frame_idxs, axis=)\n\n         x, non_empty_frame_idxs\n</code></pre>",
      "rawMarkdown": "> **Many congratulations to all the winners in this competition. Huge thanks to all the discussion topics, public notebooks and datasets - most (if not all of) of my learning came from them.**\n\nFirst I would like to thank the organizers of this wonderful competition. \n\nTo say that I learned a lot by competing in this competition would really be an understatement. I wanted to take on the challenge of using Tensorflow in this competition and boy did it test me. If there was one thing that gave me hope in this competition it was @markwijkhuizen's amazing public [notebook](https://www.kaggle.com/code/markwijkhuizen/gislr-tf-data-processing-transformer-training). It was the foundation to my entire work in this competition.\n\nInitially I was using my own Transformer architectures in this competition which were not bad themselves but as @tatamikenn pointed it, the public notebooks had really good Transformer architectures, with a few tweaks and code cleanups I was able to extract good results both on CV and LB, this gave me a significant increase in CV and LB right on my first experiment. \n\n> I have open sourced the code base on GitHub -> https://github.com/abhinand5/isolated-sign-language-recognition. \n> \n> Leave a :star: if you like my work.\n\n### What worked for me?\n- Flip augmentation (`aug_prob = 0.4`)\n- 10 Fold Cross Validation\n- `label_smoothing = 0.3`\n- `embed_dim = 256`\n- `input_size = 32`\n- `num_transformer_blocks = 2` (Using 1 was under-fitting and 3 was over fitting even after several experiments of trying to tune)\n- Attention heads: `MHA_HEADS = 4` and `MHA_HEADS = 8`\n- Train on the entire dataset before submission\n- FP16 Quantization\n- GeLU activation\n\n### What did NOT work for me?\n- `input_size = 64`\n- Knowledge Distillation (wasted a lot of time in this)\n- Rotate augmentations (degraded CV and LB)\n- Adding Gaussian noise to the inputs (degraded CV and LB)\n- Removing `LayerNorm` from Transformer blocks (although the arch is shallow it did not work for me)\n- Strip Pruning \n- Dynamic Range Quantization (increased the inference time)\n- 2D Affine transforms\n\n### Final Solution:\nThere were two different model architectures in the final solution. \n\n**Dataset:** Only Competition Data\n\n**Model 1:**\n- Training `epochs = 50`\n- Add LayerNorm\n- CV Strategy: `StratifiedGroupKFold(k=10)`\n- Optuna tuning -> Best OOF on all folds\n- `SEED = 555`\n- Ensemble of Best Folds: `[3,6,7,9]` (based on OOF score)\n- `QUANTIZE_MODEL = True`\n- `QUANT_METHOD = \"float16\"`\n- No Augmentations\n\nBelow is the complete model conf:\n\n```yaml\nmodel:\n  # Dense layer units for landmarks\n  LIPS_UNITS: 256\n  HANDS_UNITS: 256\n  POSE_UNITS: 256\n  # Num attention heads\n  MHA_HEADS: 8\n  # final embedding and transformer embedding size\n  UNITS: 256\n\n  # Transformer\n  NUM_BLOCKS: 2\n  MLP_RATIO: 2\n  ADD_LAYER_NORM: True\n  LAYER_NORM_EPS: 1.0e-6\n  # Dropout\n  EMBEDDING_DROPOUT: 0.00\n  MLP_DROPOUT_RATIO: 0.10\n  CLASSIFIER_DROPOUT_RATIO: 0.40\n  LABEL_SMOOTHING: 0.3\n  ACTIVATION_FN: gelu\n```\n\n**Model 2:**\n- Training `epochs = 100`\n- Add LayerNorm\n- Single Split -> Optuna Tuning -> Use best hyper params to train on entire train data\n- `SEED = 555`\n- `QUANTIZE_MODEL = True`\n- `QUANT_METHOD = \"float16\"`\n- Flip Augmentations (`aug_prob=0.4`)\n\nBelow is the complete model conf:\n\n```yaml\nmodel:\n  # Dense layer units for landmarks\n  LIPS_UNITS: 256\n  HANDS_UNITS: 256\n  POSE_UNITS: 256\n  # Num attention heads\n  MHA_HEADS: 4\n  # final embedding and transformer embedding size\n  UNITS: 512\n\n  # Transformer\n  NUM_BLOCKS: 2\n  MLP_RATIO: 2\n  ADD_LAYER_NORM: True\n  LAYER_NORM_EPS: 1.0e-6\n  # Dropout\n  EMBEDDING_DROPOUT: 0.00\n  MLP_DROPOUT_RATIO: 0.10\n  CLASSIFIER_DROPOUT_RATIO: 0.30\n  LABEL_SMOOTHING: 0.10\n  ACTIVATION_FN: gelu\n```\n\n**Submission:**\nThe submission is a simple ensemble of 5 models \n\n```\nModel 1 (folds=[3,6,7,9]) + Model 2\n\nPrivate Score: 0.82815\n\nPublic Score: 0.73896\n```\n\nEnsembling code can be found below:\n\n```python\nclass TFLiteEnsembleModel(tf.Module):\n    def __init__(self, models, preprocess_layer, weights=None):\n        super(TFLiteEnsembleModel, self).__init__()\n\n        # Load the feature generation and main models\n        self.preprocess_layer = preprocess_layer\n        self.models = models\n        \n        # Set weights\n        self.weights = weights or [1.0/len(models)] * len(models)\n        self.weights = tf.reshape(self.weights, (-1, 1))\n    \n    @tf.function(input_signature=[tf.TensorSpec(shape=[None, N_ROWS, N_DIMS], dtype=tf.float32, name='inputs')])\n    def __call__(self, inputs):\n        # Preprocess the inputs\n        x, non_empty_frame_idxs = self.proprocess_inputs(inputs)\n        \n        outputs = []\n        # Make Prediction\n        for _model in self.models:\n            output = _model({ 'frames': x, 'non_empty_frame_idxs': non_empty_frame_idxs })\n            outputs.append(output)\n            \n        # Weighted Average\n        outputs = tf.concat(outputs, axis=0)\n        weighted_outputs = outputs * self.weights\n        weighted_sum = tf.reduce_sum(weighted_outputs, axis=0, keepdims=True)\n\n        # Squeeze Output 1x250 -> 250\n        outputs = tf.squeeze(weighted_sum, axis=0)\n        \n        # Return a dictionary with the output tensor\n        return {'outputs': outputs}\n    \n    def proprocess_inputs(self, inputs):\n        # Preprocess Data\n        x, non_empty_frame_idxs = self.preprocess_layer(inputs)\n        # Add Batch Dimension\n        x = tf.expand_dims(x, axis=0)\n        non_empty_frame_idxs = tf.expand_dims(non_empty_frame_idxs, axis=0)\n        \n        return x, non_empty_frame_idxs\n```",
      "votes": null
    },
    {
      "id": "2244051",
      "postDate": "05/03/2023 11:54:10",
      "content": "<p>well explained</p>",
      "rawMarkdown": "well explained",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2244051,
      "author_name": "chetan8007",
      "author_url": "",
      "post_date": "05/03/2023 11:54:10",
      "content": "<p>well explained</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2242131": "> **Many congratulations to all the winners in this competition. Huge thanks to all the discussion topics, public notebooks and datasets - most (if not all of) of my learning came from them.**\n\nFirst I would like to thank the organizers of this wonderful competition. \n\nTo say that I learned a lot by competing in this competition would really be an understatement. I wanted to take on the challenge of using Tensorflow in this competition and boy did it test me. If there was one thing that gave me hope in this competition it was @markwijkhuizen's amazing public [notebook](https://www.kaggle.com/code/markwijkhuizen/gislr-tf-data-processing-transformer-training). It was the foundation to my entire work in this competition.\n\nInitially I was using my own Transformer architectures in this competition which were not bad themselves but as @tatamikenn pointed it, the public notebooks had really good Transformer architectures, with a few tweaks and code cleanups I was able to extract good results both on CV and LB, this gave me a significant increase in CV and LB right on my first experiment. \n\n> I have open sourced the code base on GitHub -> https://github.com/abhinand5/isolated-sign-language-recognition. \n> \n> Leave a :star: if you like my work.\n\n### What worked for me?\n- Flip augmentation (`aug_prob = 0.4`)\n- 10 Fold Cross Validation\n- `label_smoothing = 0.3`\n- `embed_dim = 256`\n- `input_size = 32`\n- `num_transformer_blocks = 2` (Using 1 was under-fitting and 3 was over fitting even after several experiments of trying to tune)\n- Attention heads: `MHA_HEADS = 4` and `MHA_HEADS = 8`\n- Train on the entire dataset before submission\n- FP16 Quantization\n- GeLU activation\n\n### What did NOT work for me?\n- `input_size = 64`\n- Knowledge Distillation (wasted a lot of time in this)\n- Rotate augmentations (degraded CV and LB)\n- Adding Gaussian noise to the inputs (degraded CV and LB)\n- Removing `LayerNorm` from Transformer blocks (although the arch is shallow it did not work for me)\n- Strip Pruning \n- Dynamic Range Quantization (increased the inference time)\n- 2D Affine transforms\n\n### Final Solution:\nThere were two different model architectures in the final solution. \n\n**Dataset:** Only Competition Data\n\n**Model 1:**\n- Training `epochs = 50`\n- Add LayerNorm\n- CV Strategy: `StratifiedGroupKFold(k=10)`\n- Optuna tuning -> Best OOF on all folds\n- `SEED = 555`\n- Ensemble of Best Folds: `[3,6,7,9]` (based on OOF score)\n- `QUANTIZE_MODEL = True`\n- `QUANT_METHOD = \"float16\"`\n- No Augmentations\n\nBelow is the complete model conf:\n\n```yaml\nmodel:\n  # Dense layer units for landmarks\n  LIPS_UNITS: 256\n  HANDS_UNITS: 256\n  POSE_UNITS: 256\n  # Num attention heads\n  MHA_HEADS: 8\n  # final embedding and transformer embedding size\n  UNITS: 256\n\n  # Transformer\n  NUM_BLOCKS: 2\n  MLP_RATIO: 2\n  ADD_LAYER_NORM: True\n  LAYER_NORM_EPS: 1.0e-6\n  # Dropout\n  EMBEDDING_DROPOUT: 0.00\n  MLP_DROPOUT_RATIO: 0.10\n  CLASSIFIER_DROPOUT_RATIO: 0.40\n  LABEL_SMOOTHING: 0.3\n  ACTIVATION_FN: gelu\n```\n\n**Model 2:**\n- Training `epochs = 100`\n- Add LayerNorm\n- Single Split -> Optuna Tuning -> Use best hyper params to train on entire train data\n- `SEED = 555`\n- `QUANTIZE_MODEL = True`\n- `QUANT_METHOD = \"float16\"`\n- Flip Augmentations (`aug_prob=0.4`)\n\nBelow is the complete model conf:\n\n```yaml\nmodel:\n  # Dense layer units for landmarks\n  LIPS_UNITS: 256\n  HANDS_UNITS: 256\n  POSE_UNITS: 256\n  # Num attention heads\n  MHA_HEADS: 4\n  # final embedding and transformer embedding size\n  UNITS: 512\n\n  # Transformer\n  NUM_BLOCKS: 2\n  MLP_RATIO: 2\n  ADD_LAYER_NORM: True\n  LAYER_NORM_EPS: 1.0e-6\n  # Dropout\n  EMBEDDING_DROPOUT: 0.00\n  MLP_DROPOUT_RATIO: 0.10\n  CLASSIFIER_DROPOUT_RATIO: 0.30\n  LABEL_SMOOTHING: 0.10\n  ACTIVATION_FN: gelu\n```\n\n**Submission:**\nThe submission is a simple ensemble of 5 models \n\n```\nModel 1 (folds=[3,6,7,9]) + Model 2\n\nPrivate Score: 0.82815\n\nPublic Score: 0.73896\n```\n\nEnsembling code can be found below:\n\n```python\nclass TFLiteEnsembleModel(tf.Module):\n    def __init__(self, models, preprocess_layer, weights=None):\n        super(TFLiteEnsembleModel, self).__init__()\n\n        # Load the feature generation and main models\n        self.preprocess_layer = preprocess_layer\n        self.models = models\n        \n        # Set weights\n        self.weights = weights or [1.0/len(models)] * len(models)\n        self.weights = tf.reshape(self.weights, (-1, 1))\n    \n    @tf.function(input_signature=[tf.TensorSpec(shape=[None, N_ROWS, N_DIMS], dtype=tf.float32, name='inputs')])\n    def __call__(self, inputs):\n        # Preprocess the inputs\n        x, non_empty_frame_idxs = self.proprocess_inputs(inputs)\n        \n        outputs = []\n        # Make Prediction\n        for _model in self.models:\n            output = _model({ 'frames': x, 'non_empty_frame_idxs': non_empty_frame_idxs })\n            outputs.append(output)\n            \n        # Weighted Average\n        outputs = tf.concat(outputs, axis=0)\n        weighted_outputs = outputs * self.weights\n        weighted_sum = tf.reduce_sum(weighted_outputs, axis=0, keepdims=True)\n\n        # Squeeze Output 1x250 -> 250\n        outputs = tf.squeeze(weighted_sum, axis=0)\n        \n        # Return a dictionary with the output tensor\n        return {'outputs': outputs}\n    \n    def proprocess_inputs(self, inputs):\n        # Preprocess Data\n        x, non_empty_frame_idxs = self.preprocess_layer(inputs)\n        # Add Batch Dimension\n        x = tf.expand_dims(x, axis=0)\n        non_empty_frame_idxs = tf.expand_dims(non_empty_frame_idxs, axis=0)\n        \n        return x, non_empty_frame_idxs\n```",
    "2244051": "well explained"
  },
  "source": "meta"
}