{
  "id": 393655,
  "title": "Normalize or Not",
  "url": "/competitions/asl-signs/discussion/393655",
  "author_name": "",
  "post_date": "2023-03-10T10:13:52.996195900Z",
  "votes": 2,
  "comment_count": 6,
  "views": 0,
  "content": "<p>When I normalize the data, the model starts to overfit.</p>\n<pre><code> ():\n     tf.reduce_sum(tf.where(tf.math.is_nan(x), tf.zeros_like(x), x), axis=axis) / tf.reduce_sum(tf.where(tf.math.is_nan(x), tf.zeros_like(x), tf.ones_like(x)), axis=axis)\n\n\n\n ():\n    d = x - tf_nan_mean(x, axis=axis)\n     tf.math.sqrt(tf_nan_mean(d * d, axis=axis))\n\n\n ():\n    \n    x_mean = tf_nan_mean(x, axis=)\n    x_std  = tf_nan_std(x,  axis=)\n\n    x_out = tf.concat([x_mean, x_std], axis=)\n    x_out = tf.reshape(x_out, (, INPUT_SHAPE[]*))\n    x_out = tf.where(tf.math.is_finite(x_out), x_out, tf.zeros_like(x_out))\n     x_out\n\n\n ():\n    \n    x_mean = tf_nan_mean(x, axis=)\n    x_std  = tf_nan_std(x,  axis=)\n      (x - x_mean) / x_std\n\n\n\nx = tf.gather(x_in, point_landmarks, axis=)\n\nx = normalize_data(x)\n\nx_list = []\n\nx = tf.image.resize(tf.where(tf.math.is_finite(x), x, tf_nan_mean(x, axis=)), [NUM_FRAMES, LANDMARKS])\nx = tf.reshape(x, (, INPUT_SHAPE[]*INPUT_SHAPE[]))\nx = tf.where(tf.math.is_nan(x), tf.zeros_like(x), x)\n\nx_list.append(x)\nx = tf.concat(x_list, axis=)\n</code></pre>\n<p>Did anyone get benefit from doing normalization?</p>\n<p>w/o Normalization:  CV-0.75+<br>\nwith Normalization: CV-0.46+</p>",
  "messages": [
    {
      "id": "2176004",
      "postDate": "03/10/2023 10:13:52",
      "content": "<p>When I normalize the data, the model starts to overfit.</p>\n<pre><code> ():\n     tf.reduce_sum(tf.where(tf.math.is_nan(x), tf.zeros_like(x), x), axis=axis) / tf.reduce_sum(tf.where(tf.math.is_nan(x), tf.zeros_like(x), tf.ones_like(x)), axis=axis)\n\n\n\n ():\n    d = x - tf_nan_mean(x, axis=axis)\n     tf.math.sqrt(tf_nan_mean(d * d, axis=axis))\n\n\n ():\n    \n    x_mean = tf_nan_mean(x, axis=)\n    x_std  = tf_nan_std(x,  axis=)\n\n    x_out = tf.concat([x_mean, x_std], axis=)\n    x_out = tf.reshape(x_out, (, INPUT_SHAPE[]*))\n    x_out = tf.where(tf.math.is_finite(x_out), x_out, tf.zeros_like(x_out))\n     x_out\n\n\n ():\n    \n    x_mean = tf_nan_mean(x, axis=)\n    x_std  = tf_nan_std(x,  axis=)\n      (x - x_mean) / x_std\n\n\n\nx = tf.gather(x_in, point_landmarks, axis=)\n\nx = normalize_data(x)\n\nx_list = []\n\nx = tf.image.resize(tf.where(tf.math.is_finite(x), x, tf_nan_mean(x, axis=)), [NUM_FRAMES, LANDMARKS])\nx = tf.reshape(x, (, INPUT_SHAPE[]*INPUT_SHAPE[]))\nx = tf.where(tf.math.is_nan(x), tf.zeros_like(x), x)\n\nx_list.append(x)\nx = tf.concat(x_list, axis=)\n</code></pre>\n<p>Did anyone get benefit from doing normalization?</p>\n<p>w/o Normalization:  CV-0.75+<br>\nwith Normalization: CV-0.46+</p>",
      "rawMarkdown": "When I normalize the data, the model starts to overfit.\n\n\n```python\ndef tf_nan_mean(x, axis=0):\n    return tf.reduce_sum(tf.where(tf.math.is_nan(x), tf.zeros_like(x), x), axis=axis) / tf.reduce_sum(tf.where(tf.math.is_nan(x), tf.zeros_like(x), tf.ones_like(x)), axis=axis)\n\n\n\ndef tf_nan_std(x, axis=0):\n    d = x - tf_nan_mean(x, axis=axis)\n    return tf.math.sqrt(tf_nan_mean(d * d, axis=axis))\n\n\ndef flatten_means_and_stds(x, axis=0):\n    # Get means and stds\n    x_mean = tf_nan_mean(x, axis=0)\n    x_std  = tf_nan_std(x,  axis=0)\n\n    x_out = tf.concat([x_mean, x_std], axis=0)\n    x_out = tf.reshape(x_out, (1, INPUT_SHAPE[1]*2))\n    x_out = tf.where(tf.math.is_finite(x_out), x_out, tf.zeros_like(x_out))\n    return x_out\n\n\ndef normalize_data(x, axis=0):\n    # Get means and stds\n    x_mean = tf_nan_mean(x, axis=0)\n    x_std  = tf_nan_std(x,  axis=0)\n    return  (x - x_mean) / x_std\n\n\n## Data Normalization \nx = tf.gather(x_in, point_landmarks, axis=1)\n# Normalization    \nx = normalize_data(x)\n\nx_list = []\n## Resize only dimension 0. Resize can't handle nan, so replace nan with that dimension's avg value to reduce impact.\nx = tf.image.resize(tf.where(tf.math.is_finite(x), x, tf_nan_mean(x, axis=0)), [NUM_FRAMES, LANDMARKS])\nx = tf.reshape(x, (1, INPUT_SHAPE[0]*INPUT_SHAPE[1]))\nx = tf.where(tf.math.is_nan(x), tf.zeros_like(x), x)\n        \nx_list.append(x)\nx = tf.concat(x_list, axis=1)\n```\n\nDid anyone get benefit from doing normalization?\n\nw/o Normalization:  CV-0.75+\nwith Normalization: CV-0.46+",
      "votes": null
    },
    {
      "id": "2176316",
      "postDate": "03/10/2023 14:52:09",
      "content": "<p>deleted. i am sorry that my previous comment were wrong</p>",
      "rawMarkdown": "deleted. i am sorry that my previous comment were wrong",
      "votes": null
    },
    {
      "id": "2176510",
      "postDate": "03/10/2023 17:38:02",
      "content": "<p>Although I haven't tried exactly that, I've tried a different normalization variation and had a lot of trouble with it.</p>\n<p>But your problem might be simpler?</p>\n<p>Darien's <a href=\"https://www.kaggle.com/code/dschettler8845/gislr-how-to-ensemble\" target=\"_blank\">How to Ensemble notebook</a> uses normalization, and does it for each feature individually (axis=0), but across the entire dataset, not just a single row entry across all frames.</p>\n<p>Try doing all preprocessing, and then normalizing at the end?</p>\n<p>It looks like you are doing it for a single entry individually, across all frames. Among other concerns with that method:</p>\n<ul>\n<li>normalization axis=0 is very different from, say, normalization axis=None.</li>\n<li>Normalization axis=0 means that the point cloud is no longer using unified coordinates. For example, finger landmark touching lip landmark is no longer approximately finger landmark(x,y,z) == lip landmark(x,y,z). Though if using all data, it might still be possible for the model to 'know' when finger touches lips. Not if using normalization separately for each individual.</li>\n<li>Along the same lines as above, what if in a right handed sign, 5% of the time the media pipe mistakes it for the left hand? Before, it would've been right hand thumb frames 1,2,3 -&gt; 0.433, nan, 0.432, left nan, 0.431, nan. Model, theoretically, could 'learn' how to use left hand as stand in for right hand. After norm, though, maybe more like right -1.0, nan, -1.01, left nan, 0.0, nan? Left hand only had a single frame, so everything is guaranteed to be set to 0, basically erasing that frame's data.</li>\n<li>Normalization axis=0 could have a huge impact on the model's ability to handle nan as a special case, depending on many other factors. After normalization, you could try setting nan to some non-zero value like -3 or something and see if it helps.</li>\n</ul>\n<p>Another thing, if using Darien's method, is to ensure you use the EXACT same normalization for inference. Checking… yep, Darien does it correctly:</p>\n<pre><code> (tf.keras.layers.Layer):\n     ():\n</code></pre>\n<p>See that he passes the exact mean and std values he used for train into his test data preprocessing layer directly.</p>",
      "rawMarkdown": "Although I haven't tried exactly that, I've tried a different normalization variation and had a lot of trouble with it.\n\nBut your problem might be simpler?\n\nDarien's [How to Ensemble notebook](https://www.kaggle.com/code/dschettler8845/gislr-how-to-ensemble) uses normalization, and does it for each feature individually (axis=0), but across the entire dataset, not just a single row entry across all frames.\n\nTry doing all preprocessing, and then normalizing at the end?\n\nIt looks like you are doing it for a single entry individually, across all frames. Among other concerns with that method:\n* normalization axis=0 is very different from, say, normalization axis=None.\n* Normalization axis=0 means that the point cloud is no longer using unified coordinates. For example, finger landmark touching lip landmark is no longer approximately finger landmark(x,y,z) == lip landmark(x,y,z). Though if using all data, it might still be possible for the model to 'know' when finger touches lips. Not if using normalization separately for each individual.\n* Along the same lines as above, what if in a right handed sign, 5% of the time the media pipe mistakes it for the left hand? Before, it would've been right hand thumb frames 1,2,3 -> 0.433, nan, 0.432, left nan, 0.431, nan. Model, theoretically, could 'learn' how to use left hand as stand in for right hand. After norm, though, maybe more like right -1.0, nan, -1.01, left nan, 0.0, nan? Left hand only had a single frame, so everything is guaranteed to be set to 0, basically erasing that frame's data.\n* Normalization axis=0 could have a huge impact on the model's ability to handle nan as a special case, depending on many other factors. After normalization, you could try setting nan to some non-zero value like -3 or something and see if it helps.\n\nAnother thing, if using Darien's method, is to ensure you use the EXACT same normalization for inference. Checking... yep, Darien does it correctly:\n\n```python\nclass PrepInputs(tf.keras.layers.Layer):\n    def __init__(self, lh_idx_range=(468, 489), pose_idx_range=(489, 522), rh_idx_range=(522, 543), distribution_mean=all_mean, distribution_std=all_std):\n```\nSee that he passes the exact mean and std values he used for train into his test data preprocessing layer directly.",
      "votes": null
    },
    {
      "id": "2176783",
      "postDate": "03/10/2023 22:35:28",
      "content": "<p>Mediapipe Landmark coordinates normalized to [0.0, 1.0] by the image width and height respectively. Maybe re-normalized data is the problem ? </p>",
      "rawMarkdown": "Mediapipe Landmark coordinates normalized to [0.0, 1.0] by the image width and height respectively. Maybe re-normalized data is the problem ?",
      "votes": null
    },
    {
      "id": "2176966",
      "postDate": "03/11/2023 04:58:40",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> for the detailed explanation.. </p>\n<p>Will try few methods and will post my experience later 😊</p>",
      "rawMarkdown": "Thanks @roberthatch for the detailed explanation.. \n\nWill try few methods and will post my experience later 😊",
      "votes": null
    },
    {
      "id": "2177124",
      "postDate": "03/11/2023 08:21:20",
      "content": "<p>this normalisation works for me.<br>\nthe effects is visualized in <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/391265#2174824\" target=\"_blank\">https://www.kaggle.com/competitions/asl-signs/discussion/391265#2174824</a></p>\n<pre><code>class InputNet(tf.keras.layers.Layer):\n    def __init__(self, ):\n        super(InputNet, self).__init__()\n        self.lip = tf.constant([\n            61, 185, 40, 39, 37, 0, 267, 269, 270, 409,\n            291, 146, 91, 181, 84, 17, 314, 405, 321, 375,\n            78, 191, 80, 81, 82, 13, 312, 311, 310, 415,\n            95, 88, 178, 87, 14, 317, 402, 318, 324, 308,\n        ])\n        self.lhand = (468, 489)\n        self.rhand = (522, 543)\n        self.max_length = CFG.max_length\n\n    def call(self, xyz):\n        #x = xyz\n\n        L = len(xyz)\n        if L &gt; self.max_length:\n            # xyz = xyz[:self.max_length] #first\n            # xyz = xyz[-self.max_length:] #last\n\n            i = (L-self.max_length)//2\n            xyz = xyz[i:i + self.max_length]  # center\n\n\n        not_nan_xyz = xyz[~tf.math.is_nan(xyz)]\n        xyz -= tf.math.reduce_mean (not_nan_xyz, axis=0, keepdims=True)  # noramlisation to common maen\n        xyz /= tf.math.reduce_std (not_nan_xyz, axis=0, keepdims=True)\n\n        x = tf.concat([\n            tf.gather(xyz, self.lip, axis=1),\n            xyz[:, self.lhand[0]:self.lhand[1]],\n            xyz[:, self.rhand[0]:self.rhand[1]],\n        ],1)\n        x = tf.where(tf.math.is_finite(x), x, tf.zeros_like(x))\n        return  x\n\n\nclass TFModel1(tf.Module):\n    def __init__(self,):\n        super(TFModel1, self).__init__()\n        self.input_net = InputNet()\n\n    @tf.function(input_signature=[tf.TensorSpec(shape=[None, 543, 3], dtype=tf.float32, name='inputs')])\n    def __call__(self, inputs):\n        outputs = self.input_net(inputs)\n        return outputs\n</code></pre>",
      "rawMarkdown": "this normalisation works for me.\nthe effects is visualized in https://www.kaggle.com/competitions/asl-signs/discussion/391265#2174824\n\n```\nclass InputNet(tf.keras.layers.Layer):\n\tdef __init__(self, ):\n\t\tsuper(InputNet, self).__init__()\n\t\tself.lip = tf.constant([\n\t\t\t61, 185, 40, 39, 37, 0, 267, 269, 270, 409,\n\t\t\t291, 146, 91, 181, 84, 17, 314, 405, 321, 375,\n\t\t\t78, 191, 80, 81, 82, 13, 312, 311, 310, 415,\n\t\t\t95, 88, 178, 87, 14, 317, 402, 318, 324, 308,\n\t\t])\n\t\tself.lhand = (468, 489)\n\t\tself.rhand = (522, 543)\n\t\tself.max_length = CFG.max_length\n\n\tdef call(self, xyz):\n\t\t#x = xyz\n\n\t\tL = len(xyz)\n\t\tif L > self.max_length:\n\t\t\t# xyz = xyz[:self.max_length] #first\n\t\t\t# xyz = xyz[-self.max_length:] #last\n\n\t\t\ti = (L-self.max_length)//2\n\t\t\txyz = xyz[i:i + self.max_length]  # center\n\n\n\t\tnot_nan_xyz = xyz[~tf.math.is_nan(xyz)]\n\t\txyz -= tf.math.reduce_mean (not_nan_xyz, axis=0, keepdims=True)  # noramlisation to common maen\n\t\txyz /= tf.math.reduce_std (not_nan_xyz, axis=0, keepdims=True)\n\n\t\tx = tf.concat([\n\t\t\ttf.gather(xyz, self.lip, axis=1),\n\t\t\txyz[:, self.lhand[0]:self.lhand[1]],\n\t\t\txyz[:, self.rhand[0]:self.rhand[1]],\n\t\t],1)\n\t\tx = tf.where(tf.math.is_finite(x), x, tf.zeros_like(x))\n\t\treturn  x\n\n\nclass TFModel1(tf.Module):\n\tdef __init__(self,):\n\t\tsuper(TFModel1, self).__init__()\n\t\tself.input_net = InputNet()\n\n\t@tf.function(input_signature=[tf.TensorSpec(shape=[None, 543, 3], dtype=tf.float32, name='inputs')])\n\tdef __call__(self, inputs):\n\t\toutputs = self.input_net(inputs)\n\t\treturn outputs\n\n```",
      "votes": null
    },
    {
      "id": "2177268",
      "postDate": "03/11/2023 10:33:48",
      "content": "<p>Thanks a lot for sharing <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> </p>",
      "rawMarkdown": "Thanks a lot for sharing @hengck23",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2176316,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/10/2023 14:52:09",
      "content": "<p>deleted. i am sorry that my previous comment were wrong</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2176510,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "03/10/2023 17:38:02",
      "content": "<p>Although I haven't tried exactly that, I've tried a different normalization variation and had a lot of trouble with it.</p>\n<p>But your problem might be simpler?</p>\n<p>Darien's <a href=\"https://www.kaggle.com/code/dschettler8845/gislr-how-to-ensemble\" target=\"_blank\">How to Ensemble notebook</a> uses normalization, and does it for each feature individually (axis=0), but across the entire dataset, not just a single row entry across all frames.</p>\n<p>Try doing all preprocessing, and then normalizing at the end?</p>\n<p>It looks like you are doing it for a single entry individually, across all frames. Among other concerns with that method:</p>\n<ul>\n<li>normalization axis=0 is very different from, say, normalization axis=None.</li>\n<li>Normalization axis=0 means that the point cloud is no longer using unified coordinates. For example, finger landmark touching lip landmark is no longer approximately finger landmark(x,y,z) == lip landmark(x,y,z). Though if using all data, it might still be possible for the model to 'know' when finger touches lips. Not if using normalization separately for each individual.</li>\n<li>Along the same lines as above, what if in a right handed sign, 5% of the time the media pipe mistakes it for the left hand? Before, it would've been right hand thumb frames 1,2,3 -&gt; 0.433, nan, 0.432, left nan, 0.431, nan. Model, theoretically, could 'learn' how to use left hand as stand in for right hand. After norm, though, maybe more like right -1.0, nan, -1.01, left nan, 0.0, nan? Left hand only had a single frame, so everything is guaranteed to be set to 0, basically erasing that frame's data.</li>\n<li>Normalization axis=0 could have a huge impact on the model's ability to handle nan as a special case, depending on many other factors. After normalization, you could try setting nan to some non-zero value like -3 or something and see if it helps.</li>\n</ul>\n<p>Another thing, if using Darien's method, is to ensure you use the EXACT same normalization for inference. Checking… yep, Darien does it correctly:</p>\n<pre><code> (tf.keras.layers.Layer):\n     ():\n</code></pre>\n<p>See that he passes the exact mean and std values he used for train into his test data preprocessing layer directly.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2176966,
          "author_name": "dhakshiin1601",
          "author_url": "",
          "post_date": "03/11/2023 04:58:40",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> for the detailed explanation.. </p>\n<p>Will try few methods and will post my experience later 😊</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2176783,
      "author_name": "jahsylla",
      "author_url": "",
      "post_date": "03/10/2023 22:35:28",
      "content": "<p>Mediapipe Landmark coordinates normalized to [0.0, 1.0] by the image width and height respectively. Maybe re-normalized data is the problem ? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2177124,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/11/2023 08:21:20",
      "content": "<p>this normalisation works for me.<br>\nthe effects is visualized in <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/391265#2174824\" target=\"_blank\">https://www.kaggle.com/competitions/asl-signs/discussion/391265#2174824</a></p>\n<pre><code>class InputNet(tf.keras.layers.Layer):\n    def __init__(self, ):\n        super(InputNet, self).__init__()\n        self.lip = tf.constant([\n            61, 185, 40, 39, 37, 0, 267, 269, 270, 409,\n            291, 146, 91, 181, 84, 17, 314, 405, 321, 375,\n            78, 191, 80, 81, 82, 13, 312, 311, 310, 415,\n            95, 88, 178, 87, 14, 317, 402, 318, 324, 308,\n        ])\n        self.lhand = (468, 489)\n        self.rhand = (522, 543)\n        self.max_length = CFG.max_length\n\n    def call(self, xyz):\n        #x = xyz\n\n        L = len(xyz)\n        if L &gt; self.max_length:\n            # xyz = xyz[:self.max_length] #first\n            # xyz = xyz[-self.max_length:] #last\n\n            i = (L-self.max_length)//2\n            xyz = xyz[i:i + self.max_length]  # center\n\n\n        not_nan_xyz = xyz[~tf.math.is_nan(xyz)]\n        xyz -= tf.math.reduce_mean (not_nan_xyz, axis=0, keepdims=True)  # noramlisation to common maen\n        xyz /= tf.math.reduce_std (not_nan_xyz, axis=0, keepdims=True)\n\n        x = tf.concat([\n            tf.gather(xyz, self.lip, axis=1),\n            xyz[:, self.lhand[0]:self.lhand[1]],\n            xyz[:, self.rhand[0]:self.rhand[1]],\n        ],1)\n        x = tf.where(tf.math.is_finite(x), x, tf.zeros_like(x))\n        return  x\n\n\nclass TFModel1(tf.Module):\n    def __init__(self,):\n        super(TFModel1, self).__init__()\n        self.input_net = InputNet()\n\n    @tf.function(input_signature=[tf.TensorSpec(shape=[None, 543, 3], dtype=tf.float32, name='inputs')])\n    def __call__(self, inputs):\n        outputs = self.input_net(inputs)\n        return outputs\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 2177268,
          "author_name": "dhakshiin1601",
          "author_url": "",
          "post_date": "03/11/2023 10:33:48",
          "content": "<p>Thanks a lot for sharing <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2176004": "When I normalize the data, the model starts to overfit.\n\n\n```python\ndef tf_nan_mean(x, axis=0):\n    return tf.reduce_sum(tf.where(tf.math.is_nan(x), tf.zeros_like(x), x), axis=axis) / tf.reduce_sum(tf.where(tf.math.is_nan(x), tf.zeros_like(x), tf.ones_like(x)), axis=axis)\n\n\n\ndef tf_nan_std(x, axis=0):\n    d = x - tf_nan_mean(x, axis=axis)\n    return tf.math.sqrt(tf_nan_mean(d * d, axis=axis))\n\n\ndef flatten_means_and_stds(x, axis=0):\n    # Get means and stds\n    x_mean = tf_nan_mean(x, axis=0)\n    x_std  = tf_nan_std(x,  axis=0)\n\n    x_out = tf.concat([x_mean, x_std], axis=0)\n    x_out = tf.reshape(x_out, (1, INPUT_SHAPE[1]*2))\n    x_out = tf.where(tf.math.is_finite(x_out), x_out, tf.zeros_like(x_out))\n    return x_out\n\n\ndef normalize_data(x, axis=0):\n    # Get means and stds\n    x_mean = tf_nan_mean(x, axis=0)\n    x_std  = tf_nan_std(x,  axis=0)\n    return  (x - x_mean) / x_std\n\n\n## Data Normalization \nx = tf.gather(x_in, point_landmarks, axis=1)\n# Normalization    \nx = normalize_data(x)\n\nx_list = []\n## Resize only dimension 0. Resize can't handle nan, so replace nan with that dimension's avg value to reduce impact.\nx = tf.image.resize(tf.where(tf.math.is_finite(x), x, tf_nan_mean(x, axis=0)), [NUM_FRAMES, LANDMARKS])\nx = tf.reshape(x, (1, INPUT_SHAPE[0]*INPUT_SHAPE[1]))\nx = tf.where(tf.math.is_nan(x), tf.zeros_like(x), x)\n        \nx_list.append(x)\nx = tf.concat(x_list, axis=1)\n```\n\nDid anyone get benefit from doing normalization?\n\nw/o Normalization:  CV-0.75+\nwith Normalization: CV-0.46+",
    "2176316": "deleted. i am sorry that my previous comment were wrong",
    "2176510": "Although I haven't tried exactly that, I've tried a different normalization variation and had a lot of trouble with it.\n\nBut your problem might be simpler?\n\nDarien's [How to Ensemble notebook](https://www.kaggle.com/code/dschettler8845/gislr-how-to-ensemble) uses normalization, and does it for each feature individually (axis=0), but across the entire dataset, not just a single row entry across all frames.\n\nTry doing all preprocessing, and then normalizing at the end?\n\nIt looks like you are doing it for a single entry individually, across all frames. Among other concerns with that method:\n* normalization axis=0 is very different from, say, normalization axis=None.\n* Normalization axis=0 means that the point cloud is no longer using unified coordinates. For example, finger landmark touching lip landmark is no longer approximately finger landmark(x,y,z) == lip landmark(x,y,z). Though if using all data, it might still be possible for the model to 'know' when finger touches lips. Not if using normalization separately for each individual.\n* Along the same lines as above, what if in a right handed sign, 5% of the time the media pipe mistakes it for the left hand? Before, it would've been right hand thumb frames 1,2,3 -> 0.433, nan, 0.432, left nan, 0.431, nan. Model, theoretically, could 'learn' how to use left hand as stand in for right hand. After norm, though, maybe more like right -1.0, nan, -1.01, left nan, 0.0, nan? Left hand only had a single frame, so everything is guaranteed to be set to 0, basically erasing that frame's data.\n* Normalization axis=0 could have a huge impact on the model's ability to handle nan as a special case, depending on many other factors. After normalization, you could try setting nan to some non-zero value like -3 or something and see if it helps.\n\nAnother thing, if using Darien's method, is to ensure you use the EXACT same normalization for inference. Checking... yep, Darien does it correctly:\n\n```python\nclass PrepInputs(tf.keras.layers.Layer):\n    def __init__(self, lh_idx_range=(468, 489), pose_idx_range=(489, 522), rh_idx_range=(522, 543), distribution_mean=all_mean, distribution_std=all_std):\n```\nSee that he passes the exact mean and std values he used for train into his test data preprocessing layer directly.",
    "2176783": "Mediapipe Landmark coordinates normalized to [0.0, 1.0] by the image width and height respectively. Maybe re-normalized data is the problem ?",
    "2176966": "Thanks @roberthatch for the detailed explanation.. \n\nWill try few methods and will post my experience later 😊",
    "2177124": "this normalisation works for me.\nthe effects is visualized in https://www.kaggle.com/competitions/asl-signs/discussion/391265#2174824\n\n```\nclass InputNet(tf.keras.layers.Layer):\n\tdef __init__(self, ):\n\t\tsuper(InputNet, self).__init__()\n\t\tself.lip = tf.constant([\n\t\t\t61, 185, 40, 39, 37, 0, 267, 269, 270, 409,\n\t\t\t291, 146, 91, 181, 84, 17, 314, 405, 321, 375,\n\t\t\t78, 191, 80, 81, 82, 13, 312, 311, 310, 415,\n\t\t\t95, 88, 178, 87, 14, 317, 402, 318, 324, 308,\n\t\t])\n\t\tself.lhand = (468, 489)\n\t\tself.rhand = (522, 543)\n\t\tself.max_length = CFG.max_length\n\n\tdef call(self, xyz):\n\t\t#x = xyz\n\n\t\tL = len(xyz)\n\t\tif L > self.max_length:\n\t\t\t# xyz = xyz[:self.max_length] #first\n\t\t\t# xyz = xyz[-self.max_length:] #last\n\n\t\t\ti = (L-self.max_length)//2\n\t\t\txyz = xyz[i:i + self.max_length]  # center\n\n\n\t\tnot_nan_xyz = xyz[~tf.math.is_nan(xyz)]\n\t\txyz -= tf.math.reduce_mean (not_nan_xyz, axis=0, keepdims=True)  # noramlisation to common maen\n\t\txyz /= tf.math.reduce_std (not_nan_xyz, axis=0, keepdims=True)\n\n\t\tx = tf.concat([\n\t\t\ttf.gather(xyz, self.lip, axis=1),\n\t\t\txyz[:, self.lhand[0]:self.lhand[1]],\n\t\t\txyz[:, self.rhand[0]:self.rhand[1]],\n\t\t],1)\n\t\tx = tf.where(tf.math.is_finite(x), x, tf.zeros_like(x))\n\t\treturn  x\n\n\nclass TFModel1(tf.Module):\n\tdef __init__(self,):\n\t\tsuper(TFModel1, self).__init__()\n\t\tself.input_net = InputNet()\n\n\t@tf.function(input_signature=[tf.TensorSpec(shape=[None, 543, 3], dtype=tf.float32, name='inputs')])\n\tdef __call__(self, inputs):\n\t\toutputs = self.input_net(inputs)\n\t\treturn outputs\n\n```",
    "2177268": "Thanks a lot for sharing @hengck23"
  },
  "source": "meta"
}