{
  "id": 390702,
  "title": "Keras preprocessing layers documentation",
  "url": "/competitions/asl-signs/discussion/390702",
  "author_name": "",
  "post_date": "2023-02-26T20:33:58.626454500Z",
  "votes": 4,
  "comment_count": 10,
  "views": 0,
  "content": "<p><a href=\"https://www.tensorflow.org/guide/keras/preprocessing_layers\" target=\"_blank\">https://www.tensorflow.org/guide/keras/preprocessing_layers</a></p>\n<p>\"With Keras preprocessing layers, you can build and export models that are truly end-to-end: models that accept raw images or raw structured data as input; models that handle feature normalization or feature value indexing on their own.\"</p>\n<p>For better or for worse, we are REQUIRED to do exactly this in this competition. I am calling it a \"Model Competition\", it is NOT a \"Code Competition\", you are required to build a model that handles shape ([n],543,3) and produces the proper 250 output prediction. One successful <a href=\"https://www.kaggle.com/code/lonnieqin/isolated-sign-language-recognition-with-dnn/notebook\" target=\"_blank\">example is here</a>. Unlike a Code Competition, you don't get a chance to process the (n, 543, 3) directly, for example you can't even drop the z values as suggested in the dataset description! (\"The MediaPipe model is not fully trained to predict depth so you may wish to ignore the z values.\")</p>\n<p>So what do you do? Well, I don't know, but it's all in the above link, which I'll be reading. Hopefully a few example notebooks will start teaching some of the building blocks, as well.</p>\n<p>Note that TensorFlow Lite shouldn't be any FURTHER handicap. So far as I am aware, you can think of it as a simple compilation and/or compression step, any and all TensorFlow models can be converted to TensorFlow Lite.</p>",
  "messages": [
    {
      "id": "2160652",
      "postDate": "02/26/2023 20:33:58",
      "content": "<p><a href=\"https://www.tensorflow.org/guide/keras/preprocessing_layers\" target=\"_blank\">https://www.tensorflow.org/guide/keras/preprocessing_layers</a></p>\n<p>\"With Keras preprocessing layers, you can build and export models that are truly end-to-end: models that accept raw images or raw structured data as input; models that handle feature normalization or feature value indexing on their own.\"</p>\n<p>For better or for worse, we are REQUIRED to do exactly this in this competition. I am calling it a \"Model Competition\", it is NOT a \"Code Competition\", you are required to build a model that handles shape ([n],543,3) and produces the proper 250 output prediction. One successful <a href=\"https://www.kaggle.com/code/lonnieqin/isolated-sign-language-recognition-with-dnn/notebook\" target=\"_blank\">example is here</a>. Unlike a Code Competition, you don't get a chance to process the (n, 543, 3) directly, for example you can't even drop the z values as suggested in the dataset description! (\"The MediaPipe model is not fully trained to predict depth so you may wish to ignore the z values.\")</p>\n<p>So what do you do? Well, I don't know, but it's all in the above link, which I'll be reading. Hopefully a few example notebooks will start teaching some of the building blocks, as well.</p>\n<p>Note that TensorFlow Lite shouldn't be any FURTHER handicap. So far as I am aware, you can think of it as a simple compilation and/or compression step, any and all TensorFlow models can be converted to TensorFlow Lite.</p>",
      "rawMarkdown": "https://www.tensorflow.org/guide/keras/preprocessing_layers\n\n\"With Keras preprocessing layers, you can build and export models that are truly end-to-end: models that accept raw images or raw structured data as input; models that handle feature normalization or feature value indexing on their own.\"\n\nFor better or for worse, we are REQUIRED to do exactly this in this competition. I am calling it a \"Model Competition\", it is NOT a \"Code Competition\", you are required to build a model that handles shape ([n],543,3) and produces the proper 250 output prediction. One successful [example is here](https://www.kaggle.com/code/lonnieqin/isolated-sign-language-recognition-with-dnn/notebook). Unlike a Code Competition, you don't get a chance to process the (n, 543, 3) directly, for example you can't even drop the z values as suggested in the dataset description! (\"The MediaPipe model is not fully trained to predict depth so you may wish to ignore the z values.\")\n\nSo what do you do? Well, I don't know, but it's all in the above link, which I'll be reading. Hopefully a few example notebooks will start teaching some of the building blocks, as well.\n\nNote that TensorFlow Lite shouldn't be any FURTHER handicap. So far as I am aware, you can think of it as a simple compilation and/or compression step, any and all TensorFlow models can be converted to TensorFlow Lite.",
      "votes": null
    },
    {
      "id": "2161108",
      "postDate": "02/27/2023 08:01:58",
      "content": "<p>We've also got all the regular tensorflow operations available to us.  So, for your example of dropping z values, we can use <a href=\"https://www.tensorflow.org/api_docs/python/tf/slice\" target=\"_blank\">tf.slice</a>.  For example…</p>\n<pre><code>vector = tf.slice(vector, [0, 0, 0], [-1, -1, 2])\n</code></pre>\n<p>The same approach can be used if, for example, you wanted to drop the face/pose data and just keep the hand data - which appears as the last 42/543 xyz values.</p>\n<pre><code>vector = tf.slice(vector, [0, 500, 0], [-1, 42, 3])\n</code></pre>\n<p>The example notebook already has examples of a couple of other useful pre-processing operations.</p>\n<pre><code>x = tf.where(tf.math.is_nan(inputs), tf.zeros_like(inputs), inputs)\nx = tf.reduce_mean(x, axis=0, keepdims=True)\n</code></pre>\n<p>The first essentially gives you the ability to have <code>if</code> statements.  The latter, obviously, calculates the mean.</p>",
      "rawMarkdown": "We've also got all the regular tensorflow operations available to us.  So, for your example of dropping z values, we can use [tf.slice](https://www.tensorflow.org/api_docs/python/tf/slice).  For example...\n\n```\nvector = tf.slice(vector, [0, 0, 0], [-1, -1, 2])\n```\n\nThe same approach can be used if, for example, you wanted to drop the face/pose data and just keep the hand data - which appears as the last 42/543 xyz values.\n\n```\nvector = tf.slice(vector, [0, 500, 0], [-1, 42, 3])\n```\n\nThe example notebook already has examples of a couple of other useful pre-processing operations.\n\n```\nx = tf.where(tf.math.is_nan(inputs), tf.zeros_like(inputs), inputs)\nx = tf.reduce_mean(x, axis=0, keepdims=True)\n```\n\nThe first essentially gives you the ability to have `if` statements.  The latter, obviously, calculates the mean.",
      "votes": null
    },
    {
      "id": "2163408",
      "postDate": "02/28/2023 19:36:06",
      "content": "<p>That's helpful, thanks.</p>\n<p>I \"think I know\" how to do some of the more complex stuff, too. But it adds quite a bit of work and complexity when trying to figure out the building blocks, the syntax of those building blocks, and also figuring out how to solve complex problems using (only) those building blocks. Reminds me of my first time using Polars.</p>\n<p>Examples of problems:<br>\nSingle model that's an ensemble of public model A and independent public model B.<br>\nSeparate preprocessing of hands vs face vs pose. -&gt; I think you can make separate tensors and then combine them after you've done all the separate preprocessing you want.</p>",
      "rawMarkdown": "That's helpful, thanks.\n\nI \"think I know\" how to do some of the more complex stuff, too. But it adds quite a bit of work and complexity when trying to figure out the building blocks, the syntax of those building blocks, and also figuring out how to solve complex problems using (only) those building blocks. Reminds me of my first time using Polars.\n\nExamples of problems:\nSingle model that's an ensemble of public model A and independent public model B.\nSeparate preprocessing of hands vs face vs pose. -> I think you can make separate tensors and then combine them after you've done all the separate preprocessing you want.",
      "votes": null
    },
    {
      "id": "2173054",
      "postDate": "03/08/2023 03:51:36",
      "content": "<p>Maybe a dumb question here, but when you say ([n],543,3), is n the batch size? Most notebooks I've seen are treating it this way, including the one you linked. But the only reason it is that shape is because in that notebook he is calling a pre-processing step like so x = tf.reduce_mean(x, axis=0, keepdims=True). The actual data when it's first loaded with the predefined function load_relevant_data_subset returns data with the shape (n, 543, 3) where n is the number of frames, not the batch size. So shouldn't it be possible to create a model that accepts (b, n, 543, 3) where b is the batch size and n is a number of frames? You would have to set n to a specific value and pre-process the data to make n match that value. But shouldn't this be possible?</p>",
      "rawMarkdown": "Maybe a dumb question here, but when you say ([n],543,3), is n the batch size? Most notebooks I've seen are treating it this way, including the one you linked. But the only reason it is that shape is because in that notebook he is calling a pre-processing step like so x = tf.reduce_mean(x, axis=0, keepdims=True). The actual data when it's first loaded with the predefined function load_relevant_data_subset returns data with the shape (n, 543, 3) where n is the number of frames, not the batch size. So shouldn't it be possible to create a model that accepts (b, n, 543, 3) where b is the batch size and n is a number of frames? You would have to set n to a specific value and pre-process the data to make n match that value. But shouldn't this be possible?",
      "votes": null
    },
    {
      "id": "2173201",
      "postDate": "03/08/2023 07:20:43",
      "content": "<p>The n in the original question is the number of frames in a single video.  That's the pre-defined input shape that we'll be given.  (At inference time, we'll be passed a single video at a time.). How you choose to process that is up to you. You can indeed convert a single video into a fixed number of frames e.g. (15, 543, 3) in a preprocessing function.  Then, for training, you send them in in batches, e.g. (64, 15, 543, 3).</p>\n<p>One thing you need to be careful with is if you mix tf functions and keras layers in your main model.  The former has to account for the batch dimension whereas the latter handles it for you and makes it look like it doesn't exist.</p>",
      "rawMarkdown": "The n in the original question is the number of frames in a single video.  That's the pre-defined input shape that we'll be given.  (At inference time, we'll be passed a single video at a time.). How you choose to process that is up to you. You can indeed convert a single video into a fixed number of frames e.g. (15, 543, 3) in a preprocessing function.  Then, for training, you send them in in batches, e.g. (64, 15, 543, 3).\n\nOne thing you need to be careful with is if you mix tf functions and keras layers in your main model.  The former has to account for the batch dimension whereas the latter handles it for you and makes it look like it doesn't exist.",
      "votes": null
    },
    {
      "id": "2174114",
      "postDate": "03/08/2023 21:04:17",
      "content": "<p>Thanks for the reply, the second part of our answer is precisely what I'm struggling with. Do you have any examples or documentation on how to handle this? Been spending hours dealing with shape errors. I'm approaching it by creating a Lambda layer that pre-processes the input, but I constantly get shape errors. Either when the pre-processing layer takes in the input or when it outputs it to the next layer. Thanks!</p>",
      "rawMarkdown": "Thanks for the reply, the second part of our answer is precisely what I'm struggling with. Do you have any examples or documentation on how to handle this? Been spending hours dealing with shape errors. I'm approaching it by creating a Lambda layer that pre-processes the input, but I constantly get shape errors. Either when the pre-processing layer takes in the input or when it outputs it to the next layer. Thanks!",
      "votes": null
    },
    {
      "id": "2174162",
      "postDate": "03/08/2023 22:30:50",
      "content": "<p><a href=\"https://www.kaggle.com/code/roberthatch/gislr-lb-0-63-on-the-shoulders\" target=\"_blank\">My approach</a> - built on examples from others - is to use a preprocessing keras layer class, <a href=\"https://www.kaggle.com/code/roberthatch/gislr-feature-data-on-the-shoulders\" target=\"_blank\">convert it to npy</a> directly for train data, and then include the same preprocessing layer into the final inference model.</p>\n<pre><code>\n(feature_converter(tf.keras.Input((, ), dtype=tf.float32, name=)))\nfeature_converter(load_relevant_data_subset())\n</code></pre>\n<p>Notice the output of this cell prints the output shape, you don't have to guess: (in both the above notebooks)</p>\n<pre><code>KerasTensor(type_spec=TensorSpec(shape=(, ), dtype=tf.float32, name=), name=, description=)\n&lt;tf.Tensor: shape=(, ), dtype=float32, numpy=\narray([[ ,  ,  , ...,\n         ,  , -]], dtype=float32)&gt;\n</code></pre>\n<p>For this competition, your preprocessing should take shape (None, 543, 3) aka tf.keras.Input((543, 3)) and output (1, [your output shape])</p>\n<p>Basically, you know that your preprocessing (if done how I did it) will always have batch size 1, so you leave it off. That way it matches inference, because for this competition you can treat it as a single entry coming in to your model with that (None, 543, 3) shape.</p>\n<p>But you need your final preprocessing output to be (1, [your shape here]), with 1 being the batch size. (Well, I didn't try without it, but I think it might be required). For example (None, 543, 3) -&gt; (1, 10, 543, 3) or my preprocessing (None, 543, 3) -&gt; (1, 5796).</p>\n<p>If you didn't convert to fixed size, let's say all you did was drop z and drop all face landmarks, you can leave it as (1, None, 75, 2). But then you will have to figure out the proper way to store the dataset, I haven't seen examples for that. </p>",
      "rawMarkdown": "[My approach](https://www.kaggle.com/code/roberthatch/gislr-lb-0-63-on-the-shoulders) - built on examples from others - is to use a preprocessing keras layer class, [convert it to npy](https://www.kaggle.com/code/roberthatch/gislr-feature-data-on-the-shoulders) directly for train data, and then include the same preprocessing layer into the final inference model.\n\n```python\n## One tests symbolic tensor, the other tests real data.\nprint(feature_converter(tf.keras.Input((543, 3), dtype=tf.float32, name=\"inputs\")))\nfeature_converter(load_relevant_data_subset(f'/kaggle/input/asl-signs/{pd.read_csv(TRAIN_FILE).path[1]}'))\n```\nNotice the output of this cell prints the output shape, you don't have to guess: (in both the above notebooks)\n```python\nKerasTensor(type_spec=TensorSpec(shape=(1, 5796), dtype=tf.float32, name=None), name='feature_gen/concat_5:0', description=\"created by layer 'feature_gen'\")\n<tf.Tensor: shape=(1, 5796), dtype=float32, numpy=\narray([[ 5.18924236e-01,  3.42620254e-01,  1.48732506e-05, ...,\n         6.35599867e-02,  5.70323110e-01, -1.20788895e-01]], dtype=float32)>\n```\nFor this competition, your preprocessing should take shape (None, 543, 3) aka tf.keras.Input((543, 3)) and output (1, [your output shape])\n\nBasically, you know that your preprocessing (if done how I did it) will always have batch size 1, so you leave it off. That way it matches inference, because for this competition you can treat it as a single entry coming in to your model with that (None, 543, 3) shape.\n\nBut you need your final preprocessing output to be (1, [your shape here]), with 1 being the batch size. (Well, I didn't try without it, but I think it might be required). For example (None, 543, 3) -> (1, 10, 543, 3) or my preprocessing (None, 543, 3) -> (1, 5796).\n\nIf you didn't convert to fixed size, let's say all you did was drop z and drop all face landmarks, you can leave it as (1, None, 75, 2). But then you will have to figure out the proper way to store the dataset, I haven't seen examples for that.",
      "votes": null
    },
    {
      "id": "2174244",
      "postDate": "03/09/2023 01:41:30",
      "content": "<p>i am not familiar with keras, i have a question to ask.</p>\n<p>in pytorch you can write:</p>\n<pre><code>def forward(xyz):\n     L=xyz.shape[0]\n     if L%2==1 :\n          xyz = F.pad(xyz, ...)  #. e.g. pad if video length even is not even \n</code></pre>\n<p>\"python if condition\" will be build into the graph when we call torch jit script function.</p>\n<p>in keras, does it still work?<br>\ni read that you must explicitly use e.g. tf.cond()</p>",
      "rawMarkdown": "i am not familiar with keras, i have a question to ask.\n\nin pytorch you can write:\n\n```\ndef forward(xyz):\n     L=xyz.shape[0]\n     if L%2==1 :\n          xyz = F.pad(xyz, ...)  #. e.g. pad if video length even is not even \n\n```\n\n\"python if condition\" will be build into the graph when we call torch jit script function.\n\nin keras, does it still work?\ni read that you must explicitly use e.g. tf.cond()",
      "votes": null
    },
    {
      "id": "2174257",
      "postDate": "03/09/2023 02:12:28",
      "content": "<p>I don't think it will work. Even tf.cond not sure if it would work as-is. I recall the shape has to be exactly the same for both sides of the if/else functions with tf.cond. Though maybe there's a way of telling keras to continue using 'None' for that variable dimension…</p>",
      "rawMarkdown": "I don't think it will work. Even tf.cond not sure if it would work as-is. I recall the shape has to be exactly the same for both sides of the if/else functions with tf.cond. Though maybe there's a way of telling keras to continue using 'None' for that variable dimension...",
      "votes": null
    },
    {
      "id": "2174342",
      "postDate": "03/09/2023 04:36:45",
      "content": "<p>my trick in pytorch is to:</p>\n<ol>\n<li>just pad one regardless of length</li>\n<li>truncate x[:(L//2)*2]</li>\n</ol>\n<p>to avoid tfReadDiv in tflite, i need to make a silly lookup table nn.Parameter to do int division</p>",
      "rawMarkdown": "my trick in pytorch is to:\n1. just pad one regardless of length\n2. truncate x[:(L//2)*2]\n\nto avoid tfReadDiv in tflite, i need to make a silly lookup table nn.Parameter to do int division",
      "votes": null
    },
    {
      "id": "2198342",
      "postDate": "03/26/2023 23:34:02",
      "content": "<p>Thank you, for this thread, it is very insightful and I have been coming back to it since I'm new to TF processing layers.<br>\none thing I noticed though is that in the second code, you are assuming that the hands' landmarks are coming sequentially when in fact it is <code>left_hand</code>, <code>pose</code>, then <code>right_hand</code>.<br>\nAm I correct or did I misunderstand this?<br>\nThanks again!</p>",
      "rawMarkdown": "Thank you, for this thread, it is very insightful and I have been coming back to it since I'm new to TF processing layers.\none thing I noticed though is that in the second code, you are assuming that the hands' landmarks are coming sequentially when in fact it is `left_hand`, `pose`, then `right_hand`.\nAm I correct or did I misunderstand this?\nThanks again!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2161108,
      "author_name": "andrewrrose",
      "author_url": "",
      "post_date": "02/27/2023 08:01:58",
      "content": "<p>We've also got all the regular tensorflow operations available to us.  So, for your example of dropping z values, we can use <a href=\"https://www.tensorflow.org/api_docs/python/tf/slice\" target=\"_blank\">tf.slice</a>.  For example…</p>\n<pre><code>vector = tf.slice(vector, [0, 0, 0], [-1, -1, 2])\n</code></pre>\n<p>The same approach can be used if, for example, you wanted to drop the face/pose data and just keep the hand data - which appears as the last 42/543 xyz values.</p>\n<pre><code>vector = tf.slice(vector, [0, 500, 0], [-1, 42, 3])\n</code></pre>\n<p>The example notebook already has examples of a couple of other useful pre-processing operations.</p>\n<pre><code>x = tf.where(tf.math.is_nan(inputs), tf.zeros_like(inputs), inputs)\nx = tf.reduce_mean(x, axis=0, keepdims=True)\n</code></pre>\n<p>The first essentially gives you the ability to have <code>if</code> statements.  The latter, obviously, calculates the mean.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2163408,
          "author_name": "roberthatch",
          "author_url": "",
          "post_date": "02/28/2023 19:36:06",
          "content": "<p>That's helpful, thanks.</p>\n<p>I \"think I know\" how to do some of the more complex stuff, too. But it adds quite a bit of work and complexity when trying to figure out the building blocks, the syntax of those building blocks, and also figuring out how to solve complex problems using (only) those building blocks. Reminds me of my first time using Polars.</p>\n<p>Examples of problems:<br>\nSingle model that's an ensemble of public model A and independent public model B.<br>\nSeparate preprocessing of hands vs face vs pose. -&gt; I think you can make separate tensors and then combine them after you've done all the separate preprocessing you want.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2198342,
          "author_name": "rashasalim",
          "author_url": "",
          "post_date": "03/26/2023 23:34:02",
          "content": "<p>Thank you, for this thread, it is very insightful and I have been coming back to it since I'm new to TF processing layers.<br>\none thing I noticed though is that in the second code, you are assuming that the hands' landmarks are coming sequentially when in fact it is <code>left_hand</code>, <code>pose</code>, then <code>right_hand</code>.<br>\nAm I correct or did I misunderstand this?<br>\nThanks again!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2173054,
      "author_name": "conweezy",
      "author_url": "",
      "post_date": "03/08/2023 03:51:36",
      "content": "<p>Maybe a dumb question here, but when you say ([n],543,3), is n the batch size? Most notebooks I've seen are treating it this way, including the one you linked. But the only reason it is that shape is because in that notebook he is calling a pre-processing step like so x = tf.reduce_mean(x, axis=0, keepdims=True). The actual data when it's first loaded with the predefined function load_relevant_data_subset returns data with the shape (n, 543, 3) where n is the number of frames, not the batch size. So shouldn't it be possible to create a model that accepts (b, n, 543, 3) where b is the batch size and n is a number of frames? You would have to set n to a specific value and pre-process the data to make n match that value. But shouldn't this be possible?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2173201,
          "author_name": "andrewrrose",
          "author_url": "",
          "post_date": "03/08/2023 07:20:43",
          "content": "<p>The n in the original question is the number of frames in a single video.  That's the pre-defined input shape that we'll be given.  (At inference time, we'll be passed a single video at a time.). How you choose to process that is up to you. You can indeed convert a single video into a fixed number of frames e.g. (15, 543, 3) in a preprocessing function.  Then, for training, you send them in in batches, e.g. (64, 15, 543, 3).</p>\n<p>One thing you need to be careful with is if you mix tf functions and keras layers in your main model.  The former has to account for the batch dimension whereas the latter handles it for you and makes it look like it doesn't exist.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2174114,
              "author_name": "conweezy",
              "author_url": "",
              "post_date": "03/08/2023 21:04:17",
              "content": "<p>Thanks for the reply, the second part of our answer is precisely what I'm struggling with. Do you have any examples or documentation on how to handle this? Been spending hours dealing with shape errors. I'm approaching it by creating a Lambda layer that pre-processes the input, but I constantly get shape errors. Either when the pre-processing layer takes in the input or when it outputs it to the next layer. Thanks!</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2174162,
                  "author_name": "roberthatch",
                  "author_url": "",
                  "post_date": "03/08/2023 22:30:50",
                  "content": "<p><a href=\"https://www.kaggle.com/code/roberthatch/gislr-lb-0-63-on-the-shoulders\" target=\"_blank\">My approach</a> - built on examples from others - is to use a preprocessing keras layer class, <a href=\"https://www.kaggle.com/code/roberthatch/gislr-feature-data-on-the-shoulders\" target=\"_blank\">convert it to npy</a> directly for train data, and then include the same preprocessing layer into the final inference model.</p>\n<pre><code>\n(feature_converter(tf.keras.Input((, ), dtype=tf.float32, name=)))\nfeature_converter(load_relevant_data_subset())\n</code></pre>\n<p>Notice the output of this cell prints the output shape, you don't have to guess: (in both the above notebooks)</p>\n<pre><code>KerasTensor(type_spec=TensorSpec(shape=(, ), dtype=tf.float32, name=), name=, description=)\n&lt;tf.Tensor: shape=(, ), dtype=float32, numpy=\narray([[ ,  ,  , ...,\n         ,  , -]], dtype=float32)&gt;\n</code></pre>\n<p>For this competition, your preprocessing should take shape (None, 543, 3) aka tf.keras.Input((543, 3)) and output (1, [your output shape])</p>\n<p>Basically, you know that your preprocessing (if done how I did it) will always have batch size 1, so you leave it off. That way it matches inference, because for this competition you can treat it as a single entry coming in to your model with that (None, 543, 3) shape.</p>\n<p>But you need your final preprocessing output to be (1, [your shape here]), with 1 being the batch size. (Well, I didn't try without it, but I think it might be required). For example (None, 543, 3) -&gt; (1, 10, 543, 3) or my preprocessing (None, 543, 3) -&gt; (1, 5796).</p>\n<p>If you didn't convert to fixed size, let's say all you did was drop z and drop all face landmarks, you can leave it as (1, None, 75, 2). But then you will have to figure out the proper way to store the dataset, I haven't seen examples for that. </p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2174244,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/09/2023 01:41:30",
      "content": "<p>i am not familiar with keras, i have a question to ask.</p>\n<p>in pytorch you can write:</p>\n<pre><code>def forward(xyz):\n     L=xyz.shape[0]\n     if L%2==1 :\n          xyz = F.pad(xyz, ...)  #. e.g. pad if video length even is not even \n</code></pre>\n<p>\"python if condition\" will be build into the graph when we call torch jit script function.</p>\n<p>in keras, does it still work?<br>\ni read that you must explicitly use e.g. tf.cond()</p>",
      "votes": null,
      "replies": [
        {
          "id": 2174257,
          "author_name": "roberthatch",
          "author_url": "",
          "post_date": "03/09/2023 02:12:28",
          "content": "<p>I don't think it will work. Even tf.cond not sure if it would work as-is. I recall the shape has to be exactly the same for both sides of the if/else functions with tf.cond. Though maybe there's a way of telling keras to continue using 'None' for that variable dimension…</p>",
          "votes": null,
          "replies": [
            {
              "id": 2174342,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "03/09/2023 04:36:45",
              "content": "<p>my trick in pytorch is to:</p>\n<ol>\n<li>just pad one regardless of length</li>\n<li>truncate x[:(L//2)*2]</li>\n</ol>\n<p>to avoid tfReadDiv in tflite, i need to make a silly lookup table nn.Parameter to do int division</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2160652": "https://www.tensorflow.org/guide/keras/preprocessing_layers\n\n\"With Keras preprocessing layers, you can build and export models that are truly end-to-end: models that accept raw images or raw structured data as input; models that handle feature normalization or feature value indexing on their own.\"\n\nFor better or for worse, we are REQUIRED to do exactly this in this competition. I am calling it a \"Model Competition\", it is NOT a \"Code Competition\", you are required to build a model that handles shape ([n],543,3) and produces the proper 250 output prediction. One successful [example is here](https://www.kaggle.com/code/lonnieqin/isolated-sign-language-recognition-with-dnn/notebook). Unlike a Code Competition, you don't get a chance to process the (n, 543, 3) directly, for example you can't even drop the z values as suggested in the dataset description! (\"The MediaPipe model is not fully trained to predict depth so you may wish to ignore the z values.\")\n\nSo what do you do? Well, I don't know, but it's all in the above link, which I'll be reading. Hopefully a few example notebooks will start teaching some of the building blocks, as well.\n\nNote that TensorFlow Lite shouldn't be any FURTHER handicap. So far as I am aware, you can think of it as a simple compilation and/or compression step, any and all TensorFlow models can be converted to TensorFlow Lite.",
    "2161108": "We've also got all the regular tensorflow operations available to us.  So, for your example of dropping z values, we can use [tf.slice](https://www.tensorflow.org/api_docs/python/tf/slice).  For example...\n\n```\nvector = tf.slice(vector, [0, 0, 0], [-1, -1, 2])\n```\n\nThe same approach can be used if, for example, you wanted to drop the face/pose data and just keep the hand data - which appears as the last 42/543 xyz values.\n\n```\nvector = tf.slice(vector, [0, 500, 0], [-1, 42, 3])\n```\n\nThe example notebook already has examples of a couple of other useful pre-processing operations.\n\n```\nx = tf.where(tf.math.is_nan(inputs), tf.zeros_like(inputs), inputs)\nx = tf.reduce_mean(x, axis=0, keepdims=True)\n```\n\nThe first essentially gives you the ability to have `if` statements.  The latter, obviously, calculates the mean.",
    "2163408": "That's helpful, thanks.\n\nI \"think I know\" how to do some of the more complex stuff, too. But it adds quite a bit of work and complexity when trying to figure out the building blocks, the syntax of those building blocks, and also figuring out how to solve complex problems using (only) those building blocks. Reminds me of my first time using Polars.\n\nExamples of problems:\nSingle model that's an ensemble of public model A and independent public model B.\nSeparate preprocessing of hands vs face vs pose. -> I think you can make separate tensors and then combine them after you've done all the separate preprocessing you want.",
    "2173054": "Maybe a dumb question here, but when you say ([n],543,3), is n the batch size? Most notebooks I've seen are treating it this way, including the one you linked. But the only reason it is that shape is because in that notebook he is calling a pre-processing step like so x = tf.reduce_mean(x, axis=0, keepdims=True). The actual data when it's first loaded with the predefined function load_relevant_data_subset returns data with the shape (n, 543, 3) where n is the number of frames, not the batch size. So shouldn't it be possible to create a model that accepts (b, n, 543, 3) where b is the batch size and n is a number of frames? You would have to set n to a specific value and pre-process the data to make n match that value. But shouldn't this be possible?",
    "2173201": "The n in the original question is the number of frames in a single video.  That's the pre-defined input shape that we'll be given.  (At inference time, we'll be passed a single video at a time.). How you choose to process that is up to you. You can indeed convert a single video into a fixed number of frames e.g. (15, 543, 3) in a preprocessing function.  Then, for training, you send them in in batches, e.g. (64, 15, 543, 3).\n\nOne thing you need to be careful with is if you mix tf functions and keras layers in your main model.  The former has to account for the batch dimension whereas the latter handles it for you and makes it look like it doesn't exist.",
    "2174114": "Thanks for the reply, the second part of our answer is precisely what I'm struggling with. Do you have any examples or documentation on how to handle this? Been spending hours dealing with shape errors. I'm approaching it by creating a Lambda layer that pre-processes the input, but I constantly get shape errors. Either when the pre-processing layer takes in the input or when it outputs it to the next layer. Thanks!",
    "2174162": "[My approach](https://www.kaggle.com/code/roberthatch/gislr-lb-0-63-on-the-shoulders) - built on examples from others - is to use a preprocessing keras layer class, [convert it to npy](https://www.kaggle.com/code/roberthatch/gislr-feature-data-on-the-shoulders) directly for train data, and then include the same preprocessing layer into the final inference model.\n\n```python\n## One tests symbolic tensor, the other tests real data.\nprint(feature_converter(tf.keras.Input((543, 3), dtype=tf.float32, name=\"inputs\")))\nfeature_converter(load_relevant_data_subset(f'/kaggle/input/asl-signs/{pd.read_csv(TRAIN_FILE).path[1]}'))\n```\nNotice the output of this cell prints the output shape, you don't have to guess: (in both the above notebooks)\n```python\nKerasTensor(type_spec=TensorSpec(shape=(1, 5796), dtype=tf.float32, name=None), name='feature_gen/concat_5:0', description=\"created by layer 'feature_gen'\")\n<tf.Tensor: shape=(1, 5796), dtype=float32, numpy=\narray([[ 5.18924236e-01,  3.42620254e-01,  1.48732506e-05, ...,\n         6.35599867e-02,  5.70323110e-01, -1.20788895e-01]], dtype=float32)>\n```\nFor this competition, your preprocessing should take shape (None, 543, 3) aka tf.keras.Input((543, 3)) and output (1, [your output shape])\n\nBasically, you know that your preprocessing (if done how I did it) will always have batch size 1, so you leave it off. That way it matches inference, because for this competition you can treat it as a single entry coming in to your model with that (None, 543, 3) shape.\n\nBut you need your final preprocessing output to be (1, [your shape here]), with 1 being the batch size. (Well, I didn't try without it, but I think it might be required). For example (None, 543, 3) -> (1, 10, 543, 3) or my preprocessing (None, 543, 3) -> (1, 5796).\n\nIf you didn't convert to fixed size, let's say all you did was drop z and drop all face landmarks, you can leave it as (1, None, 75, 2). But then you will have to figure out the proper way to store the dataset, I haven't seen examples for that.",
    "2174244": "i am not familiar with keras, i have a question to ask.\n\nin pytorch you can write:\n\n```\ndef forward(xyz):\n     L=xyz.shape[0]\n     if L%2==1 :\n          xyz = F.pad(xyz, ...)  #. e.g. pad if video length even is not even \n\n```\n\n\"python if condition\" will be build into the graph when we call torch jit script function.\n\nin keras, does it still work?\ni read that you must explicitly use e.g. tf.cond()",
    "2174257": "I don't think it will work. Even tf.cond not sure if it would work as-is. I recall the shape has to be exactly the same for both sides of the if/else functions with tf.cond. Though maybe there's a way of telling keras to continue using 'None' for that variable dimension...",
    "2174342": "my trick in pytorch is to:\n1. just pad one regardless of length\n2. truncate x[:(L//2)*2]\n\nto avoid tfReadDiv in tflite, i need to make a silly lookup table nn.Parameter to do int division",
    "2198342": "Thank you, for this thread, it is very insightful and I have been coming back to it since I'm new to TF processing layers.\none thing I noticed though is that in the second code, you are assuming that the hands' landmarks are coming sequentially when in fact it is `left_hand`, `pose`, then `right_hand`.\nAm I correct or did I misunderstand this?\nThanks again!"
  },
  "source": "meta"
}