{
  "id": 401481,
  "title": "Very slow fitting (about 10 hours for 10 epochs",
  "url": "/competitions/asl-signs/discussion/401481",
  "author_name": "",
  "post_date": "2023-04-13T12:56:55.185139400Z",
  "votes": 8,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Good afternoon.<br>\nI have 10 epochs trained in about 10 hours. So slow. I suspect it's because of the way the data is processed. But it is not exactly</p>\n<h3>Here is the code for loading the file and converting it to Numpy ndarray:</h3>\n<pre><code> ():\n    data = pd.read_parquet(os.path.join(,pq_path), columns=PARQUET_RELEVANT_ROWS).to_numpy()\n    data = data.reshape((data)//PARQUET_ROWS_PER_FRAME, PARQUET_ROWS_PER_FRAME, (PARQUET_RELEVANT_ROWS))\n    data = PrepareFrames(data)\n     data\n</code></pre>\n<h3>This is the code that I use to determine the dominant hand, remove skipping frames (without hands and face), and also bring all the data to the same number of frames:</h3>\n<pre><code> ():\n    \n    left_hand_dominate = \n     np.count_nonzero(~np.isnan(\n            frames[:,np.array(important_indexes[:]).astype(),:].flatten()\n        )) &lt; \\\n        np.count_nonzero(~np.isnan(\n            frames[:,np.array(important_indexes[-:]).astype(),:].flatten()\n        )):\n        left_hand_dominate = \n        frames[:,:] = frames[:,:,]\n        frames = frames[:,:,]\n    :\n        frames = frames[:,:,]\n\n    \n    non_nan_frames = np.array(\n        [frame  frame  frames \n          (~np.isnan(frame[np.array(important_indexes[:]),:]).(axis=)).()\n        ]\n    )\n    \n     (non_nan_frames) == :\n        non_nan_frames = np.random.uniform(low=, high=, size=(, ,))\n\n    \n     non_nan_frames.shape[] &gt; LIMIT_MAX_FRAMES:\n        interval = np.linspace(, non_nan_frames.shape[]-, LIMIT_MAX_FRAMES, dtype=)\n        non_nan_frames = non_nan_frames[np.array(interval),:,:]\n    :\n        tiling = np.tile(non_nan_frames[-], (LIMIT_MAX_FRAMES-non_nan_frames.shape[], )).reshape(-,,)\n        non_nan_frames = np.vstack([non_nan_frames,tiling])\n\n     non_nan_frames\n</code></pre>\n<h3>There is also a method by which I extract features from the data.</h3>\n<pre><code> ():\n    \n     ():\n        self._torso_size_multiplier = torso_size_multiplier\n\n        embedding = np.array([\n            self._get_angle_by_names(landmarks, , , , , degrees=),\n            self._get_distance(\n                                self._get_coordinates_by_name(landmarks, ),\n                                self._get_coordinates_by_name(landmarks, )\n                              )\n        ])\n         embedding.T\n     ():\n        lmk_from = landmarks[:, landmark_indexes[landmark_from],:]\n        lmk_to = landmarks[:, landmark_indexes[landmark_to],:]\n         (lmk_from + lmk_to) * \n\n     ():\n         [np.sqrt(np.((p1 - p2)**))  p1, p2  (lmk_from, lmk_to)]\n\n     ():\n         landmarks[:, landmark_indexes[landmark],:]\n\n     ():\n         self._get_angle(\n            landmarks[:, landmark_indexes[landmark_1],:],\n            landmarks[:, landmark_indexes[landmark_2],:],\n            landmarks[:, landmark_indexes[landmark_3],:],\n            landmarks[:, landmark_indexes[landmark_4],:],\n            degrees\n        )\n\n     ():\n        a = np.subtract(lmk_1,lmk_2)\n        b = np.subtract(lmk_3,lmk_4)\n\n        dot_products = [([v1[i] * v2[i]  i  ((v1))])  v1, v2  (a,b)]\n        mag1s = [math.sqrt(([x**  x  v1]))  v1, _  (a,b)]\n        mag2s = [math.sqrt(([x**  x  v2]))  _, v2  (a,b)]\n        cos_thetas = [dot_product / (mag1 * mag2)  dot_product, mag1, mag2  (dot_products, mag1s, mag2s)]\n        thetas = [math.acos(cos_theta)  cos_theta  cos_thetas]\n        angles = [math.degrees(theta)  theta  thetas]\n         degrees:\n             angles\n        :\n             thetas\n\nsign_embedder = SignEmbedder()\n</code></pre>\n<h3>Loading data into Tensorflow DataSet</h3>\n<pre><code> ():\n     ():\n         sign_embedder(parquet_to_numpy(ftensor.numpy().decode()))\n     tf.py_function(\n        feat_wrapper,\n        [ftensor],\n        Tout=tf.float32\n    )\n ():\n\n    \n     tf.ensure_shape(x, ( LIMIT_MAX_FRAMES, ))\n\nX = tf.data.Dataset.from_tensor_slices(\n    train_df.path.values\n).(\n    tf_get_features\n).(\n    set_shape\n)\n</code></pre>\n<h3>Here are some performance tests</h3>\n<pre><code>%%timeit\n\nBATCH_SIZE = \n\n path  train_df.path.values[:BATCH_SIZE]:\n    data = sign_embedder(parquet_to_numpy(path))\n</code></pre>\n<blockquote>\n  <p>6.28 s ± 265 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)</p>\n</blockquote>",
  "messages": [
    {
      "id": "2220496",
      "postDate": "04/13/2023 12:56:55",
      "content": "<p>Good afternoon.<br>\nI have 10 epochs trained in about 10 hours. So slow. I suspect it's because of the way the data is processed. But it is not exactly</p>\n<h3>Here is the code for loading the file and converting it to Numpy ndarray:</h3>\n<pre><code> ():\n    data = pd.read_parquet(os.path.join(,pq_path), columns=PARQUET_RELEVANT_ROWS).to_numpy()\n    data = data.reshape((data)//PARQUET_ROWS_PER_FRAME, PARQUET_ROWS_PER_FRAME, (PARQUET_RELEVANT_ROWS))\n    data = PrepareFrames(data)\n     data\n</code></pre>\n<h3>This is the code that I use to determine the dominant hand, remove skipping frames (without hands and face), and also bring all the data to the same number of frames:</h3>\n<pre><code> ():\n    \n    left_hand_dominate = \n     np.count_nonzero(~np.isnan(\n            frames[:,np.array(important_indexes[:]).astype(),:].flatten()\n        )) &lt; \\\n        np.count_nonzero(~np.isnan(\n            frames[:,np.array(important_indexes[-:]).astype(),:].flatten()\n        )):\n        left_hand_dominate = \n        frames[:,:] = frames[:,:,]\n        frames = frames[:,:,]\n    :\n        frames = frames[:,:,]\n\n    \n    non_nan_frames = np.array(\n        [frame  frame  frames \n          (~np.isnan(frame[np.array(important_indexes[:]),:]).(axis=)).()\n        ]\n    )\n    \n     (non_nan_frames) == :\n        non_nan_frames = np.random.uniform(low=, high=, size=(, ,))\n\n    \n     non_nan_frames.shape[] &gt; LIMIT_MAX_FRAMES:\n        interval = np.linspace(, non_nan_frames.shape[]-, LIMIT_MAX_FRAMES, dtype=)\n        non_nan_frames = non_nan_frames[np.array(interval),:,:]\n    :\n        tiling = np.tile(non_nan_frames[-], (LIMIT_MAX_FRAMES-non_nan_frames.shape[], )).reshape(-,,)\n        non_nan_frames = np.vstack([non_nan_frames,tiling])\n\n     non_nan_frames\n</code></pre>\n<h3>There is also a method by which I extract features from the data.</h3>\n<pre><code> ():\n    \n     ():\n        self._torso_size_multiplier = torso_size_multiplier\n\n        embedding = np.array([\n            self._get_angle_by_names(landmarks, , , , , degrees=),\n            self._get_distance(\n                                self._get_coordinates_by_name(landmarks, ),\n                                self._get_coordinates_by_name(landmarks, )\n                              )\n        ])\n         embedding.T\n     ():\n        lmk_from = landmarks[:, landmark_indexes[landmark_from],:]\n        lmk_to = landmarks[:, landmark_indexes[landmark_to],:]\n         (lmk_from + lmk_to) * \n\n     ():\n         [np.sqrt(np.((p1 - p2)**))  p1, p2  (lmk_from, lmk_to)]\n\n     ():\n         landmarks[:, landmark_indexes[landmark],:]\n\n     ():\n         self._get_angle(\n            landmarks[:, landmark_indexes[landmark_1],:],\n            landmarks[:, landmark_indexes[landmark_2],:],\n            landmarks[:, landmark_indexes[landmark_3],:],\n            landmarks[:, landmark_indexes[landmark_4],:],\n            degrees\n        )\n\n     ():\n        a = np.subtract(lmk_1,lmk_2)\n        b = np.subtract(lmk_3,lmk_4)\n\n        dot_products = [([v1[i] * v2[i]  i  ((v1))])  v1, v2  (a,b)]\n        mag1s = [math.sqrt(([x**  x  v1]))  v1, _  (a,b)]\n        mag2s = [math.sqrt(([x**  x  v2]))  _, v2  (a,b)]\n        cos_thetas = [dot_product / (mag1 * mag2)  dot_product, mag1, mag2  (dot_products, mag1s, mag2s)]\n        thetas = [math.acos(cos_theta)  cos_theta  cos_thetas]\n        angles = [math.degrees(theta)  theta  thetas]\n         degrees:\n             angles\n        :\n             thetas\n\nsign_embedder = SignEmbedder()\n</code></pre>\n<h3>Loading data into Tensorflow DataSet</h3>\n<pre><code> ():\n     ():\n         sign_embedder(parquet_to_numpy(ftensor.numpy().decode()))\n     tf.py_function(\n        feat_wrapper,\n        [ftensor],\n        Tout=tf.float32\n    )\n ():\n\n    \n     tf.ensure_shape(x, ( LIMIT_MAX_FRAMES, ))\n\nX = tf.data.Dataset.from_tensor_slices(\n    train_df.path.values\n).(\n    tf_get_features\n).(\n    set_shape\n)\n</code></pre>\n<h3>Here are some performance tests</h3>\n<pre><code>%%timeit\n\nBATCH_SIZE = \n\n path  train_df.path.values[:BATCH_SIZE]:\n    data = sign_embedder(parquet_to_numpy(path))\n</code></pre>\n<blockquote>\n  <p>6.28 s ± 265 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)</p>\n</blockquote>",
      "rawMarkdown": "Good afternoon.\nI have 10 epochs trained in about 10 hours. So slow. I suspect it's because of the way the data is processed. But it is not exactly\n\n### Here is the code for loading the file and converting it to Numpy ndarray:\n```python\ndef parquet_to_numpy(pq_path):\n    data = pd.read_parquet(os.path.join('/kaggle/input/asl-signs/',pq_path), columns=PARQUET_RELEVANT_ROWS).to_numpy()\n    data = data.reshape(len(data)//PARQUET_ROWS_PER_FRAME, PARQUET_ROWS_PER_FRAME, len(PARQUET_RELEVANT_ROWS))\n    data = PrepareFrames(data)\n    return data\n```\n\n### This is the code that I use to determine the dominant hand, remove skipping frames (without hands and face), and also bring all the data to the same number of frames:\n\n```python\ndef PrepareFrames(frames):\n    # Determining the dominant hand\n    left_hand_dominate = True\n    if np.count_nonzero(~np.isnan(\n            frames[:,np.array(important_indexes[:21]).astype(int),:].flatten()\n        )) < \\\n        np.count_nonzero(~np.isnan(\n            frames[:,np.array(important_indexes[-21:]).astype(int),:].flatten()\n        )):\n        left_hand_dominate = False\n        frames[:,468:489] = frames[:,522:,]\n        frames = frames[:,:522,]\n    else:\n        frames = frames[:,:522,]\n   \n    # Creating a new ndarray without dropping frames\n    non_nan_frames = np.array(\n        [frame for frame in frames \n         if (~np.isnan(frame[np.array(important_indexes[:28]),:]).all(axis=1)).all()\n        ]\n    )\n    # There is a gag, because I haven't found a solution for skipping empty files yet\n    if len(non_nan_frames) == 0:\n        non_nan_frames = np.random.uniform(low=0.1, high=2, size=(20, 522,3))#np.zeros((20, 522,3))\n    \n    # Here I trim the number of frames if it is greater than LIMIT_MAX_FRAMES or duplicate the last frames if it is less\n    if non_nan_frames.shape[0] > LIMIT_MAX_FRAMES:\n        interval = np.linspace(0, non_nan_frames.shape[0]-1, LIMIT_MAX_FRAMES, dtype=int)\n        non_nan_frames = non_nan_frames[np.array(interval),:,:]\n    else:\n        tiling = np.tile(non_nan_frames[-1], (LIMIT_MAX_FRAMES-non_nan_frames.shape[0], 1)).reshape(-1,522,3)\n        non_nan_frames = np.vstack([non_nan_frames,tiling])\n\n    return non_nan_frames\n```\n\n### There is also a method by which I extract features from the data.\n\n```python\nclass SignEmbedder(object):\n    \"\"\"\n    \"\"\"\n    def __init__(self, torso_size_multiplier=2.5):\n        self._torso_size_multiplier = torso_size_multiplier\n  \n        embedding = np.array([\n            self._get_angle_by_names(landmarks, 'left_hand-1', 'left_hand-2', 'left_hand-2', 'left_hand-3', degrees=False),\n            self._get_distance(\n                                self._get_coordinates_by_name(landmarks, 'left_hand-4'),\n                                self._get_coordinates_by_name(landmarks, 'left_hand-0')\n                              )\n        ])\n        return embedding.T\n    def _get_average_by_names(self, landmarks, landmark_from, landmark_to):\n        lmk_from = landmarks[:, landmark_indexes[landmark_from],:]\n        lmk_to = landmarks[:, landmark_indexes[landmark_to],:]\n        return (lmk_from + lmk_to) * 0.5\n    \n    def _get_distance(self, lmk_from, lmk_to):\n        return [np.sqrt(np.sum((p1 - p2)**2)) for p1, p2 in zip(lmk_from, lmk_to)]\n    \n    def _get_coordinates_by_name(self, landmarks, landmark):\n        return landmarks[:, landmark_indexes[landmark],:]\n    \n    def _get_angle_by_names(self,landmarks,landmark_1, landmark_2,landmark_3, landmark_4, degrees = True):\n        return self._get_angle(\n            landmarks[:, landmark_indexes[landmark_1],:],\n            landmarks[:, landmark_indexes[landmark_2],:],\n            landmarks[:, landmark_indexes[landmark_3],:],\n            landmarks[:, landmark_indexes[landmark_4],:],\n            degrees\n        )\n    \n    def _get_angle(self,lmk_1, lmk_2, lmk_3, lmk_4, degrees = True):\n        a = np.subtract(lmk_1,lmk_2)\n        b = np.subtract(lmk_3,lmk_4)\n        \n        dot_products = [sum([v1[i] * v2[i] for i in range(len(v1))]) for v1, v2 in zip(a,b)]\n        mag1s = [math.sqrt(sum([x**2 for x in v1])) for v1, _ in zip(a,b)]\n        mag2s = [math.sqrt(sum([x**2 for x in v2])) for _, v2 in zip(a,b)]\n        cos_thetas = [dot_product / (mag1 * mag2) for dot_product, mag1, mag2 in zip(dot_products, mag1s, mag2s)]\n        thetas = [math.acos(cos_theta) for cos_theta in cos_thetas]\n        angles = [math.degrees(theta) for theta in thetas]\n        if degrees:\n            return angles\n        else:\n            return thetas\n        \nsign_embedder = SignEmbedder()\n```\n\n### Loading data into Tensorflow DataSet\n\n```python\ndef tf_get_features(ftensor):\n    def feat_wrapper(ftensor):\n        return sign_embedder(parquet_to_numpy(ftensor.numpy().decode('utf-8')))\n    return tf.py_function(\n        feat_wrapper,\n        [ftensor],\n        Tout=tf.float32\n    )\ndef set_shape(x):\n    \n    # None dimensions can be of any length\n    return tf.ensure_shape(x, ( LIMIT_MAX_FRAMES, 37))\n\nX = tf.data.Dataset.from_tensor_slices(\n    train_df.path.values\n).map(\n    tf_get_features\n).map(\n    set_shape\n)\n\n```\n\n### Here are some performance tests\n\n```python\n%%timeit\n\nBATCH_SIZE = 512\n\nfor path in train_df.path.values[:BATCH_SIZE]:\n    data = sign_embedder(parquet_to_numpy(path))\n```\n\n>6.28 s ± 265 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)",
      "votes": null
    },
    {
      "id": "2221073",
      "postDate": "04/14/2023 01:03:09",
      "content": "<p>You are preprocessing data by numpy  and converting data into tensorflow right?<br>\nConverting data from numpy into tensorflow may be very very slow, so you should rewrite your preprocessing layer all by tensorflow.</p>\n<p>Anyway, we must have preprocessing layer in our tflite model, so we can't use numpy I think.<br>\nConverting Preprocessing layer written by numpy into tflite model failed in my case.</p>",
      "rawMarkdown": "You are preprocessing data by numpy  and converting data into tensorflow right?\nConverting data from numpy into tensorflow may be very very slow, so you should rewrite your preprocessing layer all by tensorflow.\n\nAnyway, we must have preprocessing layer in our tflite model, so we can't use numpy I think.\nConverting Preprocessing layer written by numpy into tflite model failed in my case.",
      "votes": null
    },
    {
      "id": "2221220",
      "postDate": "04/14/2023 04:56:48",
      "content": "<p>check nvidia-smi gpu utilization/power consumption during training, <br>\nfrom the looks of it your batch generation is a huge bottleneck, <br>\nwhat you can do<br>\n1) remove python loops and convert the code to numpy/tf only ops, it will speed up things a lot<br>\n2) during prototyping you can leave code as is, however save the prepared dataset to disk and load it during training, with the fast enough disk (ssd/m2) it'll be very cheap, and you'll observe huge speed-up</p>",
      "rawMarkdown": "check nvidia-smi gpu utilization/power consumption during training, \nfrom the looks of it your batch generation is a huge bottleneck, \nwhat you can do\n1) remove python loops and convert the code to numpy/tf only ops, it will speed up things a lot\n2) during prototyping you can leave code as is, however save the prepared dataset to disk and load it during training, with the fast enough disk (ssd/m2) it'll be very cheap, and you'll observe huge speed-up",
      "votes": null
    },
    {
      "id": "2226916",
      "postDate": "04/19/2023 11:07:35",
      "content": "<p>I tried changing the data processing method from numpy to tensorflow. I can't say it really helped. I do not want to create any new datasets from the processed data, so I will try the shard dataset method and save it in tensor flow format </p>",
      "rawMarkdown": "I tried changing the data processing method from numpy to tensorflow. I can't say it really helped. I do not want to create any new datasets from the processed data, so I will try the shard dataset method and save it in tensor flow format",
      "votes": null
    },
    {
      "id": "2228950",
      "postDate": "04/21/2023 00:49:26",
      "content": "<p>I would recommend optimizing the code a bit and maybe trying to utilize TensorFlow in the whole preprocessing workflow as TensorFlow greatly improves data manipulation speed (if used correctly).</p>\n<p>Also, take advantage of vectorization in general. You probably can speed up the whole process a bit just by eliminating the loops</p>\n<p>The Devastator.</p>",
      "rawMarkdown": "I would recommend optimizing the code a bit and maybe trying to utilize TensorFlow in the whole preprocessing workflow as TensorFlow greatly improves data manipulation speed (if used correctly).\n\nAlso, take advantage of vectorization in general. You probably can speed up the whole process a bit just by eliminating the loops\n\n\nThe Devastator.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2221073,
      "author_name": "clearwaterkzk",
      "author_url": "",
      "post_date": "04/14/2023 01:03:09",
      "content": "<p>You are preprocessing data by numpy  and converting data into tensorflow right?<br>\nConverting data from numpy into tensorflow may be very very slow, so you should rewrite your preprocessing layer all by tensorflow.</p>\n<p>Anyway, we must have preprocessing layer in our tflite model, so we can't use numpy I think.<br>\nConverting Preprocessing layer written by numpy into tflite model failed in my case.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2226916,
          "author_name": "ch3rkasov",
          "author_url": "",
          "post_date": "04/19/2023 11:07:35",
          "content": "<p>I tried changing the data processing method from numpy to tensorflow. I can't say it really helped. I do not want to create any new datasets from the processed data, so I will try the shard dataset method and save it in tensor flow format </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2221220,
      "author_name": "martynoveduard",
      "author_url": "",
      "post_date": "04/14/2023 04:56:48",
      "content": "<p>check nvidia-smi gpu utilization/power consumption during training, <br>\nfrom the looks of it your batch generation is a huge bottleneck, <br>\nwhat you can do<br>\n1) remove python loops and convert the code to numpy/tf only ops, it will speed up things a lot<br>\n2) during prototyping you can leave code as is, however save the prepared dataset to disk and load it during training, with the fast enough disk (ssd/m2) it'll be very cheap, and you'll observe huge speed-up</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2228950,
      "author_name": "thedevastator",
      "author_url": "",
      "post_date": "04/21/2023 00:49:26",
      "content": "<p>I would recommend optimizing the code a bit and maybe trying to utilize TensorFlow in the whole preprocessing workflow as TensorFlow greatly improves data manipulation speed (if used correctly).</p>\n<p>Also, take advantage of vectorization in general. You probably can speed up the whole process a bit just by eliminating the loops</p>\n<p>The Devastator.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2220496": "Good afternoon.\nI have 10 epochs trained in about 10 hours. So slow. I suspect it's because of the way the data is processed. But it is not exactly\n\n### Here is the code for loading the file and converting it to Numpy ndarray:\n```python\ndef parquet_to_numpy(pq_path):\n    data = pd.read_parquet(os.path.join('/kaggle/input/asl-signs/',pq_path), columns=PARQUET_RELEVANT_ROWS).to_numpy()\n    data = data.reshape(len(data)//PARQUET_ROWS_PER_FRAME, PARQUET_ROWS_PER_FRAME, len(PARQUET_RELEVANT_ROWS))\n    data = PrepareFrames(data)\n    return data\n```\n\n### This is the code that I use to determine the dominant hand, remove skipping frames (without hands and face), and also bring all the data to the same number of frames:\n\n```python\ndef PrepareFrames(frames):\n    # Determining the dominant hand\n    left_hand_dominate = True\n    if np.count_nonzero(~np.isnan(\n            frames[:,np.array(important_indexes[:21]).astype(int),:].flatten()\n        )) < \\\n        np.count_nonzero(~np.isnan(\n            frames[:,np.array(important_indexes[-21:]).astype(int),:].flatten()\n        )):\n        left_hand_dominate = False\n        frames[:,468:489] = frames[:,522:,]\n        frames = frames[:,:522,]\n    else:\n        frames = frames[:,:522,]\n   \n    # Creating a new ndarray without dropping frames\n    non_nan_frames = np.array(\n        [frame for frame in frames \n         if (~np.isnan(frame[np.array(important_indexes[:28]),:]).all(axis=1)).all()\n        ]\n    )\n    # There is a gag, because I haven't found a solution for skipping empty files yet\n    if len(non_nan_frames) == 0:\n        non_nan_frames = np.random.uniform(low=0.1, high=2, size=(20, 522,3))#np.zeros((20, 522,3))\n    \n    # Here I trim the number of frames if it is greater than LIMIT_MAX_FRAMES or duplicate the last frames if it is less\n    if non_nan_frames.shape[0] > LIMIT_MAX_FRAMES:\n        interval = np.linspace(0, non_nan_frames.shape[0]-1, LIMIT_MAX_FRAMES, dtype=int)\n        non_nan_frames = non_nan_frames[np.array(interval),:,:]\n    else:\n        tiling = np.tile(non_nan_frames[-1], (LIMIT_MAX_FRAMES-non_nan_frames.shape[0], 1)).reshape(-1,522,3)\n        non_nan_frames = np.vstack([non_nan_frames,tiling])\n\n    return non_nan_frames\n```\n\n### There is also a method by which I extract features from the data.\n\n```python\nclass SignEmbedder(object):\n    \"\"\"\n    \"\"\"\n    def __init__(self, torso_size_multiplier=2.5):\n        self._torso_size_multiplier = torso_size_multiplier\n  \n        embedding = np.array([\n            self._get_angle_by_names(landmarks, 'left_hand-1', 'left_hand-2', 'left_hand-2', 'left_hand-3', degrees=False),\n            self._get_distance(\n                                self._get_coordinates_by_name(landmarks, 'left_hand-4'),\n                                self._get_coordinates_by_name(landmarks, 'left_hand-0')\n                              )\n        ])\n        return embedding.T\n    def _get_average_by_names(self, landmarks, landmark_from, landmark_to):\n        lmk_from = landmarks[:, landmark_indexes[landmark_from],:]\n        lmk_to = landmarks[:, landmark_indexes[landmark_to],:]\n        return (lmk_from + lmk_to) * 0.5\n    \n    def _get_distance(self, lmk_from, lmk_to):\n        return [np.sqrt(np.sum((p1 - p2)**2)) for p1, p2 in zip(lmk_from, lmk_to)]\n    \n    def _get_coordinates_by_name(self, landmarks, landmark):\n        return landmarks[:, landmark_indexes[landmark],:]\n    \n    def _get_angle_by_names(self,landmarks,landmark_1, landmark_2,landmark_3, landmark_4, degrees = True):\n        return self._get_angle(\n            landmarks[:, landmark_indexes[landmark_1],:],\n            landmarks[:, landmark_indexes[landmark_2],:],\n            landmarks[:, landmark_indexes[landmark_3],:],\n            landmarks[:, landmark_indexes[landmark_4],:],\n            degrees\n        )\n    \n    def _get_angle(self,lmk_1, lmk_2, lmk_3, lmk_4, degrees = True):\n        a = np.subtract(lmk_1,lmk_2)\n        b = np.subtract(lmk_3,lmk_4)\n        \n        dot_products = [sum([v1[i] * v2[i] for i in range(len(v1))]) for v1, v2 in zip(a,b)]\n        mag1s = [math.sqrt(sum([x**2 for x in v1])) for v1, _ in zip(a,b)]\n        mag2s = [math.sqrt(sum([x**2 for x in v2])) for _, v2 in zip(a,b)]\n        cos_thetas = [dot_product / (mag1 * mag2) for dot_product, mag1, mag2 in zip(dot_products, mag1s, mag2s)]\n        thetas = [math.acos(cos_theta) for cos_theta in cos_thetas]\n        angles = [math.degrees(theta) for theta in thetas]\n        if degrees:\n            return angles\n        else:\n            return thetas\n        \nsign_embedder = SignEmbedder()\n```\n\n### Loading data into Tensorflow DataSet\n\n```python\ndef tf_get_features(ftensor):\n    def feat_wrapper(ftensor):\n        return sign_embedder(parquet_to_numpy(ftensor.numpy().decode('utf-8')))\n    return tf.py_function(\n        feat_wrapper,\n        [ftensor],\n        Tout=tf.float32\n    )\ndef set_shape(x):\n    \n    # None dimensions can be of any length\n    return tf.ensure_shape(x, ( LIMIT_MAX_FRAMES, 37))\n\nX = tf.data.Dataset.from_tensor_slices(\n    train_df.path.values\n).map(\n    tf_get_features\n).map(\n    set_shape\n)\n\n```\n\n### Here are some performance tests\n\n```python\n%%timeit\n\nBATCH_SIZE = 512\n\nfor path in train_df.path.values[:BATCH_SIZE]:\n    data = sign_embedder(parquet_to_numpy(path))\n```\n\n>6.28 s ± 265 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)",
    "2221073": "You are preprocessing data by numpy  and converting data into tensorflow right?\nConverting data from numpy into tensorflow may be very very slow, so you should rewrite your preprocessing layer all by tensorflow.\n\nAnyway, we must have preprocessing layer in our tflite model, so we can't use numpy I think.\nConverting Preprocessing layer written by numpy into tflite model failed in my case.",
    "2221220": "check nvidia-smi gpu utilization/power consumption during training, \nfrom the looks of it your batch generation is a huge bottleneck, \nwhat you can do\n1) remove python loops and convert the code to numpy/tf only ops, it will speed up things a lot\n2) during prototyping you can leave code as is, however save the prepared dataset to disk and load it during training, with the fast enough disk (ssd/m2) it'll be very cheap, and you'll observe huge speed-up",
    "2226916": "I tried changing the data processing method from numpy to tensorflow. I can't say it really helped. I do not want to create any new datasets from the processed data, so I will try the shard dataset method and save it in tensor flow format",
    "2228950": "I would recommend optimizing the code a bit and maybe trying to utilize TensorFlow in the whole preprocessing workflow as TensorFlow greatly improves data manipulation speed (if used correctly).\n\nAlso, take advantage of vectorization in general. You probably can speed up the whole process a bit just by eliminating the loops\n\n\nThe Devastator."
  },
  "source": "meta"
}