{
  "id": 392381,
  "title": "Drop the empty hand?",
  "url": "/competitions/asl-signs/discussion/392381",
  "author_name": "",
  "post_date": "2023-03-05T03:31:31.861017500Z",
  "votes": 3,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi,</p>\n<p>We've seen that in about 96% of the data, one of the hands is NaN. I've been working on a model where I drop the NaN hand instead of converting it to zeros.</p>\n<p>It goes pretty much like:</p>\n<pre><code>left_hand_inputs = inputs[:, :, 468:489, :]\nright_hand_inputs = inputs[:, :,522:,:]\nhand_inputs = tf.where(tf.reduce_all(tf.equal(left_hand_inputs, 0)), right_hand_inputs, left_hand_inputs)\n</code></pre>\n<p>First I extract the hands, and then I load into hand_inputs the one hand that is not NaN.</p>\n<p>This, of course, in the 4% of the cases, will always bring the left_hand_inputs.</p>\n<p>Result have not been too good so far, but I think it may be due to other issues in the model.</p>\n<p>Anyone has any thoughts / reactions on this?</p>\n<p>Thanks!</p>",
  "messages": [
    {
      "id": "2169345",
      "postDate": "03/05/2023 03:31:31",
      "content": "<p>Hi,</p>\n<p>We've seen that in about 96% of the data, one of the hands is NaN. I've been working on a model where I drop the NaN hand instead of converting it to zeros.</p>\n<p>It goes pretty much like:</p>\n<pre><code>left_hand_inputs = inputs[:, :, 468:489, :]\nright_hand_inputs = inputs[:, :,522:,:]\nhand_inputs = tf.where(tf.reduce_all(tf.equal(left_hand_inputs, 0)), right_hand_inputs, left_hand_inputs)\n</code></pre>\n<p>First I extract the hands, and then I load into hand_inputs the one hand that is not NaN.</p>\n<p>This, of course, in the 4% of the cases, will always bring the left_hand_inputs.</p>\n<p>Result have not been too good so far, but I think it may be due to other issues in the model.</p>\n<p>Anyone has any thoughts / reactions on this?</p>\n<p>Thanks!</p>",
      "rawMarkdown": "Hi,\n\nWe've seen that in about 96% of the data, one of the hands is NaN. I've been working on a model where I drop the NaN hand instead of converting it to zeros.\n\nIt goes pretty much like:\n\n    left_hand_inputs = inputs[:, :, 468:489, :]\n    right_hand_inputs = inputs[:, :,522:,:]\n    hand_inputs = tf.where(tf.reduce_all(tf.equal(left_hand_inputs, 0)), right_hand_inputs, left_hand_inputs)\n\nFirst I extract the hands, and then I load into hand_inputs the one hand that is not NaN.\n\nThis, of course, in the 4% of the cases, will always bring the left_hand_inputs.\n\nResult have not been too good so far, but I think it may be due to other issues in the model.\n\nAnyone has any thoughts / reactions on this?\n\nThanks!",
      "votes": null
    },
    {
      "id": "2169373",
      "postDate": "03/05/2023 04:07:35",
      "content": "<p>One challenge is ensuring equivalence. Is it as simple as flipping the x-axis? For all the data if using left hand, or only the hand? What if the there's 7 frames with right hand, 2 frames with left hand?</p>\n<p>It is a good idea, but may not work without at least a little extra tweaking. </p>",
      "rawMarkdown": "One challenge is ensuring equivalence. Is it as simple as flipping the x-axis? For all the data if using left hand, or only the hand? What if the there's 7 frames with right hand, 2 frames with left hand?\n\nIt is a good idea, but may not work without at least a little extra tweaking.",
      "votes": null
    },
    {
      "id": "2169842",
      "postDate": "03/05/2023 13:39:59",
      "content": "<p>What I am doing is picking the active hand frame by frame. From my review, I could not see cases where the sign would contain one hand and then the other.  There are videos with both hands, though, and that may cause noise. I will share any outcomes of this experiment.</p>",
      "rawMarkdown": "What I am doing is picking the active hand frame by frame. From my review, I could not see cases where the sign would contain one hand and then the other.  There are videos with both hands, though, and that may cause noise. I will share any outcomes of this experiment.",
      "votes": null
    },
    {
      "id": "2169937",
      "postDate": "03/05/2023 15:06:27",
      "content": "<p>Frame by frame makes sense. But it might be harder for the model if you have right hand and left hand using the same input, which means same NN model parameters?</p>\n<p>You could compare your result with keeping active hand per frame, but storing in the right hand location if most frames have right hand, and storing in the left hand location if most frames have left hand. </p>\n<p>And one signer in train might be using both hands, that signer has a ton of noisy hand data, so I'm not sure. </p>",
      "rawMarkdown": "Frame by frame makes sense. But it might be harder for the model if you have right hand and left hand using the same input, which means same NN model parameters?\n\nYou could compare your result with keeping active hand per frame, but storing in the right hand location if most frames have right hand, and storing in the left hand location if most frames have left hand. \n\nAnd one signer in train might be using both hands, that signer has a ton of noisy hand data, so I'm not sure.",
      "votes": null
    },
    {
      "id": "2170680",
      "postDate": "03/06/2023 07:42:09",
      "content": "<p>I haven't tried this, but I did try something similar that gave me a 5 percentage point boost on the public LB.</p>\n<ol>\n<li>First, when calculating mean &amp; standard deviation, properly account for NaNs.  (Instead of Lonnie's original approach which was to set all the NaNs to 0 at the start, I basically implemented a TF equivalent of np.nanmean and np.nanstd.)  I think Robert's public notebook also does this, so you can look at that for the code.  (My implementation is different and I haven't measured which is faster yet, but I don't think I'm near the time limit anyway, so probably doesn't matter.)</li>\n<li>As you noted, some signs are NaN in all frames.  Alongside the mean plane and the standard deviation plane, I created an additional indicator plane that was 1 if there were any non-NaN values for that landmark and 0 if all values were NaN.</li>\n</ol>",
      "rawMarkdown": "I haven't tried this, but I did try something similar that gave me a 5 percentage point boost on the public LB.\n\n1. First, when calculating mean & standard deviation, properly account for NaNs.  (Instead of Lonnie's original approach which was to set all the NaNs to 0 at the start, I basically implemented a TF equivalent of np.nanmean and np.nanstd.)  I think Robert's public notebook also does this, so you can look at that for the code.  (My implementation is different and I haven't measured which is faster yet, but I don't think I'm near the time limit anyway, so probably doesn't matter.)\n1. As you noted, some signs are NaN in all frames.  Alongside the mean plane and the standard deviation plane, I created an additional indicator plane that was 1 if there were any non-NaN values for that landmark and 0 if all values were NaN.",
      "votes": null
    },
    {
      "id": "2170991",
      "postDate": "03/06/2023 12:48:07",
      "content": "<p>Thank you for the feedback Andrew! I'll try these hints :)</p>",
      "rawMarkdown": "Thank you for the feedback Andrew! I'll try these hints :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2169373,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "03/05/2023 04:07:35",
      "content": "<p>One challenge is ensuring equivalence. Is it as simple as flipping the x-axis? For all the data if using left hand, or only the hand? What if the there's 7 frames with right hand, 2 frames with left hand?</p>\n<p>It is a good idea, but may not work without at least a little extra tweaking. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2169842,
          "author_name": "juanolano",
          "author_url": "",
          "post_date": "03/05/2023 13:39:59",
          "content": "<p>What I am doing is picking the active hand frame by frame. From my review, I could not see cases where the sign would contain one hand and then the other.  There are videos with both hands, though, and that may cause noise. I will share any outcomes of this experiment.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2169937,
              "author_name": "roberthatch",
              "author_url": "",
              "post_date": "03/05/2023 15:06:27",
              "content": "<p>Frame by frame makes sense. But it might be harder for the model if you have right hand and left hand using the same input, which means same NN model parameters?</p>\n<p>You could compare your result with keeping active hand per frame, but storing in the right hand location if most frames have right hand, and storing in the left hand location if most frames have left hand. </p>\n<p>And one signer in train might be using both hands, that signer has a ton of noisy hand data, so I'm not sure. </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2170680,
      "author_name": "andrewrrose",
      "author_url": "",
      "post_date": "03/06/2023 07:42:09",
      "content": "<p>I haven't tried this, but I did try something similar that gave me a 5 percentage point boost on the public LB.</p>\n<ol>\n<li>First, when calculating mean &amp; standard deviation, properly account for NaNs.  (Instead of Lonnie's original approach which was to set all the NaNs to 0 at the start, I basically implemented a TF equivalent of np.nanmean and np.nanstd.)  I think Robert's public notebook also does this, so you can look at that for the code.  (My implementation is different and I haven't measured which is faster yet, but I don't think I'm near the time limit anyway, so probably doesn't matter.)</li>\n<li>As you noted, some signs are NaN in all frames.  Alongside the mean plane and the standard deviation plane, I created an additional indicator plane that was 1 if there were any non-NaN values for that landmark and 0 if all values were NaN.</li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 2170991,
          "author_name": "juanolano",
          "author_url": "",
          "post_date": "03/06/2023 12:48:07",
          "content": "<p>Thank you for the feedback Andrew! I'll try these hints :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2169345": "Hi,\n\nWe've seen that in about 96% of the data, one of the hands is NaN. I've been working on a model where I drop the NaN hand instead of converting it to zeros.\n\nIt goes pretty much like:\n\n    left_hand_inputs = inputs[:, :, 468:489, :]\n    right_hand_inputs = inputs[:, :,522:,:]\n    hand_inputs = tf.where(tf.reduce_all(tf.equal(left_hand_inputs, 0)), right_hand_inputs, left_hand_inputs)\n\nFirst I extract the hands, and then I load into hand_inputs the one hand that is not NaN.\n\nThis, of course, in the 4% of the cases, will always bring the left_hand_inputs.\n\nResult have not been too good so far, but I think it may be due to other issues in the model.\n\nAnyone has any thoughts / reactions on this?\n\nThanks!",
    "2169373": "One challenge is ensuring equivalence. Is it as simple as flipping the x-axis? For all the data if using left hand, or only the hand? What if the there's 7 frames with right hand, 2 frames with left hand?\n\nIt is a good idea, but may not work without at least a little extra tweaking.",
    "2169842": "What I am doing is picking the active hand frame by frame. From my review, I could not see cases where the sign would contain one hand and then the other.  There are videos with both hands, though, and that may cause noise. I will share any outcomes of this experiment.",
    "2169937": "Frame by frame makes sense. But it might be harder for the model if you have right hand and left hand using the same input, which means same NN model parameters?\n\nYou could compare your result with keeping active hand per frame, but storing in the right hand location if most frames have right hand, and storing in the left hand location if most frames have left hand. \n\nAnd one signer in train might be using both hands, that signer has a ton of noisy hand data, so I'm not sure.",
    "2170680": "I haven't tried this, but I did try something similar that gave me a 5 percentage point boost on the public LB.\n\n1. First, when calculating mean & standard deviation, properly account for NaNs.  (Instead of Lonnie's original approach which was to set all the NaNs to 0 at the start, I basically implemented a TF equivalent of np.nanmean and np.nanstd.)  I think Robert's public notebook also does this, so you can look at that for the code.  (My implementation is different and I haven't measured which is faster yet, but I don't think I'm near the time limit anyway, so probably doesn't matter.)\n1. As you noted, some signs are NaN in all frames.  Alongside the mean plane and the standard deviation plane, I created an additional indicator plane that was 1 if there were any non-NaN values for that landmark and 0 if all values were NaN.",
    "2170991": "Thank you for the feedback Andrew! I'll try these hints :)"
  },
  "source": "meta"
}