{
  "id": 391265,
  "title": "[LB0.76 one fold !!!! ] my pytorch transformer experiment results",
  "url": "/competitions/asl-signs/discussion/391265",
  "author_name": "hengck23",
  "post_date": "2023-02-28T21:05:45.123000",
  "votes": 128,
  "comment_count": 180,
  "views": 0,
  "content": "<p>experiments in progress …</p>\n<p><img src=\"https://i.ibb.co/9TvfrhW/Selection-999-1855.png\" alt=\"https://i.ibb.co/9TvfrhW/Selection-999-1855.png\"></p>\n<hr>\n<p>previous results:<br>\n<a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/391265#2197663\" target=\"_blank\">https://www.kaggle.com/competitions/asl-signs/discussion/391265#2197663</a></p>\n<hr>\n<p>after some experiments for the first week, here is the master plan:</p>\n<p>[1] one (or two) layer transformer is sufficient. for one fold, memory is only 2mb ( tf.lite.Optimize.DEFAULT, i.e. int8 in storage) and 25min for all 40_000 test videos (or 30 msec per video). Future because of missing frame, long term attention has its advantage</p>\n<p>in fact two layer transformer would lead to over fitting.</p>\n<p>[2] next step is to spend more time on data representation, especially normalization</p>\n<p>[3] use of external data, maybe need to use mediapipe to extract landmarks on external ASL video. <br>\nOther ways to use increase data increases, adversarial/diffusion augmentation, masked token prediction, etc</p>\n<p>I feel that SSL with masked token prediction (e.g. generative pretraining) fits this problem well as there are high proportion of missing frames.</p>\n<p>[4], There no much of algorithm design because network is quite shallow. if there is, it would be more on increasing parameters for accuracy and speeds. There are many VIT tricks that can be used here, e.g. window attention, poolforming, convolutional positional encoding, grouped conv, …</p>\n<hr>\n<p>First results </p>\n<ul>\n<li>one fold, max video length=60 (crop first) : LB 0.65 </li>\n<li>1 fold, max video length=512 (crop center) :  LB  0.65  25 min</li>\n<li>4 fold, max video length=512 (crop center) :  LB  0.68  32 min</li>\n<li>5 fold, max video length=512 (crop center) :  LB  (better) 0.68  35 min</li>\n</ul>\n<p>CV one fold using train/validation random split:  cv 0.72~0.74 (probably the limit if good generalisation)<br>\nCV one fold using train/validation participant non-overlap split:  cv 0.61 (compared with above, show the severity of lack of train data) </p>\n<p>inference:<br>\n<a href=\"https://www.kaggle.com/code/hengck23/lb-0-65-one-pytorch-transformer-solution\" target=\"_blank\">https://www.kaggle.com/code/hengck23/lb-0-65-one-pytorch-transformer-solution</a></p>\n<p>training code:<br>\n<a href=\"https://www.kaggle.com/datasets/hengck23/asl-demo\" target=\"_blank\">https://www.kaggle.com/datasets/hengck23/asl-demo</a></p>",
  "messages": [
    {
      "id": 2163485,
      "postDate": "2023-02-28T21:05:45.123Z",
      "content": "<p>experiments in progress …</p>\n<p><img src=\"https://i.ibb.co/9TvfrhW/Selection-999-1855.png\" alt=\"https://i.ibb.co/9TvfrhW/Selection-999-1855.png\"></p>\n<hr>\n<p>previous results:<br>\n<a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/391265#2197663\" target=\"_blank\">https://www.kaggle.com/competitions/asl-signs/discussion/391265#2197663</a></p>\n<hr>\n<p>after some experiments for the first week, here is the master plan:</p>\n<p>[1] one (or two) layer transformer is sufficient. for one fold, memory is only 2mb ( tf.lite.Optimize.DEFAULT, i.e. int8 in storage) and 25min for all 40_000 test videos (or 30 msec per video). Future because of missing frame, long term attention has its advantage</p>\n<p>in fact two layer transformer would lead to over fitting.</p>\n<p>[2] next step is to spend more time on data representation, especially normalization</p>\n<p>[3] use of external data, maybe need to use mediapipe to extract landmarks on external ASL video. <br>\nOther ways to use increase data increases, adversarial/diffusion augmentation, masked token prediction, etc</p>\n<p>I feel that SSL with masked token prediction (e.g. generative pretraining) fits this problem well as there are high proportion of missing frames.</p>\n<p>[4], There no much of algorithm design because network is quite shallow. if there is, it would be more on increasing parameters for accuracy and speeds. There are many VIT tricks that can be used here, e.g. window attention, poolforming, convolutional positional encoding, grouped conv, …</p>\n<hr>\n<p>First results </p>\n<ul>\n<li>one fold, max video length=60 (crop first) : LB 0.65 </li>\n<li>1 fold, max video length=512 (crop center) :  LB  0.65  25 min</li>\n<li>4 fold, max video length=512 (crop center) :  LB  0.68  32 min</li>\n<li>5 fold, max video length=512 (crop center) :  LB  (better) 0.68  35 min</li>\n</ul>\n<p>CV one fold using train/validation random split:  cv 0.72~0.74 (probably the limit if good generalisation)<br>\nCV one fold using train/validation participant non-overlap split:  cv 0.61 (compared with above, show the severity of lack of train data) </p>\n<p>inference:<br>\n<a href=\"https://www.kaggle.com/code/hengck23/lb-0-65-one-pytorch-transformer-solution\" target=\"_blank\">https://www.kaggle.com/code/hengck23/lb-0-65-one-pytorch-transformer-solution</a></p>\n<p>training code:<br>\n<a href=\"https://www.kaggle.com/datasets/hengck23/asl-demo\" target=\"_blank\">https://www.kaggle.com/datasets/hengck23/asl-demo</a></p>",
      "rawMarkdown": "experiments in progress ...\n\n![https://i.ibb.co/9TvfrhW/Selection-999-1855.png](https://i.ibb.co/9TvfrhW/Selection-999-1855.png)\n\n---\n\nprevious results:\nhttps://www.kaggle.com/competitions/asl-signs/discussion/391265#2197663\n\n\n---\n\nafter some experiments for the first week, here is the master plan:\n\n[1] one (or two) layer transformer is sufficient. for one fold, memory is only 2mb ( tf.lite.Optimize.DEFAULT, i.e. int8 in storage) and 25min for all 40\\_000 test videos (or 30 msec per video). Future because of missing frame, long term attention has its advantage\n\nin fact two layer transformer would lead to over fitting.\n\n[2] next step is to spend more time on data representation, especially normalization\n\n[3] use of external data, maybe need to use mediapipe to extract landmarks on external ASL video. \nOther ways to use increase data increases, adversarial/diffusion augmentation, masked token prediction, etc\n\nI feel that SSL with masked token prediction (e.g. generative pretraining) fits this problem well as there are high proportion of missing frames.\n\n[4], There no much of algorithm design because network is quite shallow. if there is, it would be more on increasing parameters for accuracy and speeds. There are many VIT tricks that can be used here, e.g. window attention, poolforming, convolutional positional encoding, grouped conv, ...\n\n\n----\nFirst results \n- one fold, max video length=60 (crop first) : LB 0.65 \n- 1 fold, max video length=512 (crop center) :  LB  0.65  25 min\n- 4 fold, max video length=512 (crop center) :  LB  0.68  32 min\n- 5 fold, max video length=512 (crop center) :  LB  (better) 0.68  35 min\n\nCV one fold using train/validation random split:  cv 0.72~0.74 (probably the limit if good generalisation)\nCV one fold using train/validation participant non-overlap split:  cv 0.61 (compared with above, show the severity of lack of train data) \n\ninference:\nhttps://www.kaggle.com/code/hengck23/lb-0-65-one-pytorch-transformer-solution\n\ntraining code:\nhttps://www.kaggle.com/datasets/hengck23/asl-demo\n\n\n",
      "votes": 126
    },
    {
      "id": 2179326,
      "postDate": "2023-03-13T04:40:49.920Z",
      "content": "<p>i have been playing with controlnet and i come across this:<br>\n<img src=\"https://i.ibb.co/0cdwBSL/Selection-999-1366.png\" alt=\"https://i.ibb.co/0cdwBSL/Selection-999-1366.png\"></p>\n<p>maybe can make some fun animation video while waiting for training</p>",
      "rawMarkdown": "i have been playing with controlnet and i come across this:\n![https://i.ibb.co/0cdwBSL/Selection-999-1366.png](https://i.ibb.co/0cdwBSL/Selection-999-1366.png)\n\nmaybe can make some fun animation video while waiting for training",
      "votes": 14,
      "replies": [
        {
          "id": 2180654,
          "postDate": "2023-03-14T02:26:27.493Z",
          "content": "<p>landmark --&gt; controlnet --&gt;stable diffusion --&gt; image+style+augment --&gt; media pose --&gt;landmark+augment ???</p>\n<p>note that you can have depthmap as control etc and you can change viewpoint, etc</p>",
          "rawMarkdown": "landmark --> controlnet -->stable diffusion --> image+style+augment --> media pose -->landmark+augment ???\n\nnote that you can have depthmap as control etc and you can change viewpoint, etc",
          "votes": 2
        },
        {
          "id": 2202379,
          "postDate": "2023-03-30T02:22:09.557Z",
          "content": "<p>controlnet based on face landmark<br>\n<a href=\"https://huggingface.co/georgefen/Face-Landmark-ControlNet\" target=\"_blank\">https://huggingface.co/georgefen/Face-Landmark-ControlNet</a></p>",
          "rawMarkdown": "controlnet based on face landmark\nhttps://huggingface.co/georgefen/Face-Landmark-ControlNet"
        },
        {
          "id": 2483568,
          "postDate": "2023-10-15T19:40:02.877Z",
          "content": "<p>Hi, that is really interesting. I am guessing this will create images frame by frame and we merge all the frames to create a video. Can you give me gist of it? I'm still learning, but seems like a fun learning project. Thank you!</p>",
          "rawMarkdown": "Hi, that is really interesting. I am guessing this will create images frame by frame and we merge all the frames to create a video. Can you give me gist of it? I'm still learning, but seems like a fun learning project. Thank you!"
        }
      ]
    },
    {
      "id": 2201909,
      "postDate": "2023-03-29T16:19:45.363Z",
      "content": "<p>flip augmentation magic:<br>\nCV 0.703 (single fold)<br>\nLB 0.72</p>\n<pre><code>LHAND = np.arange(468, 489).tolist() # 21\nRHAND = np.arange(522, 543).tolist() # 21\nPOSE  = np.arange(489, 522).tolist() # 33\nFACE  = np.arange(0,468).tolist()    #468\n\nREYE = [\n    33, 7, 163, 144, 145, 153, 154, 155, 133,\n    246, 161, 160, 159, 158, 157, 173,\n]\nLEYE = [\n    263, 249, 390, 373, 374, 380, 381, 382, 362,\n    466, 388, 387, 386, 385, 384, 398,\n]\nNOSE=[\n    1,2,98,327\n]\nSLIP = [\n    78, 95, 88, 178, 87, 14, 317, 402, 318, 324, 308,\n    191, 80, 81, 82, 13, 312, 311, 310, 415,\n]\nSPOSE = (np.array([\n    11,13,15,12,14,16,23,24,\n])+489).tolist()\n\n\n...\n## assume zero mean xyz. so flip can be implement by multuiplication of -1\ndef do_hflip_hand(lhand, rhand):\n    rhand[...,0] *= -1\n    lhand[...,0] *= -1\n    rhand, lhand = lhand,rhand\n    return lhand, rhand\n\ndef do_hflip_spose(spose):\n    spose[...,0] *= -1\n    spose = spose[:,[3,4,5,0,1,2,7,6]]\n    return spose\n\ndef do_hflip_slip(slip):\n    slip[...,0] *= -1\n    slip = slip[:,[10,9,8,7,6,5,4,3,2,1,0]+[19,18,17,16,15,14,13,12,11]]\n    return slip\n\n...\n...\n        xyz = load_relevant_data_subset(pq_file)\n        lhand = xyz[:,LHAND]\n        rhand = xyz[:,RHAND]\n        spose = xyz[:,SPOSE]\n        leye = xyz[:,LEYE]\n        reye = xyz[:,REYE]\n        slip = xyz[:,SLIP]\n        nose = xyz[:,NOSE]\n\n\n    if is_aug==1:\n        if np.random.rand()&lt;0.5:\n            lhand, rhand = do_hflip_hand(lhand, rhand)\n            spose = do_hflip_spose(spose)\n            leye, reye = do_hflip_eye(leye, reye)\n            slip = do_hflip_slip(slip)\n            nose = do_hflip_nose(nose)\n</code></pre>\n<p>train log:</p>\n<pre><code>fold_type = kaggle-part\nfold = 0\ntrain_dataset : \n    len = 71518\n    num_participant_id = 16\n\nvalid_dataset : \n    len = 22959\n    num_participant_id = 5\n...\n   batch_size = 64 \n   experiment = ['tx025-no-z-flip', 'run_train_fold0.py']\n                           |---------------- VALID---------|---- TRAIN/BATCH ----------------------\nrate      iter       epoch | loss   top1   top2    top5    | loss                 | time           \n---------------------------------------------------------------------------------------------------\n...\n1.00e-4   00050265*  45.00 | 2.025  0.698  0.7992  0.870   | 3.887  0.000  0.000  |  1 hr 48 min\n1.00e-4   00053616*  48.00 | 2.042  0.697  0.7969  0.870   | 3.879  0.000  0.000  |  1 hr 56 min\n</code></pre>",
      "rawMarkdown": "flip augmentation magic:\nCV 0.703 (single fold)\nLB 0.72\n\n```\nLHAND = np.arange(468, 489).tolist() # 21\nRHAND = np.arange(522, 543).tolist() # 21\nPOSE  = np.arange(489, 522).tolist() # 33\nFACE  = np.arange(0,468).tolist()    #468\n\nREYE = [\n\t33, 7, 163, 144, 145, 153, 154, 155, 133,\n\t246, 161, 160, 159, 158, 157, 173,\n]\nLEYE = [\n\t263, 249, 390, 373, 374, 380, 381, 382, 362,\n\t466, 388, 387, 386, 385, 384, 398,\n]\nNOSE=[\n\t1,2,98,327\n]\nSLIP = [\n\t78, 95, 88, 178, 87, 14, 317, 402, 318, 324, 308,\n\t191, 80, 81, 82, 13, 312, 311, 310, 415,\n]\nSPOSE = (np.array([\n\t11,13,15,12,14,16,23,24,\n])+489).tolist()\n\n\n...\n## assume zero mean xyz. so flip can be implement by multuiplication of -1\ndef do_hflip_hand(lhand, rhand):\n\trhand[...,0] *= -1\n\tlhand[...,0] *= -1\n\trhand, lhand = lhand,rhand\n\treturn lhand, rhand\n\ndef do_hflip_spose(spose):\n\tspose[...,0] *= -1\n\tspose = spose[:,[3,4,5,0,1,2,7,6]]\n\treturn spose\n\ndef do_hflip_slip(slip):\n\tslip[...,0] *= -1\n\tslip = slip[:,[10,9,8,7,6,5,4,3,2,1,0]+[19,18,17,16,15,14,13,12,11]]\n\treturn slip\n\n...\n...\n\t\txyz = load_relevant_data_subset(pq_file)\n\t\tlhand = xyz[:,LHAND]\n\t\trhand = xyz[:,RHAND]\n\t\tspose = xyz[:,SPOSE]\n\t\tleye = xyz[:,LEYE]\n\t\treye = xyz[:,REYE]\n\t\tslip = xyz[:,SLIP]\n\t\tnose = xyz[:,NOSE]\n\n\n\tif is_aug==1:\n\t\tif np.random.rand()<0.5:\n\t\t\tlhand, rhand = do_hflip_hand(lhand, rhand)\n\t\t\tspose = do_hflip_spose(spose)\n\t\t\tleye, reye = do_hflip_eye(leye, reye)\n\t\t\tslip = do_hflip_slip(slip)\n\t\t\tnose = do_hflip_nose(nose)\n\n```\n\ntrain log:\n\n```\nfold_type = kaggle-part\nfold = 0\ntrain_dataset : \n\tlen = 71518\n\tnum_participant_id = 16\n\nvalid_dataset : \n\tlen = 22959\n\tnum_participant_id = 5\n...\n   batch_size = 64 \n   experiment = ['tx025-no-z-flip', 'run_train_fold0.py']\n                           |---------------- VALID---------|---- TRAIN/BATCH ----------------------\nrate      iter       epoch | loss   top1   top2    top5    | loss                 | time           \n---------------------------------------------------------------------------------------------------\n...\n1.00e-4   00050265*  45.00 | 2.025  0.698  0.7992  0.870   | 3.887  0.000  0.000  |  1 hr 48 min\n1.00e-4   00053616*  48.00 | 2.042  0.697  0.7969  0.870   | 3.879  0.000  0.000  |  1 hr 56 min\n\n```\n\n\n",
      "votes": 8,
      "replies": [
        {
          "id": 2202177,
          "postDate": "2023-03-29T19:49:19.067Z",
          "content": "<p>one may want to try \"cut and paste augment\": assemble body parts from different signers</p>",
          "rawMarkdown": "one may want to try \"cut and paste augment\": assemble body parts from different signers",
          "votes": 2,
          "replies": [
            {
              "id": 2202195,
              "postDate": "2023-03-29T20:06:22.660Z",
              "content": "<p>I've already tried something similar when I changed body parts. On my pipeline, the result is controversial. On the one hand, this gave an increase of 1 fold, but on the other hand, the ensemble did not improve. Of course I could have made a mistake somewhere in the code. And this reduced val_loss</p>",
              "rawMarkdown": "I've already tried something similar when I changed body parts. On my pipeline, the result is controversial. On the one hand, this gave an increase of 1 fold, but on the other hand, the ensemble did not improve. Of course I could have made a mistake somewhere in the code. And this reduced val_loss",
              "votes": 1
            },
            {
              "id": 2202204,
              "postDate": "2023-03-29T20:12:24.370Z",
              "content": "<p>the parts have to be synchronised correctly in time. a better way is to use gan or pddm</p>",
              "rawMarkdown": "the parts have to be synchronised correctly in time. a better way is to use gan or pddm",
              "votes": 1
            },
            {
              "id": 2202381,
              "postDate": "2023-03-30T02:30:20.223Z",
              "content": "<p>instead of cut and paste, imagine:</p>\n<ol>\n<li>there is a canoical signer</li>\n<li>you are use distance constrain and inverse kinemetics to transfer kaggle signer motion to this canoical signer</li>\n<li>check if we use only motion of a sinle canoical signer, classification is better</li>\n<li>think of a way to predict :  canoical signer motion given  transfer kaggle signer motion<br>\n(e.g. pre-processing or learn a simple net)</li>\n</ol>\n<p>basically, we are learning a better normalisation<br>\n<a href=\"https://www.youtube.com/watch?v=U9G2VlG9HbA\" target=\"_blank\">https://www.youtube.com/watch?v=U9G2VlG9HbA</a></p>",
              "rawMarkdown": "instead of cut and paste, imagine:\n1. there is a canoical signer\n2. you are use distance constrain and inverse kinemetics to transfer kaggle signer motion to this canoical signer\n3. check if we use only motion of a sinle canoical signer, classification is better\n4. think of a way to predict :  canoical signer motion given  transfer kaggle signer motion\n(e.g. pre-processing or learn a simple net)\n\nbasically, we are learning a better normalisation\nhttps://www.youtube.com/watch?v=U9G2VlG9HbA"
            }
          ]
        },
        {
          "id": 2202388,
          "postDate": "2023-03-30T02:48:43.213Z",
          "content": "<p>🤔Do other landmarks really help? In my experiments, event simply add pose landmarks (arms position) decrease CV.</p>",
          "rawMarkdown": "🤔Do other landmarks really help? In my experiments, event simply add pose landmarks (arms position) decrease CV.",
          "votes": 1,
          "replies": [
            {
              "id": 2202395,
              "postDate": "2023-03-30T03:08:03.123Z",
              "content": "<p>it depends on how you normalised i think.</p>\n<p>there is many noise in the data becuase each signer is of different size and their camera may have different frame rate. They also move differently.</p>\n<p>in theory:<br>\neye:  should help. for sign words  \"awake\", the eye is closed then opened.<br>\npose (arm): should help. this is global motion of hand.</p>\n<p>but if your nomralisation is poor or your model cannot filter out the noise, then there will be easy overfitting to noise</p>",
              "rawMarkdown": "it depends on how you normalised i think.\n\nthere is many noise in the data becuase each signer is of different size and their camera may have different frame rate. They also move differently.\n\nin theory:\neye:  should help. for sign words  \"awake\", the eye is closed then opened.\npose (arm): should help. this is global motion of hand.\n\nbut if your nomralisation is poor or your model cannot filter out the noise, then there will be easy overfitting to noise\n",
              "votes": 4
            },
            {
              "id": 2202398,
              "postDate": "2023-03-30T03:12:11.127Z",
              "content": "<p>yes, for me, normalization method also matters</p>",
              "rawMarkdown": "yes, for me, normalization method also matters"
            }
          ]
        },
        {
          "id": 2202578,
          "postDate": "2023-03-30T06:40:03.967Z",
          "content": "<p>this is not too difficult if you think about it</p>\n<p><img src=\"https://i.ibb.co/5FLVDDw/Selection-999-1630.png\" alt=\"https://i.ibb.co/5FLVDDw/Selection-999-1630.png\"></p>",
          "rawMarkdown": "this is not too difficult if you think about it\n\n![https://i.ibb.co/5FLVDDw/Selection-999-1630.png](https://i.ibb.co/5FLVDDw/Selection-999-1630.png)",
          "votes": 1,
          "replies": [
            {
              "id": 2202731,
              "postDate": "2023-03-30T09:06:33.303Z",
              "content": "<p>I tried a very similar approach to this, but unfortunately, I didn't see any performance improvement compared to simple augmentation(maybe my approach was too naive). However, I think it's a fascinating idea, and I'm excited to hear about your progress if you decide to give it a go:).</p>",
              "rawMarkdown": "I tried a very similar approach to this, but unfortunately, I didn't see any performance improvement compared to simple augmentation(maybe my approach was too naive). However, I think it's a fascinating idea, and I'm excited to hear about your progress if you decide to give it a go:)."
            }
          ]
        },
        {
          "id": 2229296,
          "postDate": "2023-04-21T08:27:47.767Z",
          "content": "<p>According to LANDMARKS_REFINEMENT_LEFT_EYE_CONFIG and LANDMARKS_REFINEMENT_RIGHT_EYE_CONFIG (from <a href=\"https://github.com/tensorflow/tfjs-models/blob/master/face-landmarks-detection/src/tfjs/constants.ts\" target=\"_blank\">https://github.com/tensorflow/tfjs-models/blob/master/face-landmarks-detection/src/tfjs/constants.ts</a>) , your REYE if a left eye, and LEYE is a right eye. Are you sure that's correct?</p>",
          "rawMarkdown": "According to LANDMARKS_REFINEMENT_LEFT_EYE_CONFIG and LANDMARKS_REFINEMENT_RIGHT_EYE_CONFIG (from https://github.com/tensorflow/tfjs-models/blob/master/face-landmarks-detection/src/tfjs/constants.ts) , your REYE if a left eye, and LEYE is a right eye. Are you sure that's correct?"
        }
      ]
    },
    {
      "id": 2193356,
      "postDate": "2023-03-23T08:28:19.040Z",
      "content": "<p>I love your experiments records table! thanks for sharing this!</p>",
      "rawMarkdown": "I love your experiments records table! thanks for sharing this!",
      "votes": 8,
      "replies": [
        {
          "id": 2193995,
          "postDate": "2023-03-23T16:50:13.230Z",
          "content": "<p>+1 <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> </p>",
          "rawMarkdown": "+1 @hengck23 "
        }
      ]
    },
    {
      "id": 2178946,
      "postDate": "2023-03-12T19:46:44.203Z",
      "content": "<p>[unorganized] some proven intermediate results … will write in more detail later</p>\n<ul>\n<li><p>add motion (xyz at current frame - xyz at previous frame). i.e. each point now has 6 dim. need to handle NaN motion correctly<br>\nimprove CV by 0.015</p></li>\n<li><p>add framerate interpolation augmentation scipy.interpolate.interp1d()<br>\nimprove CV by 0.015</p></li>\n</ul>",
      "rawMarkdown": "[unorganized] some proven intermediate results ... will write in more detail later\n\n- add motion (xyz at current frame - xyz at previous frame). i.e. each point now has 6 dim. need to handle NaN motion correctly\nimprove CV by 0.015\n\n\n- add framerate interpolation augmentation scipy.interpolate.interp1d()\nimprove CV by 0.015\n",
      "votes": 7,
      "replies": [
        {
          "id": 2180147,
          "postDate": "2023-03-13T16:08:00.600Z",
          "content": "<p>ideally, if there is lots and lots of training samples, you don't have to do anything, just add more parameters and the model would learn scale, rotation invariant features.</p>\n<p>if there is not much training samples, you can do augmentation, e.g. rotate/shift the hands</p>\n<p>you can also handcraft features, e.g. pairwise distance, pairwise angles  of the hand points (these are invariant to rotation)</p>\n<p>better still,  graph network CGN  or pointnet features are for this.</p>",
          "rawMarkdown": "ideally, if there is lots and lots of training samples, you don't have to do anything, just add more parameters and the model would learn scale, rotation invariant features.\n\nif there is not much training samples, you can do augmentation, e.g. rotate/shift the hands\n\nyou can also handcraft features, e.g. pairwise distance, pairwise angles  of the hand points (these are invariant to rotation)\n\nbetter still,  graph network CGN  or pointnet features are for this.\n ",
          "votes": 3,
          "replies": [
            {
              "id": 2180165,
              "postDate": "2023-03-13T16:21:10.243Z",
              "content": "<p>convergence is magical if you do things right !</p>\n<p>train/validation participant id does not overlap shown below:<br>\n<img src=\"https://i.ibb.co/g3C1Gq6/Selection-999-1379.png\" alt=\"https://i.ibb.co/g3C1Gq6/Selection-999-1379.png\"></p>\n<p>distance matrix is the poor man's graph network D x A x inv(D) …<br>\ni need to google for better joint kinematics features ….</p>",
              "rawMarkdown": "convergence is magical if you do things right !\n\ntrain/validation participant id does not overlap shown below:\n![https://i.ibb.co/g3C1Gq6/Selection-999-1379.png](https://i.ibb.co/g3C1Gq6/Selection-999-1379.png)\n\ndistance matrix is the poor man's graph network D x A x inv(D) ...\ni need to google for better joint kinematics features ....",
              "votes": 4
            },
            {
              "id": 2180629,
              "postDate": "2023-03-14T01:49:37.657Z",
              "content": "<p>i wondered if i renderd the point into 48x48 images and use 3dCNN, what will be the speed and accuracy …</p>",
              "rawMarkdown": "i wondered if i renderd the point into 48x48 images and use 3dCNN, what will be the speed and accuracy ...",
              "votes": 3
            },
            {
              "id": 2181090,
              "postDate": "2023-03-14T10:11:12.587Z",
              "content": "<p>possible of aux loss:</p>\n<ul>\n<li>predict if word is one-hand or two-hand word</li>\n<li>predict if the signer is using left, right or both hands</li>\n<li>predict signer identity</li>\n</ul>\n<p>….</p>",
              "rawMarkdown": "possible of aux loss:\n- predict if word is one-hand or two-hand word\n- predict if the signer is using left, right or both hands\n- predict signer identity\n\n....",
              "votes": 1
            },
            {
              "id": 2182554,
              "postDate": "2023-03-15T06:51:09.113Z",
              "content": "<p>I didn’t try 3D CNN yet, but I have tried 2D CNN. I use a (batch_size, num_frame, num_landmark, 3) tensor as input. I didn’t do much experiments, the best val acc I got is around 0.53. I didn’t record the speed. There are many skeleton-based motion recognition papers that use CNN, but according to the paper, they all achieve a fantastic result, which is much higher than 0.53… I guess we need to process the data in some ways to make it suitable to CNN.</p>",
              "rawMarkdown": "I didn’t try 3D CNN yet, but I have tried 2D CNN. I use a (batch_size, num_frame, num_landmark, 3) tensor as input. I didn’t do much experiments, the best val acc I got is around 0.53. I didn’t record the speed. There are many skeleton-based motion recognition papers that use CNN, but according to the paper, they all achieve a fantastic result, which is much higher than 0.53… I guess we need to process the data in some ways to make it suitable to CNN."
            },
            {
              "id": 2182562,
              "postDate": "2023-03-15T06:59:03.907Z",
              "content": "<p><a href=\"https://ibb.co/85Qp2qv\"><img src=\"https://i.ibb.co/1q45K1y/Selection-999-1426.png\" alt=\"Selection-999-1426\"></a><br>\n<a href=\"https://ibb.co/RQ2vLxC\"><img src=\"https://i.ibb.co/SVNsqhK/Selection-999-1422.png\" alt=\"Selection-999-1422\"></a></p>",
              "rawMarkdown": "<a href=\"https://ibb.co/85Qp2qv\"><img src=\"https://i.ibb.co/1q45K1y/Selection-999-1426.png\" alt=\"Selection-999-1426\" border=\"0\"></a>\n<a href=\"https://ibb.co/RQ2vLxC\"><img src=\"https://i.ibb.co/SVNsqhK/Selection-999-1422.png\" alt=\"Selection-999-1422\" border=\"0\"></a>",
              "votes": 1
            },
            {
              "id": 2182671,
              "postDate": "2023-03-15T08:41:20.820Z",
              "content": "<p><a href=\"https://www.kaggle.com/jay2333\" target=\"_blank\">@jay2333</a> Papers in the skeleton-based action recognition domain rely on high-quality pose data. In most cases, they use HRNet pose estimator to extract the skeleton landmarks. Furthermore, there are only a few missing values. </p>\n<p>The landmarks in our dataset are accurate, but just not as good as landmarks from HRNet. Plus, we have a lot of missing values and we need to find a way to deal with them. </p>",
              "rawMarkdown": "@jay2333 Papers in the skeleton-based action recognition domain rely on high-quality pose data. In most cases, they use HRNet pose estimator to extract the skeleton landmarks. Furthermore, there are only a few missing values. \n\nThe landmarks in our dataset are accurate, but just not as good as landmarks from HRNet. Plus, we have a lot of missing values and we need to find a way to deal with them. "
            },
            {
              "id": 2183125,
              "postDate": "2023-03-15T13:29:50.920Z",
              "content": "<p><a href=\"https://www.kaggle.com/gregorlied\" target=\"_blank\">@gregorlied</a> Totally agree with that. According to the current public notebook, a simple fully connected network can perform well. So the main issue is the data process part.</p>",
              "rawMarkdown": "@gregorlied Totally agree with that. According to the current public notebook, a simple fully connected network can perform well. So the main issue is the data process part.",
              "votes": 1
            },
            {
              "id": 2183239,
              "postDate": "2023-03-15T14:42:51.547Z",
              "content": "<p>\" fully connected network can perform well. \"</p>\n<p>the reason is that the seq are not really long</p>",
              "rawMarkdown": "\" fully connected network can perform well. \"\n\nthe reason is that the seq are not really long",
              "votes": 1
            },
            {
              "id": 2183267,
              "postDate": "2023-03-15T15:03:07.460Z",
              "content": "<p>there is a trick to speedup 3d CNN or 2d CNN</p>\n<p>basically it is like CLIP (aligned text and image)</p>\n<p>here you trained joint encoder for (pose, rendered image) pair.<br>\nrendered image has more information so they have better embedding.</p>\n<p>you can think of we are distilling image embedding to point embedding.<br>\nsince both embedding are \"aligned\", in testing we are only using point embedding (though we trained them jointly)</p>\n<hr>\n<p>try to use different color for part, shade for depth when rendering ….</p>",
              "rawMarkdown": "there is a trick to speedup 3d CNN or 2d CNN\n\nbasically it is like CLIP (aligned text and image)\n\nhere you trained joint encoder for (pose, rendered image) pair.\nrendered image has more information so they have better embedding.\n\nyou can think of we are distilling image embedding to point embedding.\nsince both embedding are \"aligned\", in testing we are only using point embedding (though we trained them jointly)\n\n---\ntry to use different color for part, shade for depth when rendering ....\n",
              "votes": 1
            },
            {
              "id": 2185509,
              "postDate": "2023-03-17T05:45:15.100Z",
              "content": "<p>yet  another aux loss,</p>\n<ul>\n<li>if word are noun (object) or verb (action, motion), etc</li>\n</ul>",
              "rawMarkdown": "yet  another aux loss,\n- if word are noun (object) or verb (action, motion), etc"
            },
            {
              "id": 2203259,
              "postDate": "2023-03-30T16:45:01.383Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        },
        {
          "id": 2224756,
          "postDate": "2023-04-17T16:22:28.987Z",
          "content": "<p>May I ask you how you actually apply interpolation? I'm trying to apply it, but even with soft parameters the score is lower</p>",
          "rawMarkdown": "May I ask you how you actually apply interpolation? I'm trying to apply it, but even with soft parameters the score is lower"
        }
      ]
    },
    {
      "id": 2228313,
      "postDate": "2023-04-20T12:56:37.840Z",
      "content": "<p>watch this thread!</p>\n<p>receipe for LB0.76 one fold:</p>\n<hr>\n<p>i seems that there is somthing very strange about the dataset … maybe very noisy.  i have to use unconventional strong regularisation and training method ….</p>\n<p>hint: <br>\n[1] Dropout Reduces Underfitting<br>\n<a href=\"https://arxiv.org/abs/2303.01500\" target=\"_blank\">https://arxiv.org/abs/2303.01500</a></p>\n<p>i find that if i reduce learning rate, the model easily overfits. instead, according to paper[1], by setting dropout to different values at different training epoches, we can reduce validation loss (it has the same effects as reducing learning rate) without overfitting:</p>\n<ul>\n<li>early: train without dropout</li>\n<li>late: train with large dropout</li>\n<li>final: tune with small dropout</li>\n</ul>\n<p><img src=\"https://i.ibb.co/mz15RwQ/Selection-999-1847.png\" alt=\"https://i.ibb.co/mz15RwQ/Selection-999-1847.png\"></p>\n<p><img src=\"https://i.ibb.co/GvkZB5j/Selection-999-1848.png\" alt=\"https://i.ibb.co/GvkZB5j/Selection-999-1848.png\"></p>\n<pre><code>fold_type = kaggle-random (20% validation, 80% train. pure random)\nfold = 0\nvalid_dataset : \n    len = 18896\n    num_participant_id = 21\n    num_external_video = 0\n\n\ntime_taken =  0 min 27 sec\nloss = 1.7932500839233398\ntopk[0] = 0.8607112616426758 (LB 0.76)\ntopk[1] = 0.9180249788314987\ntopk[2] = 0.9374470787468248\ntopk[3] = 0.9457557154953429\ntopk[4] = 0.9513653683319221\ncheckpoint_file = /home/titanx/hengck/share1/kaggle/2022/hand-sign/result/run202/tx90-face-norm-480-0b/fold-0-kaggle-random/checkpoint/swa.model.pth\n</code></pre>",
      "rawMarkdown": "watch this thread!\n\nreceipe for LB0.76 one fold:\n\n---\n\ni seems that there is somthing very strange about the dataset ... maybe very noisy.  i have to use unconventional strong regularisation and training method ....\n \n\nhint: \n[1] Dropout Reduces Underfitting\nhttps://arxiv.org/abs/2303.01500\n\ni find that if i reduce learning rate, the model easily overfits. instead, according to paper[1], by setting dropout to different values at different training epoches, we can reduce validation loss (it has the same effects as reducing learning rate) without overfitting:\n- early: train without dropout\n- late: train with large dropout\n- final: tune with small dropout\n\n![https://i.ibb.co/mz15RwQ/Selection-999-1847.png](https://i.ibb.co/mz15RwQ/Selection-999-1847.png)\n\n![https://i.ibb.co/GvkZB5j/Selection-999-1848.png](https://i.ibb.co/GvkZB5j/Selection-999-1848.png)\n\n```\nfold_type = kaggle-random (20% validation, 80% train. pure random)\nfold = 0\nvalid_dataset : \n\tlen = 18896\n\tnum_participant_id = 21\n\tnum_external_video = 0\n\n\ntime_taken =  0 min 27 sec\nloss = 1.7932500839233398\ntopk[0] = 0.8607112616426758 (LB 0.76)\ntopk[1] = 0.9180249788314987\ntopk[2] = 0.9374470787468248\ntopk[3] = 0.9457557154953429\ntopk[4] = 0.9513653683319221\ncheckpoint_file = /home/titanx/hengck/share1/kaggle/2022/hand-sign/result/run202/tx90-face-norm-480-0b/fold-0-kaggle-random/checkpoint/swa.model.pth\n\n```",
      "votes": 7,
      "replies": [
        {
          "id": 2228775,
          "postDate": "2023-04-20T20:02:19.243Z",
          "content": "<p>Have you tried using ViT or are you still using Conv1d + MHA?</p>",
          "rawMarkdown": "Have you tried using ViT or are you still using Conv1d + MHA?",
          "replies": [
            {
              "id": 2229582,
              "postDate": "2023-04-21T13:40:39.923Z",
              "content": "<p>dropout has the effect of flatten the valley of loss landscape.<br>\n(this is the same as SWA, deeply supervised loss,  etc)</p>\n<p>Of course, adeverisal training SAM (sharpness aware optimizer, aka sample perturbation via weight perturbation) also flattens the valley.<br>\nBelow is results of SAM, which is comparable to late dropout</p>\n<pre><code>fold_type = kaggle-part\nfold = 0\nvalid_dataset : \n    len = 22959\n    num_participant_id = 5\n    num_external_video = 0\n\n\n    22959 / 22959   0 min 30 sec\n\ntime_taken =  0 min 30 sec\nloss = 2.5410375595092773\ntopk[0] = 0.7265560346704996\ntopk[1] = 0.8181105448843591\ntopk[2] = 0.856265516790801\ntopk[3] = 0.8760398972080665\ntopk[4] = 0.8882791062328499\ncheckpoint_file = /home/user/share1/kaggle/2022/hand-sign/result/run202/tx90-face-norm-sam-fine-rho0.1/fold-0-kaggle-part/checkpoint/swa.model.pth\n----- end -----\n</code></pre>",
              "rawMarkdown": "dropout has the effect of flatten the valley of loss landscape.\n(this is the same as SWA, deeply supervised loss,  etc)\n\nOf course, adeverisal training SAM (sharpness aware optimizer, aka sample perturbation via weight perturbation) also flattens the valley.\nBelow is results of SAM, which is comparable to late dropout\n\n```\nfold_type = kaggle-part\nfold = 0\nvalid_dataset : \n\tlen = 22959\n\tnum_participant_id = 5\n\tnum_external_video = 0\n\n\n    22959 / 22959   0 min 30 sec\n\ntime_taken =  0 min 30 sec\nloss = 2.5410375595092773\ntopk[0] = 0.7265560346704996\ntopk[1] = 0.8181105448843591\ntopk[2] = 0.856265516790801\ntopk[3] = 0.8760398972080665\ntopk[4] = 0.8882791062328499\ncheckpoint_file = /home/user/share1/kaggle/2022/hand-sign/result/run202/tx90-face-norm-sam-fine-rho0.1/fold-0-kaggle-part/checkpoint/swa.model.pth\n----- end -----\n \n```\n ",
              "votes": 1
            },
            {
              "id": 2229794,
              "postDate": "2023-04-21T17:30:00.240Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2229796,
              "postDate": "2023-04-21T17:30:20.373Z",
              "content": "<p>It may be worthwhile to try additional parameters in conjunction with a 'late' configuration</p>",
              "rawMarkdown": "It may be worthwhile to try additional parameters in conjunction with a 'late' configuration"
            }
          ]
        }
      ]
    },
    {
      "id": 2163630,
      "postDate": "2023-03-01T00:49:08.950Z",
      "content": "<p>[reference]</p>\n<ul>\n<li>'Sign Pose-based Transformer for Word-level Sign Language Recognition' - Matyáš Boháček</li>\n<li>Combining Efficient and Precise Sign Language Recognition: Good pose estimation library is all you need- Matyáš Boháček</li>\n</ul>\n<p><a href=\"https://github.com/matyasbohacek/spoter\" target=\"_blank\">https://github.com/matyasbohacek/spoter</a><br>\n<a href=\"https://www.matyasbohacek.com/\" target=\"_blank\">https://www.matyasbohacek.com/</a><br>\n(uses MediaPipe library pose too)</p>\n<p><a href=\"https://huggingface.co/spaces/matyasbohacek/spoter-demo-test\" target=\"_blank\">https://huggingface.co/spaces/matyasbohacek/spoter-demo-test</a></p>\n<p><a href=\"https://www.youtube.com/watch?v=6i5ODpHMny8\" target=\"_blank\">https://www.youtube.com/watch?v=6i5ODpHMny8</a></p>\n<p><img src=\"https://i.ibb.co/r2NHnkh/Selection-999-1166.png\" alt=\"https://i.ibb.co/r2NHnkh/Selection-999-1166.png\"></p>\n<hr>\n<p>SLGTFORMER: AN ATTENTION-BASED APPROACH TO SIGN LANGUAGE RECOGNITION<br>\n<a href=\"https://github.com/neilsong/SLGTformer\" target=\"_blank\">https://github.com/neilsong/SLGTformer</a></p>",
      "rawMarkdown": "[reference]\n- 'Sign Pose-based Transformer for Word-level Sign Language Recognition' - Matyáš Boháček\n- Combining Efficient and Precise Sign Language Recognition: Good pose estimation library is all you need- Matyáš Boháček\n\nhttps://github.com/matyasbohacek/spoter\nhttps://www.matyasbohacek.com/\n(uses MediaPipe library pose too)\n\nhttps://huggingface.co/spaces/matyasbohacek/spoter-demo-test\n\nhttps://www.youtube.com/watch?v=6i5ODpHMny8\n\n![https://i.ibb.co/r2NHnkh/Selection-999-1166.png](https://i.ibb.co/r2NHnkh/Selection-999-1166.png)\n\n\n---\nSLGTFORMER: AN ATTENTION-BASED APPROACH TO SIGN LANGUAGE RECOGNITION\nhttps://github.com/neilsong/SLGTformer\n",
      "votes": 5
    },
    {
      "id": 2163638,
      "postDate": "2023-03-01T00:59:53.167Z",
      "content": "<p>[external data]<br>\n… to be updated …</p>\n<p>HOW TO SIGN<br>\n<a href=\"https://how2sign.github.io/\" target=\"_blank\">https://how2sign.github.io/</a><br>\n<a href=\"https://openaccess.thecvf.com/content/CVPR2021/papers/Duarte_How2Sign_A_Large-Scale_Multimodal_Dataset_for_Continuous_American_Sign_Language_CVPR_2021_paper.pdf\" target=\"_blank\">https://openaccess.thecvf.com/content/CVPR2021/papers/Duarte_How2Sign_A_Large-Scale_Multimodal_Dataset_for_Continuous_American_Sign_Language_CVPR_2021_paper.pdf</a><br>\n(vocab=16k, time=79 hr, signer=11)</p>\n<p>MS-ASL American Sign Language Dataset<br>\n<a href=\"https://microsoft.github.io/data-for-society/dataset?d=MS-ASL-American-Sign-Language-Dataset\" target=\"_blank\">https://microsoft.github.io/data-for-society/dataset?d=MS-ASL-American-Sign-Language-Dataset</a><br>\n<a href=\"https://arxiv.org/pdf/1812.01053.pdf\" target=\"_blank\">https://arxiv.org/pdf/1812.01053.pdf</a><br>\n(vocab=1000, num video=25513, signer=222)</p>\n<p>BSL-1K: Scaling up co-articulated sign language recognition using mouthing cues<br>\n<a href=\"https://www.robots.ox.ac.uk/~vgg/research/bsl1k/\" target=\"_blank\">https://www.robots.ox.ac.uk/~vgg/research/bsl1k/</a><br>\n(vocab=1000, num video=1000)</p>\n<p>ChaLearn LAP Large Scale Signer Independent Isolated Sign Language<br>\nRecognition Challenge: Design, Results and Future Research</p>\n<p>ASL-Skeleton3D and ASL-Phono: Two Novel<br>\nDatasets for the American Sign Language</p>\n<hr>\n<ul>\n<li><p>how to read media pipline to extra keypoint for external data <br>\n… to be udpated …</p></li>\n<li><p>more about PopSign learning app<br>\n<a href=\"https://devpost.com/software/pop-sign-learning\" target=\"_blank\">https://devpost.com/software/pop-sign-learning</a><br>\n<a href=\"https://www.youtube.com/watch?v=nbQAfiJzp2g&amp;t=430s\" target=\"_blank\">https://www.youtube.com/watch?v=nbQAfiJzp2g&amp;t=430s</a><br>\n<a href=\"https://github.com/Accessible-Technology-in-Sign/PopSign\" target=\"_blank\">https://github.com/Accessible-Technology-in-Sign/PopSign</a></p></li>\n</ul>\n<p><a href=\"https://github.com/Accessible-Technology-in-Sign/ASLRT\" target=\"_blank\">https://github.com/Accessible-Technology-in-Sign/ASLRT</a><br>\n\"A pipeline to run tracking Kinect, MediaPipe and AlphaPose experiments on sign language feature data using either HMMs, SBHMMs, or Transformers.\"</p>\n<p>collect data you from youtube?</p>",
      "rawMarkdown": "[external data]\n... to be updated ...\n\nHOW TO SIGN\nhttps://how2sign.github.io/\nhttps://openaccess.thecvf.com/content/CVPR2021/papers/Duarte_How2Sign_A_Large-Scale_Multimodal_Dataset_for_Continuous_American_Sign_Language_CVPR_2021_paper.pdf\n(vocab=16k, time=79 hr, signer=11)\n\nMS-ASL American Sign Language Dataset\nhttps://microsoft.github.io/data-for-society/dataset?d=MS-ASL-American-Sign-Language-Dataset\nhttps://arxiv.org/pdf/1812.01053.pdf\n(vocab=1000, num video=25513, signer=222)\n\nBSL-1K: Scaling up co-articulated sign language recognition using mouthing cues\nhttps://www.robots.ox.ac.uk/~vgg/research/bsl1k/\n(vocab=1000, num video=1000)\n\n\nChaLearn LAP Large Scale Signer Independent Isolated Sign Language\nRecognition Challenge: Design, Results and Future Research\n\nASL-Skeleton3D and ASL-Phono: Two Novel\nDatasets for the American Sign Language\n\n---\n\n- how to read media pipline to extra keypoint for external data \n... to be udpated ...\n\n- more about PopSign learning app\nhttps://devpost.com/software/pop-sign-learning\nhttps://www.youtube.com/watch?v=nbQAfiJzp2g&t=430s\nhttps://github.com/Accessible-Technology-in-Sign/PopSign\n\nhttps://github.com/Accessible-Technology-in-Sign/ASLRT\n\"A pipeline to run tracking Kinect, MediaPipe and AlphaPose experiments on sign language feature data using either HMMs, SBHMMs, or Transformers.\"\n\ncollect data you from youtube?\n\n\n\n\n\n\n\n",
      "votes": 6,
      "replies": [
        {
          "id": 2176272,
          "postDate": "2023-03-10T14:03:14.067Z",
          "content": "<p>pytorch mediapipe … hmm<br>\n<a href=\"https://github.com/zmurez/MediaPipePyTorch\" target=\"_blank\">https://github.com/zmurez/MediaPipePyTorch</a></p>",
          "rawMarkdown": "pytorch mediapipe ... hmm\nhttps://github.com/zmurez/MediaPipePyTorch"
        },
        {
          "id": 2183164,
          "postDate": "2023-03-15T13:53:17.117Z",
          "content": "<p>this image explain why hand are missing.<br>\nyou probably can ignore the legs</p>\n<p><img src=\"https://i.ibb.co/237HvmF/original.png\" alt=\"https://i.ibb.co/237HvmF/original.png\"></p>",
          "rawMarkdown": "this image explain why hand are missing.\nyou probably can ignore the legs\n\n![https://i.ibb.co/237HvmF/original.png](https://i.ibb.co/237HvmF/original.png)",
          "votes": 2
        },
        {
          "id": 2183363,
          "postDate": "2023-03-15T15:52:54.357Z",
          "content": "<p>there is a related product called copycat:<br>\nCopyCat: Using Sign Language Recognition to Help Deaf Children Acquire Language Skills<br>\n<a href=\"https://gvu.gatech.edu/research/projects/copycat-using-sign-language-recognition-help-deaf-children-acquire-language-skills#:~:text=CopyCat%20is%20a%20game%20where,language%20skills%20and%20working%20memory\" target=\"_blank\">https://gvu.gatech.edu/research/projects/copycat-using-sign-language-recognition-help-deaf-children-acquire-language-skills#:~:text=CopyCat%20is%20a%20game%20where,language%20skills%20and%20working%20memory</a>.</p>\n<p><a href=\"https://www.youtube.com/watch?v=WMFZVcey8FU\" target=\"_blank\">https://www.youtube.com/watch?v=WMFZVcey8FU</a><br>\n<a href=\"https://github.com/Accessible-Technology-in-Sign/ASLRT\" target=\"_blank\">https://github.com/Accessible-Technology-in-Sign/ASLRT</a></p>\n<p>much information on HMM + handcraft feature here</p>",
          "rawMarkdown": "there is a related product called copycat:\nCopyCat: Using Sign Language Recognition to Help Deaf Children Acquire Language Skills\nhttps://gvu.gatech.edu/research/projects/copycat-using-sign-language-recognition-help-deaf-children-acquire-language-skills#:~:text=CopyCat%20is%20a%20game%20where,language%20skills%20and%20working%20memory.\n\nhttps://www.youtube.com/watch?v=WMFZVcey8FU\nhttps://github.com/Accessible-Technology-in-Sign/ASLRT\n\nmuch information on HMM + handcraft feature here"
        }
      ]
    },
    {
      "id": 2226810,
      "postDate": "2023-04-19T08:43:02.580Z",
      "content": "<p>i forget an import detail at the data card. suggestions for augmentation:</p>\n<p>\"While the app provided a video example of the sign desired, signers routinely made variants of the sign based on their background and region.\"  </p>\n<ul>\n<li>you need external data or create your variant</li>\n</ul>\n<p>\"More rarely, signers might fingerspell a sign, miss it completely, or produce the wrong sign.\"  </p>\n<ul>\n<li>need to measure label noise and decide if you want to exclude some data</li>\n</ul>\n<p>\" Extraneous movements, such as scratching an itch, or the ending movement from the previous sign or the onset of the next sign, are sometimes included.\"  </p>\n<ul>\n<li>random add hand (from other video at start and end). actually CAM heat map is about to hightlight the useful frame.</li>\n</ul>\n<p>\" Conversely, some signers pressed the button late or released the button early, causing cropping in some sign examples. \"  </p>\n<ul>\n<li>cut and paste (mix up)</li>\n</ul>\n<p>\"Some signers sign with their left hand; others sign with their right. Some signers switch their signing hand. All of these situations must be handled by the game’s recognition system.\"  </p>\n<ul>\n<li>flip</li>\n</ul>\n<hr>\n<p>a also note that some signer make repeated signs in one video</p>",
      "rawMarkdown": "i forget an import detail at the data card. suggestions for augmentation:\n\n\"While the app provided a video example of the sign desired, signers routinely made variants of the sign based on their background and region.\"  \n- you need external data or create your variant\n\n\"More rarely, signers might fingerspell a sign, miss it completely, or produce the wrong sign.\"  \n- need to measure label noise and decide if you want to exclude some data\n\n\" Extraneous movements, such as scratching an itch, or the ending movement from the previous sign or the onset of the next sign, are sometimes included.\"  \n- random add hand (from other video at start and end). actually CAM heat map is about to hightlight the useful frame.\n\n\" Conversely, some signers pressed the button late or released the button early, causing cropping in some sign examples. \"  \n- cut and paste (mix up)\n\n\"Some signers sign with their left hand; others sign with their right. Some signers switch their signing hand. All of these situations must be handled by the game’s recognition system.\"  \n- flip\n\n\n---\n\na also note that some signer make repeated signs in one video\n\n",
      "votes": 3,
      "replies": [
        {
          "id": 2229212,
          "postDate": "2023-04-21T06:49:13.603Z",
          "content": "<p>\"need to measure label noise and decide if you want to exclude some data\"</p>\n<p>So we need to specify some kind of \"ideal\" signs, and compare train data with it?</p>",
          "rawMarkdown": "\"need to measure label noise and decide if you want to exclude some data\"\n\nSo we need to specify some kind of \"ideal\" signs, and compare train data with it?",
          "replies": [
            {
              "id": 2229217,
              "postDate": "2023-04-21T06:51:29.157Z",
              "content": "<p>those  that have validation/train error higher than normal  are supsicous outliers  </p>",
              "rawMarkdown": "those  that have validation/train error higher than normal  are supsicous outliers  ",
              "votes": 1
            },
            {
              "id": 2229256,
              "postDate": "2023-04-21T07:37:43.810Z",
              "content": "<p>If the hidden dataset is very cleanly drawn, I think data with a low frame count can be left out.</p>",
              "rawMarkdown": "If the hidden dataset is very cleanly drawn, I think data with a low frame count can be left out."
            },
            {
              "id": 2229468,
              "postDate": "2023-04-21T11:49:10.123Z",
              "content": "<p>there must be some reason why some video are very long (much longer than average).<br>\nat the signer moving very slowly? or are they not cropping the video (press start and end) correctly? or are they signing wrongly?</p>",
              "rawMarkdown": "there must be some reason why some video are very long (much longer than average).\nat the signer moving very slowly? or are they not cropping the video (press start and end) correctly? or are they signing wrongly?"
            },
            {
              "id": 2229484,
              "postDate": "2023-04-21T12:07:25.227Z",
              "content": "<p>You're right, long frame data has some problems. </p>",
              "rawMarkdown": "You're right, long frame data has some problems. "
            },
            {
              "id": 2229604,
              "postDate": "2023-04-21T14:07:24.983Z",
              "content": "<p>this is explain why conv1d does not work but conv1d+max pool works (about LB 0.72, compared to transformer of 0.74)</p>\n<pre><code>        self.conv2 = nn.Sequential(\n            nn.Conv1d(1024, 768, kernel_size=3, padding=1, stride=1),\n            LayerNorm1D(768),\n            nn.ReLU(inplace=True),\n            nn.Dropout(p=0.1),\n            nn.MaxPool1d(kernel_size=3, padding=1, stride=2),\n        )\n</code></pre>\n<p>max pool is poor man's attention pool(softmax versus hard max)</p>\n<p>also, why adding noise (dropout, dropframe, random values, mix frame …) before mha works. it forces the model to select the correct signal.</p>\n<p>i am very surprise of the small num of paramaters  and performance of transformer compared to other model design (like conv1d, dense, etc)</p>",
              "rawMarkdown": "this is explain why conv1d does not work but conv1d+max pool works (about LB 0.72, compared to transformer of 0.74)\n\n``` \n\t\tself.conv2 = nn.Sequential(\n\t\t\tnn.Conv1d(1024, 768, kernel_size=3, padding=1, stride=1),\n\t\t\tLayerNorm1D(768),\n\t\t\tnn.ReLU(inplace=True),\n\t\t\tnn.Dropout(p=0.1),\n\t\t\tnn.MaxPool1d(kernel_size=3, padding=1, stride=2),\n\t\t)\n\n```\n\nmax pool is poor man's attention pool(softmax versus hard max)\n\nalso, why adding noise (dropout, dropframe, random values, mix frame ...) before mha works. it forces the model to select the correct signal.\n\ni am very surprise of the small num of paramaters  and performance of transformer compared to other model design (like conv1d, dense, etc)"
            }
          ]
        }
      ]
    },
    {
      "id": 2224078,
      "postDate": "2023-04-17T01:26:37.653Z",
      "content": "<p>it is suprised for me to achieve 0.75 in one fold with participant split train data(num participant = 5 in validation data). This makes me think that 0.76 is possible. </p>\n<pre><code>fold_type = kaggle-part\nfold = 0\nvalid_dataset : \n    len = 22959\n    num_participant_id = 5 \n\n    22959 / 22959   0 min 42 sec\n\ntime_taken =  0 min 42 sec\nloss = 2.406057596206665\ntopk[0] = 0.7277320440785748\ntopk[1] = 0.8219870203406072\ntopk[2] = 0.858530423798946\ntopk[3] = 0.8757785617840498\ntopk[4] = 0.8888888888888888\ncheckpoint_file = /home/titanx/hengck/share1/kaggle/2022/hand-sign/result/run201/tx61a-1-frame-first-scale-2/fold-0-kaggle-part/checkpoint/swa.model.pth\n</code></pre>\n<p>and i hit the jackpot 888888888</p>",
      "rawMarkdown": "it is suprised for me to achieve 0.75 in one fold with participant split train data(num participant = 5 in validation data). This makes me think that 0.76 is possible. \n\n```\n\nfold_type = kaggle-part\nfold = 0\nvalid_dataset : \n\tlen = 22959\n\tnum_participant_id = 5 \n \n    22959 / 22959   0 min 42 sec\n\ntime_taken =  0 min 42 sec\nloss = 2.406057596206665\ntopk[0] = 0.7277320440785748\ntopk[1] = 0.8219870203406072\ntopk[2] = 0.858530423798946\ntopk[3] = 0.8757785617840498\ntopk[4] = 0.8888888888888888\ncheckpoint_file = /home/titanx/hengck/share1/kaggle/2022/hand-sign/result/run201/tx61a-1-frame-first-scale-2/fold-0-kaggle-part/checkpoint/swa.model.pth\n \n\n```\n\nand i hit the jackpot 888888888",
      "votes": 3,
      "replies": [
        {
          "id": 2224095,
          "postDate": "2023-04-17T02:08:50.303Z",
          "content": "<p>888888888!</p>",
          "rawMarkdown": "888888888!",
          "votes": 1
        },
        {
          "id": 2225136,
          "postDate": "2023-04-17T23:29:35.293Z",
          "content": "<p>Are these accuracies shown for the 5 holdout participants?</p>",
          "rawMarkdown": "Are these accuracies shown for the 5 holdout participants?"
        }
      ]
    },
    {
      "id": 2206909,
      "postDate": "2023-04-03T02:01:33.027Z",
      "content": "<p>a couple on conv1d (as spatial temporal encoder) + MHA attention seems to have good results.<br>\nsome good ideas to design efficient transformer.</p>\n<p>[1] FastViT: A Fast Hybrid Vision Transformer using Structural Reparameterization<br>\n<a href=\"https://arxiv.org/pdf/2303.14189.pdf\" target=\"_blank\">https://arxiv.org/pdf/2303.14189.pdf</a></p>\n<p>[2] Next-ViT: Next Generation Vision Transformer for Efficient Deployment in Realistic Industrial Scenarios<br>\n<a href=\"https://arxiv.org/abs/2207.05501\" target=\"_blank\">https://arxiv.org/abs/2207.05501</a></p>\n<p>to avoid uncessary permute reshape, i can write my q,v,k linear operation in terms of conv1d.<br>\ni will be removing absolute position embedding and use convolutional position embedding instead (convolutional FFN, i.e. position information is implicitly embedding on 1d conv)</p>\n<p>actually it is aslo possible to use 2d convolution, where (W,H) = (x,y,z of spatial, t of temporal).</p>",
      "rawMarkdown": "a couple on conv1d (as spatial temporal encoder) + MHA attention seems to have good results.\nsome good ideas to design efficient transformer.\n\n[1] FastViT: A Fast Hybrid Vision Transformer using Structural Reparameterization\nhttps://arxiv.org/pdf/2303.14189.pdf\n\n[2] Next-ViT: Next Generation Vision Transformer for Efficient Deployment in Realistic Industrial Scenarios\nhttps://arxiv.org/abs/2207.05501\n\nto avoid uncessary permute reshape, i can write my q,v,k linear operation in terms of conv1d.\ni will be removing absolute position embedding and use convolutional position embedding instead (convolutional FFN, i.e. position information is implicitly embedding on 1d conv)\n\n\nactually it is aslo possible to use 2d convolution, where (W,H) = (x,y,z of spatial, t of temporal).",
      "votes": 3,
      "replies": [
        {
          "id": 2206947,
          "postDate": "2023-04-03T02:55:46.753Z",
          "content": "<p>Won't conv bring much more parameters?</p>\n<p>I suppose what you mean is :</p>\n<pre><code>\ninputs_embeds = inputs_embeds + position_embeds \n\n\ninputs_embeds = inputs_embeds.permute(, , ) \ninputs_embeds = conv1d(inputs_embeds)          \ninputs_embeds = inputs_embeds.permute(, , ) \n</code></pre>\n<p>where the number of parameters in original is TxH, and the conv has HxHx3</p>\n<p>I also tried this idea, but found that the conv is much harder to train than MHSA, maybe because of more parameters.</p>",
          "rawMarkdown": "Won't conv bring much more parameters?\n\nI suppose what you mean is :\n\n```python\n# original Transformer\ninputs_embeds = inputs_embeds + position_embeds # BxTxH <- BxTxH + BxTxH \n\n# conv as position\ninputs_embeds = inputs_embeds.permute(0, 2, 1) # BxHxT <- BxTxH\ninputs_embeds = conv1d(inputs_embeds)          # BxHxT <- BxHxT @ HxHx3 (e.g. conv1d kernel=3)\ninputs_embeds = inputs_embeds.permute(0, 2, 1) # BxTxH <- BxHxT\n```\n\nwhere the number of parameters in original is TxH, and the conv has HxHx3\n\nI also tried this idea, but found that the conv is much harder to train than MHSA, maybe because of more parameters.",
          "votes": 1,
          "replies": [
            {
              "id": 2207339,
              "postDate": "2023-04-03T10:49:08.177Z",
              "content": "<p><a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">@carnozhao</a> </p>\n<p>1d convolution multi-head-attention:</p>\n<pre><code>class MyMultiHeadAttention(nn.Module):\n    def __init__(self,\n            embed_dim,\n            out_dim,\n            qk_dim,\n            v_dim,\n            num_head,\n        ):\n        super().__init__()\n        self.embed_dim = embed_dim\n        self.num_head  = num_head\n        self.qk_dim = qk_dim\n        self.v_dim  = v_dim\n\n        self.q = nn.Conv1d(embed_dim, qk_dim*num_head,1)\n        self.k = nn.Conv1d(embed_dim, qk_dim*num_head,1) # stride=2 for token reduction, kernel&gt;1 for mixing\n        self.v = nn.Conv1d(embed_dim, v_dim*num_head,1)\n\n        self.out = nn.Conv1d(v_dim*num_head, out_dim, 1)\n        self.scale = 1/(qk_dim**0.5)\n\n    #https://github.com/pytorch/pytorch/issues/40497\n    def forward(self, x, x_mask):\n        B,dim,L= x.shape\n\n        num_head = self.num_head\n        qk_dim = self.qk_dim\n        v_dim = self.v_dim\n\n        q = self.q(x) #B,qk_dim,L\n        k = self.k(x)\n        v = self.v(x)\n        q = q.reshape(B, num_head, qk_dim, L).permute(0,1,3,2).contiguous()\n        k = k.reshape(B, num_head, qk_dim, L)#.permute(0,1,2,3).contiguous()\n        v = v.reshape(B, num_head, v_dim,  L).permute(0,1,3,2).contiguous()\n\n        dot = torch.matmul(q, k) *self.scale  # H L L\n        x_mask = x_mask.reshape(B,1,1,L).expand(-1,num_head,L,-1)\n        #dot[x_mask]= -1e4\n        dot.masked_fill_(x_mask, -1e4)\n        attn = F.softmax(dot, -1)    # L L\n\n        v = torch.matmul(attn, v)  # L H dim\n        v = v.permute(0,1,3,2).reshape(B, v_dim*num_head,L).contiguous()\n        out = self.out(v)\n        return out\n\n#---\nembed_dim = 512\nout_dim   = 512\nqk_dim    = 512//4 #for one head\nv_dim     = 512//4\nnum_head  = 4\nmax_length = 96\nbatch_size = 4\n\nmha = MyMultiHeadAttention(\n    embed_dim,\n    out_dim,\n    qk_dim,\n    v_dim,\n    num_head,\n)\nx  = torch.from_numpy( np.random.uniform(-1,1,(batch_size, embed_dim, max_length))).float()\nx_mask  = torch.from_numpy( np.random.uniform(0,1,(batch_size, max_length)))\nx_mask = x_mask&gt;0.5\n\ny=mha(x,x_mask)\nprint(y.shape) #torch.Size([4, 512, 96])\n</code></pre>\n<p>as a comparison:</p>\n<p>original linear multi-head-attention:</p>\n<pre><code>class MyMultiHeadAttention(nn.Module):\n    def __init__(self,\n            embed_dim,\n            out_dim,\n            qk_dim,\n            v_dim,\n            num_head,\n        ):\n        super().__init__()\n        self.embed_dim = embed_dim\n        self.num_head  = num_head\n        self.qk_dim = qk_dim\n        self.v_dim  = v_dim\n\n        self.q = nn.Linear(embed_dim, qk_dim*num_head)\n        self.k = nn.Linear(embed_dim, qk_dim*num_head)\n        self.v = nn.Linear(embed_dim, v_dim*num_head)\n\n        self.out = nn.Linear(v_dim*num_head, out_dim)\n        self.scale = 1/(qk_dim**0.5)\n\n    #https://github.com/pytorch/pytorch/issues/40497\n    def forward(self, x, x_mask):\n        B,L,dim = x.shape\n        #out, _ = self.mha(x,x,x, key_padding_mask=x_mask)\n        num_head = self.num_head\n        qk_dim = self.qk_dim\n        v_dim = self.v_dim\n\n        q = self.q(x)\n        k = self.k(x)\n        v = self.v(x)\n        q = q.reshape(B, L, num_head, qk_dim).permute(0,2,1,3).contiguous()\n        k = k.reshape(B, L, num_head, qk_dim).permute(0,2,3,1).contiguous()\n        v = v.reshape(B, L, num_head, v_dim ).permute(0,2,1,3).contiguous()\n\n        dot = torch.matmul(q, k) *self.scale  # H L L\n        x_mask = x_mask.reshape(B,1,1,L).expand(-1,num_head,L,-1)\n        #dot[x_mask]= -1e4\n        dot.masked_fill_(x_mask, -1e4)\n        attn = F.softmax(dot, -1)    # L L\n\n        v = torch.matmul(attn, v)  # L H dim\n        v = v.permute(0,2,1,3).reshape(B,L, v_dim*num_head).contiguous()\n        out = self.out(v)\n        return out\n</code></pre>",
              "rawMarkdown": "@carnozhao \n\n1d convolution multi-head-attention:\n\n```\nclass MyMultiHeadAttention(nn.Module):\n    def __init__(self,\n            embed_dim,\n            out_dim,\n            qk_dim,\n            v_dim,\n            num_head,\n        ):\n        super().__init__()\n        self.embed_dim = embed_dim\n        self.num_head  = num_head\n        self.qk_dim = qk_dim\n        self.v_dim  = v_dim\n\n        self.q = nn.Conv1d(embed_dim, qk_dim*num_head,1)\n        self.k = nn.Conv1d(embed_dim, qk_dim*num_head,1) # stride=2 for token reduction, kernel>1 for mixing\n        self.v = nn.Conv1d(embed_dim, v_dim*num_head,1)\n\n        self.out = nn.Conv1d(v_dim*num_head, out_dim, 1)\n        self.scale = 1/(qk_dim**0.5)\n\n    #https://github.com/pytorch/pytorch/issues/40497\n    def forward(self, x, x_mask):\n        B,dim,L= x.shape\n\n        num_head = self.num_head\n        qk_dim = self.qk_dim\n        v_dim = self.v_dim\n\n        q = self.q(x) #B,qk_dim,L\n        k = self.k(x)\n        v = self.v(x)\n        q = q.reshape(B, num_head, qk_dim, L).permute(0,1,3,2).contiguous()\n        k = k.reshape(B, num_head, qk_dim, L)#.permute(0,1,2,3).contiguous()\n        v = v.reshape(B, num_head, v_dim,  L).permute(0,1,3,2).contiguous()\n\n        dot = torch.matmul(q, k) *self.scale  # H L L\n        x_mask = x_mask.reshape(B,1,1,L).expand(-1,num_head,L,-1)\n        #dot[x_mask]= -1e4\n        dot.masked_fill_(x_mask, -1e4)\n        attn = F.softmax(dot, -1)    # L L\n\n        v = torch.matmul(attn, v)  # L H dim\n        v = v.permute(0,1,3,2).reshape(B, v_dim*num_head,L).contiguous()\n        out = self.out(v)\n        return out\n\n#---\nembed_dim = 512\nout_dim   = 512\nqk_dim    = 512//4 #for one head\nv_dim     = 512//4\nnum_head  = 4\nmax_length = 96\nbatch_size = 4\n\nmha = MyMultiHeadAttention(\n    embed_dim,\n    out_dim,\n    qk_dim,\n    v_dim,\n    num_head,\n)\nx  = torch.from_numpy( np.random.uniform(-1,1,(batch_size, embed_dim, max_length))).float()\nx_mask  = torch.from_numpy( np.random.uniform(0,1,(batch_size, max_length)))\nx_mask = x_mask>0.5\n\ny=mha(x,x_mask)\nprint(y.shape) #torch.Size([4, 512, 96])\n```\n\nas a comparison:\n\noriginal linear multi-head-attention:\n\n```\nclass MyMultiHeadAttention(nn.Module):\n    def __init__(self,\n            embed_dim,\n            out_dim,\n            qk_dim,\n            v_dim,\n            num_head,\n        ):\n        super().__init__()\n        self.embed_dim = embed_dim\n        self.num_head  = num_head\n        self.qk_dim = qk_dim\n        self.v_dim  = v_dim\n\n        self.q = nn.Linear(embed_dim, qk_dim*num_head)\n        self.k = nn.Linear(embed_dim, qk_dim*num_head)\n        self.v = nn.Linear(embed_dim, v_dim*num_head)\n\n        self.out = nn.Linear(v_dim*num_head, out_dim)\n        self.scale = 1/(qk_dim**0.5)\n\n    #https://github.com/pytorch/pytorch/issues/40497\n    def forward(self, x, x_mask):\n        B,L,dim = x.shape\n        #out, _ = self.mha(x,x,x, key_padding_mask=x_mask)\n        num_head = self.num_head\n        qk_dim = self.qk_dim\n        v_dim = self.v_dim\n\n        q = self.q(x)\n        k = self.k(x)\n        v = self.v(x)\n        q = q.reshape(B, L, num_head, qk_dim).permute(0,2,1,3).contiguous()\n        k = k.reshape(B, L, num_head, qk_dim).permute(0,2,3,1).contiguous()\n        v = v.reshape(B, L, num_head, v_dim ).permute(0,2,1,3).contiguous()\n\n        dot = torch.matmul(q, k) *self.scale  # H L L\n        x_mask = x_mask.reshape(B,1,1,L).expand(-1,num_head,L,-1)\n        #dot[x_mask]= -1e4\n        dot.masked_fill_(x_mask, -1e4)\n        attn = F.softmax(dot, -1)    # L L\n\n        v = torch.matmul(attn, v)  # L H dim\n        v = v.permute(0,2,1,3).reshape(B,L, v_dim*num_head).contiguous()\n        out = self.out(v)\n        return out\n\n```",
              "votes": 1
            }
          ]
        },
        {
          "id": 2208680,
          "postDate": "2023-04-04T08:53:19.857Z",
          "content": "<p>My conv1d layers are not getting through the pytorch to tflite conversion. Do you still do pytorch -&gt; onnx -&gt;tflite conversion or do you have a tf/keras twin which you convert to tflite?</p>",
          "rawMarkdown": "My conv1d layers are not getting through the pytorch to tflite conversion. Do you still do pytorch -> onnx ->tflite conversion or do you have a tf/keras twin which you convert to tflite?",
          "votes": 2
        }
      ]
    },
    {
      "id": 2204321,
      "postDate": "2023-03-31T13:31:22.793Z",
      "content": "<p>this is how much the face (upper triangle = eyes and lip) and body(lower quadliteral = shoulders and hips)  have moved for one hand sign.</p>\n<p>I wonder if there will be improvement in accuracy if i remove these noise motion?</p>\n<p><img src=\"https://i.ibb.co/xgPNq9L/Selection-999-1647.png\" alt=\"https://i.ibb.co/xgPNq9L/Selection-999-1647.png\"></p>",
      "rawMarkdown": "this is how much the face (upper triangle = eyes and lip) and body(lower quadliteral = shoulders and hips)  have moved for one hand sign.\n\nI wonder if there will be improvement in accuracy if i remove these noise motion?\n\n![https://i.ibb.co/xgPNq9L/Selection-999-1647.png](https://i.ibb.co/xgPNq9L/Selection-999-1647.png)",
      "votes": 3,
      "replies": [
        {
          "id": 2204337,
          "postDate": "2023-03-31T13:42:42.513Z",
          "content": "<p>Tried to apply moving average across frames but it doesn't work for me </p>",
          "rawMarkdown": "Tried to apply moving average across frames but it doesn't work for me "
        }
      ]
    },
    {
      "id": 2206111,
      "postDate": "2023-04-02T09:49:09.440Z",
      "content": "<p>simplest normalisation code</p>\n<pre><code>REF = [500, 501, 512, 513, 159,  386, 13,]\n\ndef do_normalise_by_ref(xyz, ref):  \n    K = xyz.shape[-1]\n    xyz_flat = ref.reshape(-1,K)\n    m = np.nanmean(xyz_flat,0).reshape(1,1,K)\n    s = np.nanstd(xyz_flat, 0).mean() \n    xyz = xyz - m\n    xyz = xyz / s\n    return xyz\n\nxyz = load_relevant_data_subset(pq_file)[...,:2]\nxyz = do_normalise_by_ref(xyz,xyz[:,REF])\n</code></pre>",
      "rawMarkdown": "simplest normalisation code\n\n```\n\nREF = [500, 501, 512, 513, 159,  386, 13,]\n\ndef do_normalise_by_ref(xyz, ref):  \n\tK = xyz.shape[-1]\n\txyz_flat = ref.reshape(-1,K)\n\tm = np.nanmean(xyz_flat,0).reshape(1,1,K)\n\ts = np.nanstd(xyz_flat, 0).mean() \n\txyz = xyz - m\n\txyz = xyz / s\n\treturn xyz\n\nxyz = load_relevant_data_subset(pq_file)[...,:2]\nxyz = do_normalise_by_ref(xyz,xyz[:,REF])\n\n```\n",
      "votes": 4,
      "replies": [
        {
          "id": 2207859,
          "postDate": "2023-04-03T16:59:36.753Z",
          "rawMarkdown": "",
          "isDeleted": true,
          "replies": [
            {
              "id": 2207929,
              "postDate": "2023-04-03T17:40:00.700Z",
              "content": "<pre><code>REF = [500, 501, 512, 513, 159,  386, 13,]\nref = xyz[:,REF]\nxyz_flat = ref.reshape(-1,K)\n</code></pre>\n<p>using ref=xyz will be more accurate<br>\nyou can conduct experiments to find the subset of most stable normalising points</p>",
              "rawMarkdown": "```\nREF = [500, 501, 512, 513, 159,  386, 13,]\nref = xyz[:,REF]\nxyz_flat = ref.reshape(-1,K)\n\n```\n\nusing ref=xyz will be more accurate\nyou can conduct experiments to find the subset of most stable normalising points"
            },
            {
              "id": 2208007,
              "postDate": "2023-04-03T18:33:32.503Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 2178029,
      "postDate": "2023-03-12T02:36:03.543Z",
      "content": "<p>[self-study]</p>\n<ul>\n<li>Sign Language Translation with Transformers<br>\n<a href=\"https://www.youtube.com/watch?v=E5nKeEvoAK0\" target=\"_blank\">https://www.youtube.com/watch?v=E5nKeEvoAK0</a></li>\n</ul>",
      "rawMarkdown": "[self-study]\n- Sign Language Translation with Transformers\nhttps://www.youtube.com/watch?v=E5nKeEvoAK0\n\n",
      "votes": 3
    },
    {
      "id": 2163502,
      "postDate": "2023-02-28T21:37:38.583Z",
      "content": "<p>[ensemble trick]</p>\n<ul>\n<li>since only one model can be submitted and there is a time limit, we can distilled \"ensemble knowledge\" into one model.</li>\n<li>a better solution is distilled encoder (or multiple encoders) and multiple heads. heads are using different \"view\" of encoded features, final prediction is averaged over the heads</li>\n</ul>\n<p>more on this later</p>",
      "rawMarkdown": "[ensemble trick]\n- since only one model can be submitted and there is a time limit, we can distilled \"ensemble knowledge\" into one model.\n- a better solution is distilled encoder (or multiple encoders) and multiple heads. heads are using different \"view\" of encoded features, final prediction is averaged over the heads\n  \n\nmore on this later",
      "votes": 3
    },
    {
      "id": 2163494,
      "postDate": "2023-02-28T21:20:38.053Z",
      "content": "<p>[augmentation and tricks]</p>\n<ul>\n<li>simulate different frame length </li>\n<li>simulate different NaN (as in train data)</li>\n<li>simulate different missing frame or differnt frame rate? </li>\n<li>3d guassian noise</li>\n<li>type of landmark. use ['face', 'left_hand', 'pose', 'right_hand'] for normalisation and augmentation<br>\n(e.g. affine transform to group of point --&gt; bigger mouth, longer arm …, rotate hand)</li>\n</ul>\n<p>since landmark is xyz, we are having a 3d motion capture. would be fun to use GAN to create more data. distangle motion and  shape</p>\n<p>possible to synthetic animate motion for new person via motion transfer. But we need to find some media pipe human xyz models. </p>",
      "rawMarkdown": "[augmentation and tricks]\n- simulate different frame length \n- simulate different NaN (as in train data)\n- simulate different missing frame or differnt frame rate? \n- 3d guassian noise\n- type of landmark. use ['face', 'left_hand', 'pose', 'right_hand'] for normalisation and augmentation\n  (e.g. affine transform to group of point --> bigger mouth, longer arm ..., rotate hand)\n\nsince landmark is xyz, we are having a 3d motion capture. would be fun to use GAN to create more data. distangle motion and  shape\n\npossible to synthetic animate motion for new person via motion transfer. But we need to find some media pipe human xyz models. ",
      "votes": 3,
      "replies": [
        {
          "id": 2178001,
          "postDate": "2023-03-12T02:16:37.907Z",
          "content": "<p>augmentation paper:<br>\n[1] PoseAug: A Differentiable Pose Augmentation Framework for 3D Human Pose Estimation<br>\n<a href=\"https://arxiv.org/pdf/2105.02465.pdf\" target=\"_blank\">https://arxiv.org/pdf/2105.02465.pdf</a></p>\n<p>[2] DH-AUG: DH Forward Kinematics Model Driven Augmentation for 3D Human Pose Estimation<br>\n<a href=\"https://arxiv.org/pdf/2207.09303.pdf\" target=\"_blank\">https://arxiv.org/pdf/2207.09303.pdf</a></p>",
          "rawMarkdown": "augmentation paper:\n[1] PoseAug: A Differentiable Pose Augmentation Framework for 3D Human Pose Estimation\nhttps://arxiv.org/pdf/2105.02465.pdf\n\n\n[2] DH-AUG: DH Forward Kinematics Model Driven Augmentation for 3D Human Pose Estimation\nhttps://arxiv.org/pdf/2207.09303.pdf\n",
          "votes": 2,
          "replies": [
            {
              "id": 2178020,
              "postDate": "2023-03-12T02:29:06.657Z",
              "content": "<p>i check github for ASL generator<br>\n\"3D Avatar ASL github \"</p>",
              "rawMarkdown": "i check github for ASL generator\n\"3D Avatar ASL github \"\n"
            },
            {
              "id": 2178774,
              "postDate": "2023-03-12T17:46:15.033Z",
              "content": "<p>i wonder if motion transfer would work?<br>\ni.e. transfer from one signer to another</p>\n<p><a href=\"https://deepmotionediting.github.io/style_transfer\" target=\"_blank\">https://deepmotionediting.github.io/style_transfer</a></p>",
              "rawMarkdown": "i wonder if motion transfer would work?\ni.e. transfer from one signer to another\n\nhttps://deepmotionediting.github.io/style_transfer",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2170203,
      "postDate": "2023-03-05T18:52:03.680Z",
      "content": "<p>[normalization]</p>\n<p>normalization is all you need!</p>\n<pre><code>def pre_process(xyz):\n    lip = xyz[:, LIP]\n    lhand = xyz[:, LHAND]\n    rhand = xyz[:, RHAND]\n    xyz = torch.cat([ #(none, 82, 3)\n        lip,\n        lhand,\n        rhand,\n    ],1)\n    xyz = xyz - xyz[~torch.isnan(xyz)].mean(0,keepdims=True) #noramlisation to common maen\n    xyz = xyz / xyz[~torch.isnan(xyz)].std(0, keepdims=True)\n    xyz[torch.isnan(xyz)] = 0\n    xyz = xyz[:max_length]\n    return xyz\n</code></pre>",
      "rawMarkdown": "[normalization]\n\nnormalization is all you need!\n```\n\n\ndef pre_process(xyz):\n\tlip = xyz[:, LIP]\n\tlhand = xyz[:, LHAND]\n\trhand = xyz[:, RHAND]\n\txyz = torch.cat([ #(none, 82, 3)\n\t\tlip,\n\t\tlhand,\n\t\trhand,\n\t],1)\n\txyz = xyz - xyz[~torch.isnan(xyz)].mean(0,keepdims=True) #noramlisation to common maen\n\txyz = xyz / xyz[~torch.isnan(xyz)].std(0, keepdims=True)\n\txyz[torch.isnan(xyz)] = 0\n\txyz = xyz[:max_length]\n\treturn xyz\n``` ",
      "votes": 4,
      "replies": [
        {
          "id": 2170328,
          "postDate": "2023-03-05T21:44:12.203Z",
          "content": "<pre><code>** start training here! **\n   batch_size = 64 \n   experiment = ['transformer-pool-03-hand_lip', 'run_train_fold0.py']\n                           |---------------- VALID---------|---- TRAIN/BATCH ----------------------\nrate      iter       epoch | loss   top1   top2    top5    | loss                 | time           \n---------------------------------------------------------------------------------------------------\n0.00e+0   00000000*   0.00 | 5.544  0.005  0.0078  0.019   | 0.000  0.000  0.000  |  0 hr 00 min\n1.00e-4   00003600*   3.00 | 3.183  0.293  0.4145  0.572   | 2.815  0.000  0.000  |  0 hr 04 min\n1.00e-4   00007200*   6.00 | 2.637  0.406  0.5310  0.678   | 2.038  0.000  0.000  |  0 hr 08 min\n1.00e-4   00010800*   9.00 | 2.404  0.451  0.5784  0.719   | 1.671  0.000  0.000  |  0 hr 11 min\n1.00e-4   00014400*  12.00 | 2.363  0.475  0.6005  0.729   | 1.540  0.000  0.000  |  0 hr 15 min\n1.00e-4   00018000*  15.00 | 2.198  0.505  0.6321  0.758   | 1.296  0.000  0.000  |  0 hr 19 min\n1.00e-4   00021600*  18.00 | 2.146  0.519  0.6440  0.765   | 1.225  0.000  0.000  |  0 hr 23 min\n1.00e-4   00025200*  21.00 | 2.141  0.524  0.6510  0.765   | 1.090  0.000  0.000  |  0 hr 27 min\n1.00e-4   00028800*  24.00 | 2.101  0.529  0.6570  0.775   | 0.985  0.000  0.000  |  0 hr 31 min\n1.00e-4   00032400*  27.00 | 2.071  0.540  0.6621  0.777   | 0.934  0.000  0.000  |  0 hr 35 min\n1.00e-4   00036000*  30.00 | 2.131  0.538  0.6573  0.772   | 0.911  0.000  0.000  |  0 hr 39 min\n1.00e-4   00039600*  33.00 | 2.070  0.544  0.6674  0.778   | 0.823  0.000  0.000  |  0 hr 43 min\n1.00e-4   00043200*  36.00 | 2.104  0.543  0.6666  0.779   | 0.736  0.000  0.000  |  0 hr 47 min\n1.00e-4   00046800*  39.00 | 2.114  0.544  0.6701  0.779   | 0.738  0.000  0.000  |  0 hr 50 min\n</code></pre>\n<p>fast convergence after normalisation for one layer transformer encoder with cls pooling<br>\nvalidation split is by participant id, i.e  participant id does not overlap in test and train</p>",
          "rawMarkdown": "```\n** start training here! **\n   batch_size = 64 \n   experiment = ['transformer-pool-03-hand_lip', 'run_train_fold0.py']\n                           |---------------- VALID---------|---- TRAIN/BATCH ----------------------\nrate      iter       epoch | loss   top1   top2    top5    | loss                 | time           \n---------------------------------------------------------------------------------------------------\n0.00e+0   00000000*   0.00 | 5.544  0.005  0.0078  0.019   | 0.000  0.000  0.000  |  0 hr 00 min\n1.00e-4   00003600*   3.00 | 3.183  0.293  0.4145  0.572   | 2.815  0.000  0.000  |  0 hr 04 min\n1.00e-4   00007200*   6.00 | 2.637  0.406  0.5310  0.678   | 2.038  0.000  0.000  |  0 hr 08 min\n1.00e-4   00010800*   9.00 | 2.404  0.451  0.5784  0.719   | 1.671  0.000  0.000  |  0 hr 11 min\n1.00e-4   00014400*  12.00 | 2.363  0.475  0.6005  0.729   | 1.540  0.000  0.000  |  0 hr 15 min\n1.00e-4   00018000*  15.00 | 2.198  0.505  0.6321  0.758   | 1.296  0.000  0.000  |  0 hr 19 min\n1.00e-4   00021600*  18.00 | 2.146  0.519  0.6440  0.765   | 1.225  0.000  0.000  |  0 hr 23 min\n1.00e-4   00025200*  21.00 | 2.141  0.524  0.6510  0.765   | 1.090  0.000  0.000  |  0 hr 27 min\n1.00e-4   00028800*  24.00 | 2.101  0.529  0.6570  0.775   | 0.985  0.000  0.000  |  0 hr 31 min\n1.00e-4   00032400*  27.00 | 2.071  0.540  0.6621  0.777   | 0.934  0.000  0.000  |  0 hr 35 min\n1.00e-4   00036000*  30.00 | 2.131  0.538  0.6573  0.772   | 0.911  0.000  0.000  |  0 hr 39 min\n1.00e-4   00039600*  33.00 | 2.070  0.544  0.6674  0.778   | 0.823  0.000  0.000  |  0 hr 43 min\n1.00e-4   00043200*  36.00 | 2.104  0.543  0.6666  0.779   | 0.736  0.000  0.000  |  0 hr 47 min\n1.00e-4   00046800*  39.00 | 2.114  0.544  0.6701  0.779   | 0.738  0.000  0.000  |  0 hr 50 min\n\n```\n\nfast convergence after normalisation for one layer transformer encoder with cls pooling\nvalidation split is by participant id, i.e  participant id does not overlap in test and train",
          "votes": 5,
          "replies": [
            {
              "id": 2172332,
              "postDate": "2023-03-07T13:08:18.303Z",
              "content": "<p>learning based shape normalisation:</p>\n<ul>\n<li><p>Spatial Transformer Networks<br>\n<a href=\"https://arxiv.org/pdf/1506.02025.pdf\" target=\"_blank\">https://arxiv.org/pdf/1506.02025.pdf</a></p></li>\n<li><p>PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation<br>\n<a href=\"https://openaccess.thecvf.com/content_cvpr_2017/papers/Qi_PointNet_Deep_Learning_CVPR_2017_paper.pdf\" target=\"_blank\">https://openaccess.thecvf.com/content_cvpr_2017/papers/Qi_PointNet_Deep_Learning_CVPR_2017_paper.pdf</a></p></li>\n</ul>\n<p>\"We predict an affine transformation matrix by a mini-network (T-net in Fig 2) and directly apply this transformation to the coordinates of input points\", This is basically a 3d spatial transformer network (not the modern seq-to-seq transformer)</p>",
              "rawMarkdown": "learning based shape normalisation:\n\n- Spatial Transformer Networks\nhttps://arxiv.org/pdf/1506.02025.pdf\n\n- PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation\nhttps://openaccess.thecvf.com/content_cvpr_2017/papers/Qi_PointNet_Deep_Learning_CVPR_2017_paper.pdf\n\n\"We predict an affine transformation matrix by a mini-network (T-net in Fig 2) and directly apply this transformation to the coordinates of input points\", This is basically a 3d spatial transformer network (not the modern seq-to-seq transformer)\n\n\n\n",
              "votes": 1
            },
            {
              "id": 2174824,
              "postDate": "2023-03-09T12:24:09.203Z",
              "content": "<p>effects of normalization<br>\n<img src=\"https://i.ibb.co/bLYJG27/Selection-999-1295.png\" alt=\"https://i.ibb.co/bLYJG27/Selection-999-1295.png\"></p>",
              "rawMarkdown": "effects of normalization\n![https://i.ibb.co/bLYJG27/Selection-999-1295.png](https://i.ibb.co/bLYJG27/Selection-999-1295.png)\n\n",
              "votes": 3
            },
            {
              "id": 2176110,
              "postDate": "2023-03-10T12:08:13.087Z",
              "content": "<p>I'm fascinated by this. </p>\n<ul>\n<li>Did you use the methods you linked to do the normalization? </li>\n<li>It looks like there is an upward swinging arm (right side of left image) in the original but this looks to be shifted into the mass in the right image. Do you think that's the correct transformation?</li>\n<li>How did you visualize this? (X,Y) coordinates scatter plotted clearly, but I like the soft faded way it looks. Any tips?</li>\n</ul>\n<p>As an aside, what are your thoughts about leveraging angles between keypoint-keypoint vectors within a feature set (primarily hands and pose maybe -- face wouldn't work) instead of the points themselves? I believe, although I may be wrong, that this might make the feature set more agnostic to drift, scale, etc.</p>\n<p>I plan to try a simple model shortly where I replace the hand features with the representative keypoint-keypoint vector angles.</p>\n<hr>\n<p>Anyway, definitely following, thanks for sharing !</p>",
              "rawMarkdown": "I'm fascinated by this. \n* Did you use the methods you linked to do the normalization? \n* It looks like there is an upward swinging arm (right side of left image) in the original but this looks to be shifted into the mass in the right image. Do you think that's the correct transformation?\n* How did you visualize this? (X,Y) coordinates scatter plotted clearly, but I like the soft faded way it looks. Any tips?\n\nAs an aside, what are your thoughts about leveraging angles between keypoint-keypoint vectors within a feature set (primarily hands and pose maybe -- face wouldn't work) instead of the points themselves? I believe, although I may be wrong, that this might make the feature set more agnostic to drift, scale, etc.\n\nI plan to try a simple model shortly where I replace the hand features with the representative keypoint-keypoint vector angles.\n\n---\n\nAnyway, definitely following, thanks for sharing !"
            },
            {
              "id": 2176238,
              "postDate": "2023-03-10T13:27:31.560Z",
              "content": "<p>it is just simple mean and std normalization.<br>\nthere is probably not enough resource (100 msec) to do complicated normalisation</p>\n<pre><code>    fig, ax = plt.subplots()\n    ax.invert_yaxis()\n\n    kaggle_df = pd.read_csv(f'{root_dir}/data/asl-signs/train.ver01.csv')\n    for t,d in kaggle_df.iterrows():\n        print('\\r', f'{t}/{len(kaggle_df)}', end='', flush=True)\n        pq_file = f'{root_dir}/data/asl-signs/{d.path}'\n        xyz = load_relevant_data_subset(pq_file)\n\n\n        pose = xyz[:,POSE]\n        pose = pose.reshape(-1,3)\n\n        non_nan = np.all(~np.isnan(pose), axis=-1)\n        pose = pose[non_nan]\n\n        pose = pose-pose.mean(0,keepdims=True) \n        pose = pose/pose.std(0,keepdims=True)\n\n        ax.scatter(pose[:,0], pose[:,1], alpha=0.01)\n\n        #plt.show()\n        plt.waitforbuttonpress()\n</code></pre>",
              "rawMarkdown": "it is just simple mean and std normalization.\nthere is probably not enough resource (100 msec) to do complicated normalisation\n\n```\n\tfig, ax = plt.subplots()\n\tax.invert_yaxis()\n\n\tkaggle_df = pd.read_csv(f'{root_dir}/data/asl-signs/train.ver01.csv')\n\tfor t,d in kaggle_df.iterrows():\n\t\tprint('\\r', f'{t}/{len(kaggle_df)}', end='', flush=True)\n\t\tpq_file = f'{root_dir}/data/asl-signs/{d.path}'\n\t\txyz = load_relevant_data_subset(pq_file)\n\n\n\t\tpose = xyz[:,POSE]\n\t\tpose = pose.reshape(-1,3)\n\n\t\tnon_nan = np.all(~np.isnan(pose), axis=-1)\n\t\tpose = pose[non_nan]\n\t\t\n\t\tpose = pose-pose.mean(0,keepdims=True) \n\t\tpose = pose/pose.std(0,keepdims=True)\n\n\t\tax.scatter(pose[:,0], pose[:,1], alpha=0.01)\n\n\t\t#plt.show()\n\t\tplt.waitforbuttonpress()\n\n```",
              "votes": 2
            },
            {
              "id": 2176319,
              "postDate": "2023-03-10T14:56:49.513Z",
              "content": "<p>Awesome. Makes sense!</p>",
              "rawMarkdown": "Awesome. Makes sense!"
            },
            {
              "id": 2177309,
              "postDate": "2023-03-11T11:10:47.910Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2179776,
              "postDate": "2023-03-13T11:55:35.670Z",
              "content": "<p>Hi, I'm having trouble reproducing your results. I'm using one-layer transformer with participants 26734, 28656, 16069 as validation, and my top1 val caps at 0.35 using the following hyperparams:</p>\n<ul>\n<li>batch_size 64</li>\n<li>optimizer RAdam+Lookahead (k=5)</li>\n<li>landmarks lips, left_hand, right_hand</li>\n<li>constant learning rate 1e-4 </li>\n<li>embed_dim 1024</li>\n<li>num_heads 8</li>\n<li>cls dropout 0.4</li>\n<li>max_len 512</li>\n<li>cls pooling</li>\n<li>40 epoch<br>\nDo you have some idea why this doesn't work?</li>\n</ul>",
              "rawMarkdown": "Hi, I'm having trouble reproducing your results. I'm using one-layer transformer with participants 26734, 28656, 16069 as validation, and my top1 val caps at 0.35 using the following hyperparams:\n\n- batch_size 64\n- optimizer RAdam+Lookahead (k=5)\n- landmarks lips, left_hand, right_hand\n- constant learning rate 1e-4 \n- embed_dim 1024\n- num_heads 8\n- cls dropout 0.4\n- max_len 512\n- cls pooling\n- 40 epoch\nDo you have some idea why this doesn't work?"
            },
            {
              "id": 2179965,
              "postDate": "2023-03-13T13:56:05.057Z",
              "content": "<p><a href=\"https://www.kaggle.com/vbogach\" target=\"_blank\">@vbogach</a> </p>\n<p>most likely there is some bug in your training code</p>\n<p>the training loss is shown at:<br>\n<a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/391265#2170328\" target=\"_blank\">https://www.kaggle.com/competitions/asl-signs/discussion/391265#2170328</a></p>\n<p>possible bug ares:</p>\n<ul>\n<li>forget to normalize shape</li>\n<li>NaN values are not handled</li>\n<li>forget to set model .eval() and .train()</li>\n<li>use probability to ce loss function (instead of logit)</li>\n<li>wrong loss function</li>\n<li>wrong augmentation </li>\n</ul>",
              "rawMarkdown": "@vbogach \n\nmost likely there is some bug in your training code\n\nthe training loss is shown at:\nhttps://www.kaggle.com/competitions/asl-signs/discussion/391265#2170328\n\npossible bug ares:\n- forget to normalize shape\n- NaN values are not handled\n- forget to set model .eval() and .train()\n- use probability to ce loss function (instead of logit)\n- wrong loss function\n- wrong augmentation \n",
              "votes": 1
            },
            {
              "id": 2180400,
              "postDate": "2023-03-13T19:20:24.993Z",
              "content": "<p>It looks like I was normalizing each landmark separately and that's why it didn't work. Thanks for your help!</p>",
              "rawMarkdown": "It looks like I was normalizing each landmark separately and that's why it didn't work. Thanks for your help!"
            }
          ]
        }
      ]
    },
    {
      "id": 2240056,
      "postDate": "2023-04-30T07:22:23.977Z",
      "content": "<p>Excellent sharing！I learned a lot from this.Thank you very much!Wish you can get the top prize!</p>",
      "rawMarkdown": "Excellent sharing！I learned a lot from this.Thank you very much!Wish you can get the top prize!",
      "votes": 1
    },
    {
      "id": 2224702,
      "postDate": "2023-04-17T15:03:30.517Z",
      "content": "<p><a href=\"https://github.com/facebookresearch/dropout\" target=\"_blank\">https://github.com/facebookresearch/dropout</a><br>\nDropout Reduces Underfitting</p>\n<p><img src=\"https://user-images.githubusercontent.com/8370623/222586143-3500fa5b-c294-48c9-a5cf-5fac2659e519.png\" alt=\"https://user-images.githubusercontent.com/8370623/222586143-3500fa5b-c294-48c9-a5cf-5fac2659e519.png\"></p>\n<p>late dropout works for me</p>",
      "rawMarkdown": "https://github.com/facebookresearch/dropout\nDropout Reduces Underfitting\n\n![https://user-images.githubusercontent.com/8370623/222586143-3500fa5b-c294-48c9-a5cf-5fac2659e519.png](https://user-images.githubusercontent.com/8370623/222586143-3500fa5b-c294-48c9-a5cf-5fac2659e519.png)\n\nlate dropout works for me",
      "votes": 1
    },
    {
      "id": 2224196,
      "postDate": "2023-04-17T04:45:42.443Z",
      "content": "<p>replace the word \"academics\" with \"kaggle\"</p>\n<p><img src=\"https://i.ibb.co/rxL6Thm/Selection-999-1831.png\" alt=\"https://i.ibb.co/rxL6Thm/Selection-999-1831.png\"></p>",
      "rawMarkdown": "replace the word \"academics\" with \"kaggle\"\n\n![https://i.ibb.co/rxL6Thm/Selection-999-1831.png](https://i.ibb.co/rxL6Thm/Selection-999-1831.png)",
      "votes": 1
    },
    {
      "id": 2210622,
      "postDate": "2023-04-05T14:08:28.510Z",
      "content": "<p>i wonder if (inverse) DTW can be used as augmentation?<br>\n<a href=\"https://ai.googleblog.com/2021/01/recognizing-pose-similarity-in-images.html\" target=\"_blank\">https://ai.googleblog.com/2021/01/recognizing-pose-similarity-in-images.html</a></p>",
      "rawMarkdown": "i wonder if (inverse) DTW can be used as augmentation?\nhttps://ai.googleblog.com/2021/01/recognizing-pose-similarity-in-images.html",
      "votes": 1
    },
    {
      "id": 2204573,
      "postDate": "2023-03-31T17:21:01.103Z",
      "content": "<p>Hi, I've noticed that you use <code>xyz = xyz - xyz[~torch.isnan(xyz)].mean(0,keepdim=True)</code> to normalize the input. However, <code>xyz[~torch.isnan(xyz)]</code> (or <code>np.isnan</code>) returns a <em>flattened</em> tensor of size <code>(num_frames * num_landmarks * 3 - num_nans)</code>, which means that calling <code>mean</code> will return a scalar. In other words, what you're really doing is subtracting the mean of everything in the sequence instead of normalizing each landmark <strong>across frames</strong> (passing <code>dim=0</code> and <code>keepdim=True</code>). Am I getting it correctly?</p>\n<p>Nevertheless, I believe the correct mean normalization here is to subtract the mean of all landmark positions <strong>for each frame</strong> by doing something like <code>xyz -= xyz.nanmean(dim=1, keepdim=True)</code>. It might be even better to normalize x, y, and z separately but I haven't figured out how to do it yet.</p>",
      "rawMarkdown": "Hi, I've noticed that you use `xyz = xyz - xyz[~torch.isnan(xyz)].mean(0,keepdim=True)` to normalize the input. However, `xyz[~torch.isnan(xyz)]` (or `np.isnan`) returns a *flattened* tensor of size `(num_frames * num_landmarks * 3 - num_nans)`, which means that calling `mean` will return a scalar. In other words, what you're really doing is subtracting the mean of everything in the sequence instead of normalizing each landmark **across frames** (passing `dim=0` and `keepdim=True`). Am I getting it correctly?\n\nNevertheless, I believe the correct mean normalization here is to subtract the mean of all landmark positions **for each frame** by doing something like `xyz -= xyz.nanmean(dim=1, keepdim=True)`. It might be even better to normalize x, y, and z separately but I haven't figured out how to do it yet.",
      "votes": 1,
      "replies": [
        {
          "id": 2204581,
          "postDate": "2023-03-31T17:26:31.320Z",
          "content": "<p>do a few print statement:</p>\n<pre><code>print(xyz.shape)\nprint((torch.isnan(xyz)).shape)\nprint((xyz[~torch.isnan(xyz)]).shape)\nprint((xyz[~torch.isnan(xyz)].mean(0,keepdim=True)).shape)\n</code></pre>",
          "rawMarkdown": "do a few print statement:\n\n```\nprint(xyz.shape)\nprint((torch.isnan(xyz)).shape)\nprint((xyz[~torch.isnan(xyz)]).shape)\nprint((xyz[~torch.isnan(xyz)].mean(0,keepdim=True)).shape)\n\n```",
          "replies": [
            {
              "id": 2204811,
              "postDate": "2023-04-01T01:44:46.860Z",
              "content": "<p>Here are the sizes of everything, which are literally the same as what I said above. Was it your original intention to normalize by the single mean of all landmarks in the sequence? Just curious about this as <code>dim=0, keepdim=True</code> means the opposite.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6782247%2F733adc2538a07c8859e8a7b41bcb027a%2FScreenshot%20from%202023-04-01%2008-36-47.png?generation=1680313038017988&amp;alt=media\" alt=\"\"></p>",
              "rawMarkdown": "Here are the sizes of everything, which are literally the same as what I said above. Was it your original intention to normalize by the single mean of all landmarks in the sequence? Just curious about this as `dim=0, keepdim=True` means the opposite.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6782247%2F733adc2538a07c8859e8a7b41bcb027a%2FScreenshot%20from%202023-04-01%2008-36-47.png?generation=1680313038017988&alt=media)"
            },
            {
              "id": 2204855,
              "postDate": "2023-04-01T03:27:25.140Z",
              "content": "<p><br>\n</p>",
              "rawMarkdown": "~~nanmen is correct. seems to be a bug for xyz[~torch.isnan(xyz)]~~\n~~Thanks.~~\n",
              "votes": 1
            },
            {
              "id": 2204858,
              "postDate": "2023-04-01T03:29:00.130Z",
              "content": "<p><br>\n</p>",
              "rawMarkdown": "~~but i not sure why the bugged code work~~\n~~i now redo my experiments with nanmean~~"
            },
            {
              "id": 2204886,
              "postDate": "2023-04-01T04:17:58.660Z",
              "content": "<p>Glad to help. I'm not sure how to do the same for unit variance as there is no such thing as <code>nanstd</code>. Please share if you know.</p>",
              "rawMarkdown": "Glad to help. I'm not sure how to do the same for unit variance as there is no such thing as `nanstd`. Please share if you know."
            },
            {
              "id": 2204946,
              "postDate": "2023-04-01T05:07:46.860Z",
              "content": "<p>i make the mistake becuase of <br>\n<a href=\"https://pytorch.org/docs/stable/generated/torch.nanmean.html\" target=\"_blank\">https://pytorch.org/docs/stable/generated/torch.nanmean.html</a></p>\n<p>quote; \"(torch.nanmean(a) is equivalent to torch.mean(a[~a.isnan()])).\"</p>\n<p>i suggest you use np for preparaing pytorch train dataset.<br>\nfor tflite, you will be using tf functions, so it is not a problem<br>\n(tf.experimental.numpy.std)</p>",
              "rawMarkdown": "i make the mistake becuase of \nhttps://pytorch.org/docs/stable/generated/torch.nanmean.html\n\nquote; \"(torch.nanmean(a) is equivalent to torch.mean(a[~a.isnan()])).\"\n\ni suggest you use np for preparaing pytorch train dataset.\nfor tflite, you will be using tf functions, so it is not a problem\n(tf.experimental.numpy.std)"
            },
            {
              "id": 2205036,
              "postDate": "2023-04-01T07:33:15.507Z",
              "content": "<p><a href=\"https://www.kaggle.com/notnitsuj\" target=\"_blank\">@notnitsuj</a> </p>\n<p>sorry for the confusion. the original implementation is \"almost correct\"<br>\n<img src=\"https://i.ibb.co/095CCVD/Selection-999-1671.png\" alt=\"https://i.ibb.co/095CCVD/Selection-999-1671.png\"></p>",
              "rawMarkdown": "@notnitsuj \n\nsorry for the confusion. the original implementation is \"almost correct\"\n![https://i.ibb.co/095CCVD/Selection-999-1671.png](https://i.ibb.co/095CCVD/Selection-999-1671.png)\n"
            },
            {
              "id": 2205057,
              "postDate": "2023-04-01T07:52:39.387Z",
              "content": "<p>I wonder if it's better to normalize per frame, though. I'll test that out. <br>\nSince you're normalizing everything at once, I suggest you can just delete the arguments for <code>mean</code> as they have no effects on a single scalar and only create confusion.<br>\n[Edit] Wait, what do you mean by \"almost correct\"? Can you clarify what kind of normalization are we using here?</p>",
              "rawMarkdown": "I wonder if it's better to normalize per frame, though. I'll test that out. \nSince you're normalizing everything at once, I suggest you can just delete the arguments for `mean` as they have no effects on a single scalar and only create confusion.\n[Edit] Wait, what do you mean by \"almost correct\"? Can you clarify what kind of normalization are we using here?"
            },
            {
              "id": 2205212,
              "postDate": "2023-04-01T10:44:11.257Z",
              "content": "<p>if your xyz is already zero mean, then the average distance of points from center(0,0,0) should be</p>\n<pre><code>var  =  np.nanstd(xyz.reshape(-1,3))**2 # mean of x**2, y**2, zz*2\ndistance_sq  = var.sum()\ndistance = distance_sq**0.5\n</code></pre>\n<p>this should be the normalising factor</p>",
              "rawMarkdown": "if your xyz is already zero mean, then the average distance of points from center(0,0,0) should be\n```\nvar  =  np.nanstd(xyz.reshape(-1,3))**2 # mean of x**2, y**2, zz*2\ndistance_sq  = var.sum()\ndistance = distance_sq**0.5\n```\n\nthis should be the normalising factor",
              "votes": 2
            }
          ]
        },
        {
          "id": 2206355,
          "postDate": "2023-04-02T13:54:54.773Z",
          "content": "<p>By normalising x, y, z separately you meant sth like this? <br>\n<code>xyz.nanmean(dim=(0, 1), keepdim=True)</code>'</p>\n<p>It results in normalisation of each of dimensions separately, so that mean and std for x, y, and z are the same (notice that means are around 0):<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2423410%2F20c87321212b07a3e5d3c1be4eb9ff3f%2FScreenshot%202023-04-02%20at%2015.59.16.png?generation=1680443969118291&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "By normalising x, y, z separately you meant sth like this? \n`xyz.nanmean(dim=(0, 1), keepdim=True)`'\n\nIt results in normalisation of each of dimensions separately, so that mean and std for x, y, and z are the same (notice that means are around 0):\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2423410%2F20c87321212b07a3e5d3c1be4eb9ff3f%2FScreenshot%202023-04-02%20at%2015.59.16.png?generation=1680443969118291&alt=media)\n\n",
          "replies": [
            {
              "id": 2206439,
              "postDate": "2023-04-02T14:49:32.373Z",
              "content": "<p>Yes, but I also wanted to normalize the spatial coordinates for each frame, which means to only reduce dim=1. But it's just my intuition and I'm not sure if it will benefit the model. You can check out hengck23's replies above.</p>\n<p>We're getting the means close to 0 and the stds close to 1 as the original data are already normalized. This step is just for those <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/392286#2171542\" target=\"_blank\">artifacts</a> that are outside of [0,1].</p>",
              "rawMarkdown": "Yes, but I also wanted to normalize the spatial coordinates for each frame, which means to only reduce dim=1. But it's just my intuition and I'm not sure if it will benefit the model. You can check out hengck23's replies above.\n\nWe're getting the means close to 0 and the stds close to 1 as the original data are already normalized. This step is just for those [artifacts](https://www.kaggle.com/competitions/asl-signs/discussion/392286#2171542) that are outside of [0,1]."
            }
          ]
        }
      ]
    },
    {
      "id": 2197663,
      "postDate": "2023-03-26T09:48:45.067Z",
      "content": "<p>[previous results]</p>\n<p>15-mar<br>\n<img src=\"https://i.ibb.co/Yf52njT/Selection-999-1544.png\" alt=\"https://i.ibb.co/Yf52njT/Selection-999-1544.png\"></p>\n<p>03-apr<br>\n<img src=\"https://i.ibb.co/TgWJk32/Selection-999-1677.png\" alt=\"https://i.ibb.co/TgWJk32/Selection-999-1677.png\"></p>",
      "rawMarkdown": "[previous results]\n\n15-mar\n![https://i.ibb.co/Yf52njT/Selection-999-1544.png](https://i.ibb.co/Yf52njT/Selection-999-1544.png)\n\n03-apr\n![https://i.ibb.co/TgWJk32/Selection-999-1677.png](https://i.ibb.co/TgWJk32/Selection-999-1677.png)\n",
      "votes": 1
    },
    {
      "id": 2185116,
      "postDate": "2023-03-16T20:24:55.600Z",
      "content": "<p>i just realise my work flow is:</p>\n<p>1.click edit notebook  <br>\n2.click open data in new tab  </p>\n<ul>\n<li>upload new tflite file and wait  </li>\n<li>back to notebook page click check for update </li>\n</ul>\n<p>3.click submit  </p>\n<p>i wish to hvae one button just update dataset directly from the notebook page.<br>\nno need all the \"click check for update\"  </p>",
      "rawMarkdown": "i just realise my work flow is:\n\n1.click edit notebook  \n2.click open data in new tab  \n - upload new tflite file and wait  \n - back to notebook page click check for update \n \n3.click submit  \n\ni wish to hvae one button just update dataset directly from the notebook page.\nno need all the \"click check for update\"  ",
      "votes": 1,
      "replies": [
        {
          "id": 2185307,
          "postDate": "2023-03-16T23:47:50.637Z",
          "content": "<p>I wish you could just submit the model.tflite directly like csv submissions. But it's handy enough😗<br>\nI like this new kind of kaggle competitions focused on edge computing. Kudos to the kaggle team making it happen!</p>",
          "rawMarkdown": "I wish you could just submit the model.tflite directly like csv submissions. But it's handy enough😗\nI like this new kind of kaggle competitions focused on edge computing. Kudos to the kaggle team making it happen!",
          "votes": 1,
          "replies": [
            {
              "id": 2185324,
              "postDate": "2023-03-17T00:43:51.023Z",
              "content": "<p><a href=\"https://www.kaggle.com/discussions/general/395353\" target=\"_blank\">https://www.kaggle.com/discussions/general/395353</a></p>\n<p>maybe one day we can have a chatgpt AI bot to help us with that:<br>\nprompt 'upload new tflitemodel and submit'</p>\n<p>and i think this is useful<br>\nprompt 'check license issue for this external dataset at <a href=\"http://www.xxx.xxx\" target=\"_blank\">www.xxx.xxx</a>. if there is an issue, send email to moderator and ask'</p>",
              "rawMarkdown": "https://www.kaggle.com/discussions/general/395353\n\nmaybe one day we can have a chatgpt AI bot to help us with that:\nprompt 'upload new tflitemodel and submit'\n\nand i think this is useful\nprompt 'check license issue for this external dataset at www.xxx.xxx. if there is an issue, send email to moderator and ask'"
            }
          ]
        }
      ]
    },
    {
      "id": 2182036,
      "postDate": "2023-03-14T22:33:06.623Z",
      "content": "<p>did another try mirror augmentation?<br>\nif yes, can i have your feedbacks?</p>",
      "rawMarkdown": "did another try mirror augmentation?\nif yes, can i have your feedbacks?",
      "votes": 1,
      "replies": [
        {
          "id": 2184692,
          "postDate": "2023-03-16T14:48:04.883Z",
          "content": "<p>I've tried mirror augmentation with a simple MLP model, but didn't get better results for now, I'm still stuck at 0.63LB with and without this augmentation.<br>\nI'll experiment further this weekend and I'm interested in other feedbacks on this topic as I supposed it would mitigate the right-handed/left-handed issue.</p>",
          "rawMarkdown": "I've tried mirror augmentation with a simple MLP model, but didn't get better results for now, I'm still stuck at 0.63LB with and without this augmentation.\nI'll experiment further this weekend and I'm interested in other feedbacks on this topic as I supposed it would mitigate the right-handed/left-handed issue.",
          "replies": [
            {
              "id": 2192090,
              "postDate": "2023-03-22T12:21:33.837Z",
              "content": "<p>How do you implement the mirror augmentation?</p>",
              "rawMarkdown": "How do you implement the mirror augmentation?"
            },
            {
              "id": 2210891,
              "postDate": "2023-04-05T17:13:06.600Z",
              "content": "<p>I simply multiply the x coordinates by -1 after normalization</p>",
              "rawMarkdown": "I simply multiply the x coordinates by -1 after normalization"
            }
          ]
        }
      ]
    },
    {
      "id": 2176454,
      "postDate": "2023-03-10T17:03:18.613Z",
      "content": "<p>[pretrain model]</p>\n<p><a href=\"https://github.com/AI4Bharat/OpenHands\" target=\"_blank\">https://github.com/AI4Bharat/OpenHands</a><br>\npretrain model based on mediapipe !!!!</p>",
      "rawMarkdown": "[pretrain model]\n\nhttps://github.com/AI4Bharat/OpenHands\npretrain model based on mediapipe !!!!",
      "votes": 1,
      "replies": [
        {
          "id": 2176596,
          "postDate": "2023-03-10T19:06:17.637Z",
          "content": "<p>this is one thing what i want to do:<br>\ngenerative pretraining<br>\n[1]SignBERT: Pre-Training of Hand-Model-Aware Representation for Sign Language Recognition<br>\n<a href=\"https://arxiv.org/pdf/2110.05382.pdf\" target=\"_blank\">https://arxiv.org/pdf/2110.05382.pdf</a></p>\n<p><img src=\"https://i.ibb.co/bd6YCGS/Selection-999-1327.png\" alt=\"https://i.ibb.co/bd6YCGS/Selection-999-1327.png\"></p>",
          "rawMarkdown": "this is one thing what i want to do:\ngenerative pretraining\n[1]SignBERT: Pre-Training of Hand-Model-Aware Representation for Sign Language Recognition\nhttps://arxiv.org/pdf/2110.05382.pdf\n\n\n![https://i.ibb.co/bd6YCGS/Selection-999-1327.png](https://i.ibb.co/bd6YCGS/Selection-999-1327.png)",
          "votes": 5
        },
        {
          "id": 2177220,
          "postDate": "2023-03-11T09:34:14.987Z",
          "content": "<p>It is interesting! Do you have some thoughts about creating baseline using that pretrained model?</p>",
          "rawMarkdown": "It is interesting! Do you have some thoughts about creating baseline using that pretrained model?",
          "replies": [
            {
              "id": 2178093,
              "postDate": "2023-03-12T05:19:08.527Z",
              "content": "<p>related: tokenization as pretraining<br>\n<img src=\"https://i.ibb.co/ZYrHjCv/Selection-999-1357.png\" alt=\"https://i.ibb.co/ZYrHjCv/Selection-999-1357.png\"></p>\n<p>BEST: BERT Pre-Training for Sign Language Recognition with Coupling Tokenization<br>\n<a href=\"https://arxiv.org/pdf/2302.05075.pdf\" target=\"_blank\">https://arxiv.org/pdf/2302.05075.pdf</a></p>\n<hr>\n<p>diffusion-based autoencoder (for augmentation?)<br>\nVector Quantized Diffusion Model with CodeUnet for Text-to-Sign Pose Sequences Generation<br>\n<a href=\"https://arxiv.org/pdf/2208.09141.pdf\" target=\"_blank\">https://arxiv.org/pdf/2208.09141.pdf</a></p>",
              "rawMarkdown": "related: tokenization as pretraining\n![https://i.ibb.co/ZYrHjCv/Selection-999-1357.png](https://i.ibb.co/ZYrHjCv/Selection-999-1357.png)\n\nBEST: BERT Pre-Training for Sign Language Recognition with Coupling Tokenization\nhttps://arxiv.org/pdf/2302.05075.pdf\n\n---\n\ndiffusion-based autoencoder (for augmentation?)\nVector Quantized Diffusion Model with CodeUnet for Text-to-Sign Pose Sequences Generation\nhttps://arxiv.org/pdf/2208.09141.pdf\n",
              "votes": 3
            },
            {
              "id": 2178130,
              "postDate": "2023-03-12T06:24:05.213Z",
              "content": "<p>It's too large architecture, I think.</p>",
              "rawMarkdown": "It's too large architecture, I think."
            },
            {
              "id": 2178214,
              "postDate": "2023-03-12T08:16:53.277Z",
              "content": "<p>you can distill later</p>",
              "rawMarkdown": "you can distill later",
              "votes": 1
            },
            {
              "id": 2181076,
              "postDate": "2023-03-14T09:48:18.840Z",
              "content": "<p>in theory, the max numbers of hand pattern is fixed(e.g. limited by rotation angle of finger joints, etc). hand you can actually build a \"tokenizer\" (aka lookup table)</p>\n<p>it is interesting to make a tsne plot of the the hands</p>",
              "rawMarkdown": "in theory, the max numbers of hand pattern is fixed(e.g. limited by rotation angle of finger joints, etc). hand you can actually build a \"tokenizer\" (aka lookup table)\n\nit is interesting to make a tsne plot of the the hands",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2165058,
      "postDate": "2023-03-02T00:46:44.770Z",
      "content": "<p>[framework conversion]</p>\n<ul>\n<li><p>keras, pytorch:<br>\n<a href=\"https://github.com/leondgarse/keras_cv_attention_models\" target=\"_blank\">https://github.com/leondgarse/keras_cv_attention_models</a></p></li>\n<li><p>pytorch, onnx, tf<br>\n<a href=\"https://github.com/PINTO0309/onnx2tf\" target=\"_blank\">https://github.com/PINTO0309/onnx2tf</a></p></li>\n</ul>",
      "rawMarkdown": "[framework conversion]\n\n- keras, pytorch:\nhttps://github.com/leondgarse/keras_cv_attention_models\n\n- pytorch, onnx, tf\nhttps://github.com/PINTO0309/onnx2tf",
      "votes": 1
    },
    {
      "id": 2229717,
      "postDate": "2023-04-21T16:15:47.363Z",
      "content": "<p>i check some of the confusing words …. i wonder if a second classifier for specific pairs (or groups) of confusing words would help (e.g. using more features, number of frames, etc)</p>\n<pre><code>167 pen (99)\n     43, 0.43434, 167 pen                     \n     28, 0.28283, 168 pencil                  \n      9, 0.09091,  40 chair                   \n      7, 0.07071, 176 potty                   \n      2, 0.02020,   0 TV        \n\n168 pencil (96)\n     66, 0.68750, 168 pencil                  \n     15, 0.15625, 167 pen                     \n      2, 0.02083,   6 another                 \n      2, 0.02083, 176 potty       \n\n129 kitty (103)\n     77, 0.74757, 129 kitty                   \n     10, 0.09709,  38 cat                     \n      3, 0.02913,  19 bee                     \n      3, 0.02913, 173 please     \n\n\n11 awake (96)\n     46, 0.47917,  11 awake                   \n     34, 0.35417, 233 wake                    \n      7, 0.07292, 131 later            \n</code></pre>",
      "rawMarkdown": "i check some of the confusing words .... i wonder if a second classifier for specific pairs (or groups) of confusing words would help (e.g. using more features, number of frames, etc)\n\n```\n\n167 pen (99)\n\t 43, 0.43434, 167 pen                     \n\t 28, 0.28283, 168 pencil                  \n\t  9, 0.09091,  40 chair                   \n\t  7, 0.07071, 176 potty                   \n\t  2, 0.02020,   0 TV        \n\n168 pencil (96)\n\t 66, 0.68750, 168 pencil                  \n\t 15, 0.15625, 167 pen                     \n\t  2, 0.02083,   6 another                 \n\t  2, 0.02083, 176 potty       \n\n129 kitty (103)\n\t 77, 0.74757, 129 kitty                   \n\t 10, 0.09709,  38 cat                     \n\t  3, 0.02913,  19 bee                     \n\t  3, 0.02913, 173 please     \n\n\n11 awake (96)\n\t 46, 0.47917,  11 awake                   \n\t 34, 0.35417, 233 wake                    \n\t  7, 0.07292, 131 later            \n```\n\n\n\n\n",
      "votes": 2,
      "replies": [
        {
          "id": 2230055,
          "postDate": "2023-04-22T00:56:43.143Z",
          "content": "<p><a href=\"https://arxiv.org/pdf/2303.12080.pdf\" target=\"_blank\">https://arxiv.org/pdf/2303.12080.pdf</a><br>\nNatural Language-Assisted Sign Language Recognition</p>\n<p>you can see that the confused words has similary meaning (e.g. kitty/cat, pencil/pen)<br>\ncheck the paper on how to improve on this.<br>\n(this may the the reason why label smooth work for me, but according to the paper word-based label smooth works better)</p>",
          "rawMarkdown": "https://arxiv.org/pdf/2303.12080.pdf\nNatural Language-Assisted Sign Language Recognition\n\nyou can see that the confused words has similary meaning (e.g. kitty/cat, pencil/pen)\ncheck the paper on how to improve on this.\n(this may the the reason why label smooth work for me, but according to the paper word-based label smooth works better)\n",
          "replies": [
            {
              "id": 2230115,
              "postDate": "2023-04-22T03:23:30.280Z",
              "content": "<p>Can I ask the simple question, which classes you didn't apply flip-agumentations. Thankss </p>",
              "rawMarkdown": "Can I ask the simple question, which classes you didn't apply flip-agumentations. Thankss "
            }
          ]
        }
      ]
    },
    {
      "id": 2211075,
      "postDate": "2023-04-05T19:24:21.707Z",
      "content": "<p>distangle hand shape and hand motion:<br>\nPhonologically-meaningful Subunits for Deep Learning-based Sign Language Recognition<br>\n<a href=\"https://www.slrtp.com/papers/full_papers/SLRTP.FP.02.012.paper.pdf\" target=\"_blank\">https://www.slrtp.com/papers/full_papers/SLRTP.FP.02.012.paper.pdf</a></p>\n<p>motion token?</p>\n<p>\"The large majority of sign language recognition systems based on deep learning adopt a word model approach. Here we present a system that works with subunits, rather than word models. We propose a pipelined approach to deep learning that uses a factorisation algorithm to derive hand motion features, embedded within a low-rank trajectory space.\"</p>\n<p>\"To separate the real hand motions from the camera and whole body movements of the signer, we propose to use a non-rigid structure from motion (NRSfM) technique based on the factorisation method [54].\"</p>",
      "rawMarkdown": "distangle hand shape and hand motion:\nPhonologically-meaningful Subunits for Deep Learning-based Sign Language Recognition\nhttps://www.slrtp.com/papers/full_papers/SLRTP.FP.02.012.paper.pdf\n\nmotion token?\n\n\"The large majority of sign language recognition systems based on deep learning adopt a word model approach. Here we present a system that works with subunits, rather than word models. We propose a pipelined approach to deep learning that uses a factorisation algorithm to derive hand motion features, embedded within a low-rank trajectory space.\"\n\n\n\"To separate the real hand motions from the camera and whole body movements of the signer, we propose to use a non-rigid structure from motion (NRSfM) technique based on the factorisation method [54].\"\n\n",
      "votes": 2
    },
    {
      "id": 2200211,
      "postDate": "2023-03-28T11:23:42.713Z",
      "content": "<p>i probably got it all wrong.<br>\nthe correct way to shape normalise and do feature extraction is given by:</p>\n<p><a href=\"https://google.github.io/mediapipe/solutions/pose_classification.html\" target=\"_blank\">https://google.github.io/mediapipe/solutions/pose_classification.html</a><br>\n<a href=\"https://www.tensorflow.org/lite/tutorials/pose_classification\" target=\"_blank\">https://www.tensorflow.org/lite/tutorials/pose_classification</a></p>",
      "rawMarkdown": "i probably got it all wrong.\nthe correct way to shape normalise and do feature extraction is given by:\n\nhttps://google.github.io/mediapipe/solutions/pose_classification.html\nhttps://www.tensorflow.org/lite/tutorials/pose_classification\n",
      "votes": 2,
      "replies": [
        {
          "id": 2200260,
          "postDate": "2023-03-28T12:22:35.167Z",
          "content": "<pre><code>Next, convert the landmark coordinates to a feature vector by:\n\n- Moving the pose center to the origin.\n- Scaling the pose so that the pose size becomes 1\n- Flattening these coordinates into a feature vector\n</code></pre>\n<p>I am worried about \"scaling the pose so that the pose size becomes 1\". Let's say we are focused on the hand skeleton. If we have flat palm and fist - both of them would be scaled to the size of 1. Would it be beneficial for the model? I guess without per sample scaling it would be possible to extract additional information right from a scale of a given sign. What do you think about it?</p>",
          "rawMarkdown": "```\nNext, convert the landmark coordinates to a feature vector by:\n\n- Moving the pose center to the origin.\n- Scaling the pose so that the pose size becomes 1\n- Flattening these coordinates into a feature vector\n```\n\nI am worried about \"scaling the pose so that the pose size becomes 1\". Let's say we are focused on the hand skeleton. If we have flat palm and fist - both of them would be scaled to the size of 1. Would it be beneficial for the model? I guess without per sample scaling it would be possible to extract additional information right from a scale of a given sign. What do you think about it?",
          "replies": [
            {
              "id": 2202120,
              "postDate": "2023-03-29T18:46:10.453Z",
              "content": "<p>\"i probably got it all wrong.\"</p>\n<p>becuase if your feature is pairwise distance, then your normalisation is also pairwise distance.<br>\nobjective of normalisation is to make feature values the same for similar class</p>",
              "rawMarkdown": "\"i probably got it all wrong.\"\n\nbecuase if your feature is pairwise distance, then your normalisation is also pairwise distance.\nobjective of normalisation is to make feature values the same for similar class"
            },
            {
              "id": 2205289,
              "postDate": "2023-04-01T12:07:27.503Z",
              "content": "<p>Hi, my pairwise distance calculation code reduces training speed dramatically from 20 batches/s to 4 batches/s. Do you precalculate pairwise distances or do you calculate it live? Do you have similar speed reduction?</p>",
              "rawMarkdown": "Hi, my pairwise distance calculation code reduces training speed dramatically from 20 batches/s to 4 batches/s. Do you precalculate pairwise distances or do you calculate it live? Do you have similar speed reduction?"
            },
            {
              "id": 2205319,
              "postDate": "2023-04-01T12:32:13.977Z",
              "content": "<p>If you are using tf.data.Dataset I would recommend you to use .cache() after feature calculation. It will precalculate all features and cache them. All other epochs would use cache instead of calculation all from scatch</p>",
              "rawMarkdown": "If you are using tf.data.Dataset I would recommend you to use .cache() after feature calculation. It will precalculate all features and cache them. All other epochs would use cache instead of calculation all from scatch"
            }
          ]
        }
      ]
    },
    {
      "id": 2189046,
      "postDate": "2023-03-20T07:37:41.097Z",
      "content": "<p>we have an interesting results of LB0.71 with single fold using all train data</p>",
      "rawMarkdown": "we have an interesting results of LB0.71 with single fold using all train data",
      "votes": 2,
      "replies": [
        {
          "id": 2193611,
          "postDate": "2023-03-23T11:37:21.013Z",
          "content": "<p>Yes, I felt it on the ensemble. When adding a new model, the increase in lb to 0.02 is too large. This means that unit models do not extract everything from the data. I used lb models 0.6-0.64 which resulted in 0.69. Good luck</p>",
          "rawMarkdown": "Yes, I felt it on the ensemble. When adding a new model, the increase in lb to 0.02 is too large. This means that unit models do not extract everything from the data. I used lb models 0.6-0.64 which resulted in 0.69. Good luck"
        },
        {
          "id": 2195003,
          "postDate": "2023-03-24T10:34:29.210Z",
          "content": "<p>I reach 0.71 with only one fold just now.🤓🤓</p>",
          "rawMarkdown": "I reach 0.71 with only one fold just now.🤓🤓",
          "votes": 1
        }
      ]
    },
    {
      "id": 2178024,
      "postDate": "2023-03-12T02:33:03.533Z",
      "content": "<p>[opensource implementation]</p>\n<p>[0] ST-GCN <br>\n<a href=\"https://github.com/yysijie/st-gcn\" target=\"_blank\">https://github.com/yysijie/st-gcn</a></p>\n<p>[1] SL-GCN <br>\n<a href=\"https://github.com/jackyjsy/SAM-SLR-v2\" target=\"_blank\">https://github.com/jackyjsy/SAM-SLR-v2</a></p>\n<p>[2] ECCV 2022 Sign Spotting Challenge 3rd ranked solution<br>\n<a href=\"https://chalearnlap.cvc.uab.cat/media/results/None/top3-track1_H1KGRhQ.pdf\" target=\"_blank\">https://chalearnlap.cvc.uab.cat/media/results/None/top3-track1_H1KGRhQ.pdf</a> <br>\n<a href=\"https://github.com/ycmin95/Chalearn_2022_Sign_Spotting_MSSL_track/tree/master/dataset\" target=\"_blank\">https://github.com/ycmin95/Chalearn_2022_Sign_Spotting_MSSL_track/tree/master/dataset</a></p>\n<p>include how to use mediapipe to generate data</p>",
      "rawMarkdown": "[opensource implementation]\n\n[0] ST-GCN \nhttps://github.com/yysijie/st-gcn\n\n[1] SL-GCN \nhttps://github.com/jackyjsy/SAM-SLR-v2\n\n[2] ECCV 2022 Sign Spotting Challenge 3rd ranked solution\nhttps://chalearnlap.cvc.uab.cat/media/results/None/top3-track1_H1KGRhQ.pdf \nhttps://github.com/ycmin95/Chalearn_2022_Sign_Spotting_MSSL_track/tree/master/dataset\n\ninclude how to use mediapipe to generate data\n",
      "votes": 2
    },
    {
      "id": 2163489,
      "postDate": "2023-02-28T21:14:11.653Z",
      "content": "<p>[tflite conversion]</p>\n<ul>\n<li><a href=\"https://github.com/onnx/onnx-tensorflow\" target=\"_blank\">https://github.com/onnx/onnx-tensorflow</a></li>\n<li>fallback plan (aka write my own converter):<ul>\n<li>re-code in keras. need to perform one to one check for each nn.Module(), tf.keras.layers().</li>\n<li>copy weights from pytorch to keras (trick is to give good names to layer for easier code)</li>\n<li>keras to tflite</li></ul></li>\n</ul>",
      "rawMarkdown": "[tflite conversion]\n\n-  https://github.com/onnx/onnx-tensorflow\n-  fallback plan (aka write my own converter):\n    - re-code in keras. need to perform one to one check for each nn.Module(), tf.keras.layers().\n    - copy weights from pytorch to keras (trick is to give good names to layer for easier code)\n    - keras to tflite",
      "votes": 2
    },
    {
      "id": 2230043,
      "postDate": "2023-04-22T00:40:49.027Z",
      "content": "<p>extra attribute for training (aux loss). if you have external training set, you can now use non kaggle gloss (word labels) for training too!</p>\n<p>Towards Zero-shot Sign Language Recognition<br>\n<a href=\"https://arxiv.org/pdf/2201.05914.pdf\" target=\"_blank\">https://arxiv.org/pdf/2201.05914.pdf</a><br>\n<a href=\"https://bmvc2019.org/wp-content/uploads/papers/0122-paper.pdf\" target=\"_blank\">https://bmvc2019.org/wp-content/uploads/papers/0122-paper.pdf</a></p>\n<p>\"We further annotate the datasets with high-level attributes that are gathered from American Sign Language Hand Shape Dictionary [87].\"</p>\n<p>a page from American Sign Language Hand Shape Dictionary:<br>\n<img src=\"https://i.ibb.co/Zzdt317/6e980c3584164c851f4b2d6788bb46626712a27c.jpg\" alt=\"https://i.ibb.co/Zzdt317/6e980c3584164c851f4b2d6788bb46626712a27c.jpg\"></p>\n<p>so you have attribute like:</p>\n<ul>\n<li>palm orientation (in, out, up, down, left, or right)</li>\n<li>movement (up, down, left, right, inward, outward, circular, wrist movement, finger movement)</li>\n<li>hand shape (A-hand, S-hand, 5-hand, etc.)</li>\n</ul>\n<p><img src=\"https://i.ibb.co/3TdzLM5/Selection-999-1856.png\" alt=\"https://i.ibb.co/3TdzLM5/Selection-999-1856.png\"></p>",
      "rawMarkdown": "extra attribute for training (aux loss). if you have external training set, you can now use non kaggle gloss (word labels) for training too!\n\n\nTowards Zero-shot Sign Language Recognition\nhttps://arxiv.org/pdf/2201.05914.pdf\nhttps://bmvc2019.org/wp-content/uploads/papers/0122-paper.pdf\n\n\"We further annotate the datasets with high-level attributes that are gathered from American Sign Language Hand Shape Dictionary [87].\"\n\na page from American Sign Language Hand Shape Dictionary:\n![https://i.ibb.co/Zzdt317/6e980c3584164c851f4b2d6788bb46626712a27c.jpg](https://i.ibb.co/Zzdt317/6e980c3584164c851f4b2d6788bb46626712a27c.jpg)\n\nso you have attribute like:\n- palm orientation (in, out, up, down, left, or right)\n- movement (up, down, left, right, inward, outward, circular, wrist movement, finger movement)\n- hand shape (A-hand, S-hand, 5-hand, etc.)\n\n\n![https://i.ibb.co/3TdzLM5/Selection-999-1856.png](https://i.ibb.co/3TdzLM5/Selection-999-1856.png)\n\n"
    },
    {
      "id": 2229823,
      "postDate": "2023-04-21T18:14:11.380Z",
      "content": "<p>how do you get the point_dim 1414 in the lastest update?</p>",
      "rawMarkdown": "how do you get the point_dim 1414 in the lastest update?"
    },
    {
      "id": 2222116,
      "postDate": "2023-04-14T22:59:25.597Z",
      "content": "<p>there is another way (maybe better?) to approach the problem:<br>\nsegmentation : </p>\n<ul>\n<li>classify every frame  (instead of one video), i.e. making frame features</li>\n<li>you can combine several videos during training</li>\n<li>if you want you can train an aggregator at the end  to combine frame features into video features to do video prediction</li>\n</ul>\n<hr>\n<p>an extension will be classify every N-frame intervals, etc to learn videolet features </p>\n<hr>\n<p>you can using classification to perform vector quantisation (VQ) …  kinda fake AE</p>",
      "rawMarkdown": "there is another way (maybe better?) to approach the problem:\nsegmentation : \n- classify every frame  (instead of one video), i.e. making frame features\n- you can combine several videos during training\n- if you want you can train an aggregator at the end  to combine frame features into video features to do video prediction\n\n---\n\nan extension will be classify every N-frame intervals, etc to learn videolet features \n\n---\n\nyou can using classification to perform vector quantisation (VQ) ...  kinda fake AE"
    },
    {
      "id": 2211093,
      "postDate": "2023-04-05T19:39:46.310Z",
      "content": "<p>joint angle as a feature:<br>\n<a href=\"https://github.com/google/mediapipe/issues/2999\" target=\"_blank\">https://github.com/google/mediapipe/issues/2999</a><br>\n<a href=\"https://github.com/TemugeB/joint_angles_calculate\" target=\"_blank\">https://github.com/TemugeB/joint_angles_calculate</a></p>",
      "rawMarkdown": "joint angle as a feature:\nhttps://github.com/google/mediapipe/issues/2999\nhttps://github.com/TemugeB/joint_angles_calculate"
    },
    {
      "id": 2209841,
      "postDate": "2023-04-05T01:01:04.810Z",
      "content": "<p>Could you say which of participants are in fold 0 in your table?</p>",
      "rawMarkdown": "Could you say which of participants are in fold 0 in your table?",
      "replies": [
        {
          "id": 2216136,
          "postDate": "2023-04-09T19:52:14.493Z",
          "content": "<p>I think its [49445, 61333,  4718,  2044, 37779]</p>",
          "rawMarkdown": "I think its [49445, 61333,  4718,  2044, 37779]"
        }
      ]
    },
    {
      "id": 2209724,
      "postDate": "2023-04-04T21:08:42.600Z",
      "content": "<p>how to augment:</p>\n<p>conceptually:<br>\n<img src=\"https://i.ibb.co/FsskMh7/Selection-999-1690.png\" alt=\"https://i.ibb.co/FsskMh7/Selection-999-1690.png\"></p>\n<p>need to constraint the augmentation correctly.<br>\nmore operations will be stretch shear</p>\n<p>you can observe a few coordinate images of the same word of the same/different signers</p>\n<hr>\n<p>actually one can perform classification on these coordinate images (i.e. treat it as an image problem) if you can render perform 2d convolution fast enough … </p>",
      "rawMarkdown": "how to augment:\n\nconceptually:\n![https://i.ibb.co/FsskMh7/Selection-999-1690.png](https://i.ibb.co/FsskMh7/Selection-999-1690.png)\n\nneed to constraint the augmentation correctly.\nmore operations will be stretch shear\n\nyou can observe a few coordinate images of the same word of the same/different signers\n\n---\n\nactually one can perform classification on these coordinate images (i.e. treat it as an image problem) if you can render perform 2d convolution fast enough ... "
    },
    {
      "id": 2209702,
      "postDate": "2023-04-04T20:46:35.190Z",
      "content": "<p>a possible way to select pairwise distance feature:</p>\n<p>Sequential Attention for Feature Selection<br>\n<a href=\"https://arxiv.org/abs/2209.14881\" target=\"_blank\">https://arxiv.org/abs/2209.14881</a></p>",
      "rawMarkdown": "a possible way to select pairwise distance feature:\n\n\nSequential Attention for Feature Selection\nhttps://arxiv.org/abs/2209.14881"
    },
    {
      "id": 2204403,
      "postDate": "2023-03-31T14:34:11.150Z",
      "content": "<p>Thank you for sharing. Did you use entire dataset as training data? How can you check the accuracy of a model at each checkpoint when using all train samples?</p>",
      "rawMarkdown": "Thank you for sharing. Did you use entire dataset as training data? How can you check the accuracy of a model at each checkpoint when using all train samples?",
      "replies": [
        {
          "id": 2204423,
          "postDate": "2023-03-31T14:59:33.130Z",
          "content": "<p><img src=\"https://i.ibb.co/syVRXGD/Selection-999-1648.png\" alt=\"https://i.ibb.co/syVRXGD/Selection-999-1648.png\"></p>\n<p>this is a bit tricky. but if you can set the correct regularization (and if the data has no noise), you may have this</p>",
          "rawMarkdown": "![https://i.ibb.co/syVRXGD/Selection-999-1648.png](https://i.ibb.co/syVRXGD/Selection-999-1648.png)\n\nthis is a bit tricky. but if you can set the correct regularization (and if the data has no noise), you may have this",
          "votes": 1,
          "replies": [
            {
              "id": 2204436,
              "postDate": "2023-03-31T15:03:47.210Z",
              "content": "<p>my train log:</p>\n<p>you can see my loss are almost flat</p>\n<pre><code>** start training here! **\n   batch_size = 64 \n   experiment = ['tx025', 'run_train_fold0.py']\n                           |---------------- VALID---------|---- TRAIN/BATCH ----------------------\nrate      iter       epoch | loss   top1   top2    top5    | loss                 | time           \n---------------------------------------------------------------------------------------------------\n0.00e+0   00000000*   0.00 | 0.000  0.000  0.0000  0.000   | 0.000  0.000  0.000  |  0 hr 00 min\n1.00e-4   00003351*   3.00 | 3.658  0.477  0.6162  0.760   | 5.278  0.000  0.000  |  0 hr 08 min\n1.00e-4   00006702*   6.00 | 3.250  0.574  0.7079  0.823   | 5.191  0.000  0.000  |  0 hr 16 min\n1.00e-4   00010053*   9.00 | 3.088  0.621  0.7418  0.845   | 5.148  0.000  0.000  |  0 hr 24 min\n1.00e-4   00013404*  12.00 | 3.047  0.631  0.7559  0.850   | 5.120  0.000  0.000  |  0 hr 32 min\n1.00e-4   00016755*  15.00 | 3.012  0.643  0.7622  0.854   | 5.109  0.000  0.000  |  0 hr 40 min\n1.00e-4   00020106*  18.00 | 3.002  0.645  0.7631  0.853   | 5.095  0.000  0.000  |  0 hr 48 min\n1.00e-4   00023457*  21.00 | 2.970  0.645  0.7624  0.855   | 5.083  0.000  0.000  |  0 hr 56 min\n1.00e-4   00026808*  24.00 | 2.977  0.654  0.7672  0.858   | 5.075  0.000  0.000  |  1 hr 04 min\n1.00e-4   00030159*  27.00 | 2.972  0.655  0.7690  0.856   | 5.068  0.000  0.000  |  1 hr 12 min\n1.00e-4   00033510*  30.00 | 2.953  0.660  0.7722  0.858   | 5.060  0.000  0.000  |  1 hr 20 min\n1.00e-4   00036861*  33.00 | 2.949  0.663  0.7742  0.860   | 5.058  0.000  0.000  |  1 hr 28 min\n1.00e-4   00040212*  36.00 | 2.952  0.663  0.7802  0.863   | 5.048  0.000  0.000  |  1 hr 35 min\n1.00e-4   00043563*  39.00 | 2.958  0.669  0.7794  0.859   | 5.047  0.000  0.000  |  1 hr 43 min\n1.00e-4   00046914*  42.00 | 3.010  0.663  0.7742  0.856   | 5.045  0.000  0.000  |  1 hr 51 min\n1.00e-4   00050265*  45.00 | 2.954  0.665  0.7728  0.858   | 5.040  0.000  0.000  |  1 hr 59 min\n1.00e-4   00053616*  48.00 | 2.976  0.665  0.7741  0.858   | 5.035  0.000  0.000  |  2 hr 07 min\n1.00e-4   00056967*  51.00 | 3.035  0.668  0.7744  0.857   | 5.028  0.000  0.000  |  2 hr 15 min\n1.00e-4   00060318*  54.00 | 2.952  0.669  0.7762  0.862   | 5.030  0.000  0.000  |  2 hr 23 min\n1.00e-4   00063669*  57.00 | 2.966  0.667  0.7776  0.861   | 5.023  0.000  0.000  |  2 hr 31 min\n1.00e-4   00067020*  60.00 | 2.973  0.671  0.7791  0.860   | 5.028  0.000  0.000  |  2 hr 39 min\n1.00e-4   00070371*  63.00 | 2.954  0.671  0.7773  0.861   | 5.020  0.000  0.000  |  2 hr 47 min\n1.00e-4   00073722*  66.00 | 2.966  0.666  0.7757  0.858   | 5.018  0.000  0.000  |  2 hr 55 min\n1.00e-4   00077073*  69.00 | 3.002  0.667  0.7763  0.857   | 5.016  0.000  0.000  |  3 hr 03 min\n1.00e-4   00080424*  72.00 | 2.986  0.670  0.7760  0.859   | 5.013  0.000  0.000  |  3 hr 12 min\n1.00e-4   00083775*  75.00 | 3.026  0.673  0.7779  0.858   | 5.011  0.000  0.000  |  3 hr 20 min\n1.00e-4   00087126*  78.00 | 2.970  0.671  0.7793  0.858   | 5.009  0.000  0.000  |  3 hr 28 min\n1.00e-4   00090477*  81.00 | 2.983  0.676  0.7801  0.858   | 5.008  0.000  0.000  |  3 hr 36 min\n1.00e-4   00093828*  84.00 | 2.952  0.677  0.7807  0.861   | 5.006  0.000  0.000  |  3 hr 45 min\n1.00e-4   00097179*  87.00 | 3.004  0.667  0.7775  0.858   | 5.004  0.000  0.000  |  3 hr 53 min\n1.00e-4   00100530*  90.00 | 2.985  0.674  0.7786  0.859   | 5.003  0.000  0.000  |  4 hr 01 min\n1.00e-4   00103881*  93.00 | 3.067  0.655  0.7676  0.850   | 5.001  0.000  0.000  |  4 hr 09 min\n1.00e-4   00107232*  96.00 | 2.987  0.674  0.7768  0.856   | 4.997  0.000  0.000  |  4 hr 16 min\n1.00e-4   00110583*  99.00 | 2.980  0.673  0.7795  0.858   | 5.000  0.000  0.000  |  4 hr 24 min\n1.00e-4   00113934* 102.00 | 3.001  0.669  0.7751  0.856   | 4.994  0.000  0.000  |  4 hr 32 min\n1.00e-4   00117285* 105.00 | 3.035  0.666  0.7719  0.852   | 4.994  0.000  0.000  |  4 hr 40 min\n1.00e-4   00120636* 108.00 | 3.020  0.668  0.7744  0.855   | 4.992  0.000  0.000  |  4 hr 48 min\n1.00e-4   00123987* 111.00 | 3.019  0.670  0.7756  0.858   | 4.994  0.000  0.000  |  4 hr 55 min\n1.00e-4   00127338* 114.00 | 2.999  0.676  0.7800  0.858   | 4.991  0.000  0.000  |  5 hr 03 min\n1.00e-4   00130689* 117.00 | 2.995  0.674  0.7771  0.856   | 4.990  0.000  0.000  |  5 hr 11 min\n1.00e-4   00134040* 120.00 | 3.037  0.669  0.7742  0.850   | 4.988  0.000  0.000  |  5 hr 19 min\n1.00e-4   00137391* 123.00 | 3.013  0.667  0.7762  0.855   | 4.988  0.000  0.000  |  5 hr 27 min\n1.00e-4   00140742* 126.00 | 3.045  0.668  0.7736  0.853   | 4.984  0.000  0.000  |  5 hr 35 min\n1.00e-4   00144093* 129.00 | 3.013  0.673  0.7757  0.854   | 4.985  0.000  0.000  |  5 hr 43 min\n1.00e-4   00147444* 132.00 | 3.051  0.674  0.7790  0.854   | 4.980  0.000  0.000  |  5 hr 50 min\n1.00e-4   00150795* 135.00 | 3.060  0.667  0.7693  0.849   | 4.983  0.000  0.000  |  5 hr 58 min\n1.00e-4   00154146* 138.00 | 3.054  0.669  0.7747  0.852   | 4.982  0.000  0.000  |  6 hr 06 min\n1.00e-4   00157497* 141.00 | 3.038  0.669  0.7746  0.852   | 4.981  0.000  0.000  |  6 hr 14 min\n</code></pre>",
              "rawMarkdown": "my train log:\n\nyou can see my loss are almost flat\n```\n** start training here! **\n   batch_size = 64 \n   experiment = ['tx025', 'run_train_fold0.py']\n                           |---------------- VALID---------|---- TRAIN/BATCH ----------------------\nrate      iter       epoch | loss   top1   top2    top5    | loss                 | time           \n---------------------------------------------------------------------------------------------------\n0.00e+0   00000000*   0.00 | 0.000  0.000  0.0000  0.000   | 0.000  0.000  0.000  |  0 hr 00 min\n1.00e-4   00003351*   3.00 | 3.658  0.477  0.6162  0.760   | 5.278  0.000  0.000  |  0 hr 08 min\n1.00e-4   00006702*   6.00 | 3.250  0.574  0.7079  0.823   | 5.191  0.000  0.000  |  0 hr 16 min\n1.00e-4   00010053*   9.00 | 3.088  0.621  0.7418  0.845   | 5.148  0.000  0.000  |  0 hr 24 min\n1.00e-4   00013404*  12.00 | 3.047  0.631  0.7559  0.850   | 5.120  0.000  0.000  |  0 hr 32 min\n1.00e-4   00016755*  15.00 | 3.012  0.643  0.7622  0.854   | 5.109  0.000  0.000  |  0 hr 40 min\n1.00e-4   00020106*  18.00 | 3.002  0.645  0.7631  0.853   | 5.095  0.000  0.000  |  0 hr 48 min\n1.00e-4   00023457*  21.00 | 2.970  0.645  0.7624  0.855   | 5.083  0.000  0.000  |  0 hr 56 min\n1.00e-4   00026808*  24.00 | 2.977  0.654  0.7672  0.858   | 5.075  0.000  0.000  |  1 hr 04 min\n1.00e-4   00030159*  27.00 | 2.972  0.655  0.7690  0.856   | 5.068  0.000  0.000  |  1 hr 12 min\n1.00e-4   00033510*  30.00 | 2.953  0.660  0.7722  0.858   | 5.060  0.000  0.000  |  1 hr 20 min\n1.00e-4   00036861*  33.00 | 2.949  0.663  0.7742  0.860   | 5.058  0.000  0.000  |  1 hr 28 min\n1.00e-4   00040212*  36.00 | 2.952  0.663  0.7802  0.863   | 5.048  0.000  0.000  |  1 hr 35 min\n1.00e-4   00043563*  39.00 | 2.958  0.669  0.7794  0.859   | 5.047  0.000  0.000  |  1 hr 43 min\n1.00e-4   00046914*  42.00 | 3.010  0.663  0.7742  0.856   | 5.045  0.000  0.000  |  1 hr 51 min\n1.00e-4   00050265*  45.00 | 2.954  0.665  0.7728  0.858   | 5.040  0.000  0.000  |  1 hr 59 min\n1.00e-4   00053616*  48.00 | 2.976  0.665  0.7741  0.858   | 5.035  0.000  0.000  |  2 hr 07 min\n1.00e-4   00056967*  51.00 | 3.035  0.668  0.7744  0.857   | 5.028  0.000  0.000  |  2 hr 15 min\n1.00e-4   00060318*  54.00 | 2.952  0.669  0.7762  0.862   | 5.030  0.000  0.000  |  2 hr 23 min\n1.00e-4   00063669*  57.00 | 2.966  0.667  0.7776  0.861   | 5.023  0.000  0.000  |  2 hr 31 min\n1.00e-4   00067020*  60.00 | 2.973  0.671  0.7791  0.860   | 5.028  0.000  0.000  |  2 hr 39 min\n1.00e-4   00070371*  63.00 | 2.954  0.671  0.7773  0.861   | 5.020  0.000  0.000  |  2 hr 47 min\n1.00e-4   00073722*  66.00 | 2.966  0.666  0.7757  0.858   | 5.018  0.000  0.000  |  2 hr 55 min\n1.00e-4   00077073*  69.00 | 3.002  0.667  0.7763  0.857   | 5.016  0.000  0.000  |  3 hr 03 min\n1.00e-4   00080424*  72.00 | 2.986  0.670  0.7760  0.859   | 5.013  0.000  0.000  |  3 hr 12 min\n1.00e-4   00083775*  75.00 | 3.026  0.673  0.7779  0.858   | 5.011  0.000  0.000  |  3 hr 20 min\n1.00e-4   00087126*  78.00 | 2.970  0.671  0.7793  0.858   | 5.009  0.000  0.000  |  3 hr 28 min\n1.00e-4   00090477*  81.00 | 2.983  0.676  0.7801  0.858   | 5.008  0.000  0.000  |  3 hr 36 min\n1.00e-4   00093828*  84.00 | 2.952  0.677  0.7807  0.861   | 5.006  0.000  0.000  |  3 hr 45 min\n1.00e-4   00097179*  87.00 | 3.004  0.667  0.7775  0.858   | 5.004  0.000  0.000  |  3 hr 53 min\n1.00e-4   00100530*  90.00 | 2.985  0.674  0.7786  0.859   | 5.003  0.000  0.000  |  4 hr 01 min\n1.00e-4   00103881*  93.00 | 3.067  0.655  0.7676  0.850   | 5.001  0.000  0.000  |  4 hr 09 min\n1.00e-4   00107232*  96.00 | 2.987  0.674  0.7768  0.856   | 4.997  0.000  0.000  |  4 hr 16 min\n1.00e-4   00110583*  99.00 | 2.980  0.673  0.7795  0.858   | 5.000  0.000  0.000  |  4 hr 24 min\n1.00e-4   00113934* 102.00 | 3.001  0.669  0.7751  0.856   | 4.994  0.000  0.000  |  4 hr 32 min\n1.00e-4   00117285* 105.00 | 3.035  0.666  0.7719  0.852   | 4.994  0.000  0.000  |  4 hr 40 min\n1.00e-4   00120636* 108.00 | 3.020  0.668  0.7744  0.855   | 4.992  0.000  0.000  |  4 hr 48 min\n1.00e-4   00123987* 111.00 | 3.019  0.670  0.7756  0.858   | 4.994  0.000  0.000  |  4 hr 55 min\n1.00e-4   00127338* 114.00 | 2.999  0.676  0.7800  0.858   | 4.991  0.000  0.000  |  5 hr 03 min\n1.00e-4   00130689* 117.00 | 2.995  0.674  0.7771  0.856   | 4.990  0.000  0.000  |  5 hr 11 min\n1.00e-4   00134040* 120.00 | 3.037  0.669  0.7742  0.850   | 4.988  0.000  0.000  |  5 hr 19 min\n1.00e-4   00137391* 123.00 | 3.013  0.667  0.7762  0.855   | 4.988  0.000  0.000  |  5 hr 27 min\n1.00e-4   00140742* 126.00 | 3.045  0.668  0.7736  0.853   | 4.984  0.000  0.000  |  5 hr 35 min\n1.00e-4   00144093* 129.00 | 3.013  0.673  0.7757  0.854   | 4.985  0.000  0.000  |  5 hr 43 min\n1.00e-4   00147444* 132.00 | 3.051  0.674  0.7790  0.854   | 4.980  0.000  0.000  |  5 hr 50 min\n1.00e-4   00150795* 135.00 | 3.060  0.667  0.7693  0.849   | 4.983  0.000  0.000  |  5 hr 58 min\n1.00e-4   00154146* 138.00 | 3.054  0.669  0.7747  0.852   | 4.982  0.000  0.000  |  6 hr 06 min\n1.00e-4   00157497* 141.00 | 3.038  0.669  0.7746  0.852   | 4.981  0.000  0.000  |  6 hr 14 min\n\n```",
              "votes": 3
            },
            {
              "id": 2204457,
              "postDate": "2023-03-31T15:28:02.627Z",
              "content": "<p>check also the loss curve at: <a href=\"https://www.kaggle.com/code/sabinaabdurakhmanova/transformer-training-with-learnable-weight-pe\" target=\"_blank\">https://www.kaggle.com/code/sabinaabdurakhmanova/transformer-training-with-learnable-weight-pe</a></p>",
              "rawMarkdown": "check also the loss curve at: https://www.kaggle.com/code/sabinaabdurakhmanova/transformer-training-with-learnable-weight-pe"
            },
            {
              "id": 2204464,
              "postDate": "2023-03-31T15:32:47.537Z",
              "content": "<p>Edit: I think this is possibly due to number of frames - TF code has only 32 frames and with PyTorch I was trying much higher number of frames like 256/512. </p>\n<p>This is interesting <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> , it seems one epoch is taking 8mins. I converted <a href=\"https://www.kaggle.com/code/markwijkhuizen/gislr-tf-data-processing-transformer-training\" target=\"_blank\">this code</a> to PyTorch and it ran in like &lt;30sec per epoch. (I haven't been able to create inference yet, which is a different problem) ~~</p>\n<p>My current dataset is based on your dataset and code posted <a href=\"https://www.kaggle.com/datasets/hengck23/asl-demo\" target=\"_blank\">here</a>. However when I try to use .npy file / pickle file with this best I get is ~2.5mins per epoch and I haven't understood why.  Any thoughts much appreciated, Thank you 🙏</p>",
              "rawMarkdown": "Edit: I think this is possibly due to number of frames - TF code has only 32 frames and with PyTorch I was trying much higher number of frames like 256/512. \n\nThis is interesting @hengck23 , it seems one epoch is taking 8mins. I converted [this code](https://www.kaggle.com/code/markwijkhuizen/gislr-tf-data-processing-transformer-training) to PyTorch and it ran in like <30sec per epoch. (I haven't been able to create inference yet, which is a different problem) ~~\n\nMy current dataset is based on your dataset and code posted [here](https://www.kaggle.com/datasets/hengck23/asl-demo). However when I try to use .npy file / pickle file with this best I get is ~2.5mins per epoch and I haven't understood why.  Any thoughts much appreciated, Thank you 🙏"
            },
            {
              "id": 2205038,
              "postDate": "2023-04-01T07:36:22.143Z",
              "content": "<p>it depends on if embed dim and point dim.</p>\n<p>augment that use scipy interpolation is slow, etc</p>",
              "rawMarkdown": "it depends on if embed dim and point dim.\n\naugment that use scipy interpolation is slow, etc",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2203287,
      "postDate": "2023-03-30T17:05:00.597Z",
      "content": "<p>I am getting out of memory error, even when taking emb_dim = 256, head = 4, transformer_layer = 1. any optimization i am missing out on?</p>",
      "rawMarkdown": "I am getting out of memory error, even when taking emb_dim = 256, head = 4, transformer_layer = 1. any optimization i am missing out on?",
      "replies": [
        {
          "id": 2203309,
          "postDate": "2023-03-30T17:33:06.200Z",
          "content": "<p>I had similar problem, make sure to use TF Input layer rather than pytorch <br>\nSee this - <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/390935#2176801\" target=\"_blank\">https://www.kaggle.com/competitions/asl-signs/discussion/390935#2176801</a> </p>\n<p>Edit:<a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/391265#2188739\" target=\"_blank\"> see discussion below </a></p>",
          "rawMarkdown": "I had similar problem, make sure to use TF Input layer rather than pytorch \nSee this - https://www.kaggle.com/competitions/asl-signs/discussion/390935#2176801 \n\nEdit:[ see discussion below ](https://www.kaggle.com/competitions/asl-signs/discussion/391265#2188739)",
          "replies": [
            {
              "id": 2205331,
              "postDate": "2023-04-01T12:38:25.987Z",
              "content": "<p>Thanks !!!<br>\nhow you load(reading file and preprocessing) data, it is taking a lot of time,   am loading data during training. for 1 epoch it takes 1 hour on Kaggle gpu. </p>",
              "rawMarkdown": "Thanks !!!\nhow you load(reading file and preprocessing) data, it is taking a lot of time,   am loading data during training. for 1 epoch it takes 1 hour on Kaggle gpu. "
            }
          ]
        }
      ]
    },
    {
      "id": 2202934,
      "postDate": "2023-03-30T12:07:27.730Z",
      "content": "<p>May I ask are you still using 1-layer transformer? I can't believe that one attention pooling can achieve 73 single-fold score. </p>\n<p>Or maybe your deeper embedding layer (fc-norm-act-fc-norm-act) can do the same feature processing as more transformer layers?</p>",
      "rawMarkdown": "May I ask are you still using 1-layer transformer? I can't believe that one attention pooling can achieve 73 single-fold score. \n\nOr maybe your deeper embedding layer (fc-norm-act-fc-norm-act) can do the same feature processing as more transformer layers?",
      "replies": [
        {
          "id": 2202945,
          "postDate": "2023-03-30T12:15:06.343Z",
          "content": "<p>\"using 1-layer transformer? \"<br>\nyes.</p>\n<p>the embedding (below) is not very deep either</p>\n<pre><code>class XEmbed(nn.Module):\n    def __init__(self,\n    ):\n        super().__init__()\n        self.v = nn.Sequential(\n            nn.Linear(point_dim, embed_dim*2, bias=True),\n            nn.LayerNorm(embed_dim*2),\n            nn.ReLU(inplace=True),\n            nn.Linear(embed_dim*2, embed_dim, bias=True),\n            nn.LayerNorm(embed_dim),\n            nn.ReLU(inplace=True),\n        )  \n    def forward(self, x, x_mask):\n        B,L,_ = x.shape\n        v = self.v(x)\n        x = v\n        return x, x_mask\n</code></pre>",
          "rawMarkdown": "\"using 1-layer transformer? \"\nyes.\n\nthe embedding (below) is not very deep either\n\n```\nclass XEmbed(nn.Module):\n\tdef __init__(self,\n\t):\n\t\tsuper().__init__()\n\t\tself.v = nn.Sequential(\n\t\t\tnn.Linear(point_dim, embed_dim*2, bias=True),\n\t\t\tnn.LayerNorm(embed_dim*2),\n\t\t\tnn.ReLU(inplace=True),\n\t\t\tnn.Linear(embed_dim*2, embed_dim, bias=True),\n\t\t\tnn.LayerNorm(embed_dim),\n\t\t\tnn.ReLU(inplace=True),\n\t\t)  \n\tdef forward(self, x, x_mask):\n\t\tB,L,_ = x.shape\n\t\tv = self.v(x)\n\t\tx = v\n\t\treturn x, x_mask\n```",
          "votes": 2,
          "replies": [
            {
              "id": 2202963,
              "postDate": "2023-03-30T12:25:20.450Z",
              "content": "<p>Got it, seems feature preprocessing is more important than cross-frame interaction </p>",
              "rawMarkdown": "Got it, seems feature preprocessing is more important than cross-frame interaction "
            },
            {
              "id": 2202972,
              "postDate": "2023-03-30T12:33:43.527Z",
              "content": "<p>here is the full netowrk structure for LB0.73 single fold</p>\n<p><a href=\"https://www.kaggle.com/code/hengck23/lb-0-73-single-fold-transformer-architecture\" target=\"_blank\">https://www.kaggle.com/code/hengck23/lb-0-73-single-fold-transformer-architecture</a></p>",
              "rawMarkdown": "here is the full netowrk structure for LB0.73 single fold\n\nhttps://www.kaggle.com/code/hengck23/lb-0-73-single-fold-transformer-architecture",
              "votes": 1
            },
            {
              "id": 2202977,
              "postDate": "2023-03-30T12:40:41.453Z",
              "content": "<p>\" seems feature preprocessing is more important than cross-frame interaction\"</p>\n<p>note that my input feature include frame difference motion feature.</p>\n<p>it really depends how much data we have. maybe just 21 signers (in train data) is not enough to learn the motion features. (if they all move differently) … or maybe the sign words we used don't have that much motion?</p>\n<p>i believe if we can get external data (i.e. use media pipe to generator your own data) or use GAN, PDDM, augmentation, etc. motion will become important  and i am expecting LB of 0.80</p>",
              "rawMarkdown": "\" seems feature preprocessing is more important than cross-frame interaction\"\n\nnote that my input feature include frame difference motion feature.\n\nit really depends how much data we have. maybe just 21 signers (in train data) is not enough to learn the motion features. (if they all move differently) ... or maybe the sign words we used don't have that much motion?\n\ni believe if we can get external data (i.e. use media pipe to generator your own data) or use GAN, PDDM, augmentation, etc. motion will become important  and i am expecting LB of 0.80\n"
            },
            {
              "id": 2202980,
              "postDate": "2023-03-30T12:43:23.397Z",
              "content": "<p>Amazing work! I use 9-layer transformer to reach 0.72 single fold.</p>",
              "rawMarkdown": "Amazing work! I use 9-layer transformer to reach 0.72 single fold."
            },
            {
              "id": 2203191,
              "postDate": "2023-03-30T15:44:17.403Z",
              "content": "<p>Having a Dense layer for embedding calculation, didn't we lose temporal axis here?</p>",
              "rawMarkdown": "Having a Dense layer for embedding calculation, didn't we lose temporal axis here?"
            },
            {
              "id": 2203306,
              "postDate": "2023-03-30T17:28:17.947Z",
              "content": "<p>\"Amazing work! I use 9-layer transformer\"</p>\n<p>out of curiousity, i tried a 4-layer transformer. Local CV<br>\n1-layer:  0.706302539309203<br>\n4-layer:  0.7132279280456466</p>\n<p>it is either the same or slightly better (note that slightly different hyper-parameters used).</p>",
              "rawMarkdown": "\"Amazing work! I use 9-layer transformer\"\n\nout of curiousity, i tried a 4-layer transformer. Local CV\n1-layer:  0.706302539309203\n4-layer:  0.7132279280456466\n\nit is either the same or slightly better (note that slightly different hyper-parameters used).\n "
            },
            {
              "id": 2205366,
              "postDate": "2023-04-01T13:35:07.010Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2205487,
              "postDate": "2023-04-01T15:54:15.407Z",
              "rawMarkdown": "",
              "votes": 1,
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 2202855,
      "postDate": "2023-03-30T11:24:00.570Z",
      "content": "<p>Label smoothing  and flip didn't work for me.</p>",
      "rawMarkdown": "Label smoothing  and flip didn't work for me."
    },
    {
      "id": 2202383,
      "postDate": "2023-03-30T02:40:57.767Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> , what is cls and what is all token mean pooling?</p>",
      "rawMarkdown": "Hi @hengck23 , what is cls and what is all token mean pooling?"
    },
    {
      "id": 2194498,
      "postDate": "2023-03-24T01:46:16.067Z",
      "content": "<p>i find it easier to write your own attention.<br>\nyou can control q,k,v dim differently etc.<br>\nyou can control local attention, e.g. which query is attent to which values (according hand, lip parts or joints)</p>\n<p>you can learned the affinity matrix of graph CGN if you treat affinity=attention weights</p>",
      "rawMarkdown": "i find it easier to write your own attention.\nyou can control q,k,v dim differently etc.\nyou can control local attention, e.g. which query is attent to which values (according hand, lip parts or joints)\n\nyou can learned the affinity matrix of graph CGN if you treat affinity=attention weights"
    },
    {
      "id": 2189929,
      "postDate": "2023-03-20T22:31:03.683Z",
      "content": "<p>Many thanks for you work <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>.<br>\nDo you ensemble your models or do you just take the best model through CV?<br>\nWhen I ensemble, it seems that inference exceeds an hour…</p>",
      "rawMarkdown": "Many thanks for you work @hengck23.\nDo you ensemble your models or do you just take the best model through CV?\nWhen I ensemble, it seems that inference exceeds an hour..."
    },
    {
      "id": 2189299,
      "postDate": "2023-03-20T11:43:27.953Z",
      "content": "<p>can you do validation without validation data?<br>\ne.g. if i used all the train images for training, how do i know if there is over fitting?</p>\n<p>in theory, yes ….<br>\nyou can still use the following to judge if there is overfitting:</p>\n<ol>\n<li>we usually have a training set and validation set. then we monitor the training and validation loss.</li>\n<li>But there are other indicators that is correlated to degree of overfitting. Validation loss is only one of them</li>\n<li>say if you have only train data, you can measure </li>\n</ol>\n<ul>\n<li>degree of regularization and effects on train loss (for different models of different design)</li>\n<li>rate of change of loss/prediction as data is perturbed</li>\n<li>rate of change of loss/prediction as parameters is perturbed</li>\n<li>observe attention, heatmap, etc ….</li>\n</ul>\n<p>you should think of how these indicators change for an optimum, over-fitted, under-fitted trained models.</p>\n<p>But of course having another set of data is still the best way to check over-fitting. so instead of finding methods and indicators, you may just find another set of data (external or synthetically created)</p>",
      "rawMarkdown": "can you do validation without validation data?\ne.g. if i used all the train images for training, how do i know if there is over fitting?\n\nin theory, yes ....\nyou can still use the following to judge if there is overfitting:\n1. we usually have a training set and validation set. then we monitor the training and validation loss.\n2. But there are other indicators that is correlated to degree of overfitting. Validation loss is only one of them\n3. say if you have only train data, you can measure \n- degree of regularization and effects on train loss (for different models of different design)\n- rate of change of loss/prediction as data is perturbed\n- rate of change of loss/prediction as parameters is perturbed\n- observe attention, heatmap, etc ....\n\nyou should think of how these indicators change for an optimum, over-fitted, under-fitted trained models.\n\nBut of course having another set of data is still the best way to check over-fitting. so instead of finding methods and indicators, you may just find another set of data (external or synthetically created)\n\n\n"
    },
    {
      "id": 2188739,
      "postDate": "2023-03-20T00:06:49.330Z",
      "content": "<p>Thank you so much for sharing your approach <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>  I was able to train the model from <a href=\"https://www.kaggle.com/datasets/hengck23/asl-demo\" target=\"_blank\">here</a> <br>\nI pretty much changed no settings except epochs set to 20, because I want the pipeline to work first. <br>\nI used the <a href=\"https://www.kaggle.com/code/hengck23/lb-0-65-one-pytorch-transformer-solution\" target=\"_blank\">inference</a> however I'm getting OOM even with <code>max_length=40</code>  <br>\nAny suggestions ? </p>",
      "rawMarkdown": "Thank you so much for sharing your approach @hengck23  I was able to train the model from [here](https://www.kaggle.com/datasets/hengck23/asl-demo) \nI pretty much changed no settings except epochs set to 20, because I want the pipeline to work first. \nI used the [inference](https://www.kaggle.com/code/hengck23/lb-0-65-one-pytorch-transformer-solution) however I'm getting OOM even with `max_length=40`  \nAny suggestions ? ",
      "replies": [
        {
          "id": 2188788,
          "postDate": "2023-03-20T01:40:20.897Z",
          "content": "<p><a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/390935#2176801\" target=\"_blank\">https://www.kaggle.com/competitions/asl-signs/discussion/390935#2176801</a><br>\n<a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/393655#2177124\" target=\"_blank\">https://www.kaggle.com/competitions/asl-signs/discussion/393655#2177124</a></p>\n<pre><code>class InputNet(tf.keras.layers.Layer):\n    def __init__(self, ):\n        super(InputNet, self).__init__()\n        self.lip = tf.constant([ \n......\n\n#if your input net is a keras model, you can directly use ....\nclass TFModel(tf.Module):\n    def __init__(self,):\n        super(TFModel, self).__init__()\n        self.input_net = InputNet()\n        self.single_net =  .... e.g. load from file ...\n\n\n    @tf.function(input_signature=[tf.TensorSpec(shape=[None, 543, 3], dtype=tf.float32, name='inputs')])\n    def __call__(self, inputs):\n        outputs = self.input_net(inputs) \n                .... \n</code></pre>",
          "rawMarkdown": "https://www.kaggle.com/competitions/asl-signs/discussion/390935#2176801\nhttps://www.kaggle.com/competitions/asl-signs/discussion/393655#2177124\n\n```\nclass InputNet(tf.keras.layers.Layer):\n\tdef __init__(self, ):\n\t\tsuper(InputNet, self).__init__()\n\t\tself.lip = tf.constant([ \n......\n\n#if your input net is a keras model, you can directly use ....\nclass TFModel(tf.Module):\n\tdef __init__(self,):\n\t\tsuper(TFModel, self).__init__()\n\t\tself.input_net = InputNet()\n\t\tself.single_net =  .... e.g. load from file ...\n\n\n\t@tf.function(input_signature=[tf.TensorSpec(shape=[None, 543, 3], dtype=tf.float32, name='inputs')])\n\tdef __call__(self, inputs):\n\t\toutputs = self.input_net(inputs) \n                .... \n\n```",
          "votes": 1
        },
        {
          "id": 2188797,
          "postDate": "2023-03-20T01:59:21.083Z",
          "content": "<p>i just added a new folder convert1 to <a href=\"https://www.kaggle.com/datasets/hengck23/asl-demo\" target=\"_blank\">https://www.kaggle.com/datasets/hengck23/asl-demo</a> <br>\nit shows how to use keras input model.</p>",
          "rawMarkdown": "i just added a new folder convert1 to https://www.kaggle.com/datasets/hengck23/asl-demo \nit shows how to use keras input model.",
          "votes": 1,
          "replies": [
            {
              "id": 2188822,
              "postDate": "2023-03-20T02:50:25.633Z",
              "content": "<p>Thank you so much, I'll try these ! </p>",
              "rawMarkdown": "Thank you so much, I'll try these ! "
            },
            {
              "id": 2189403,
              "postDate": "2023-03-20T13:19:55.457Z",
              "rawMarkdown": "",
              "votes": 1,
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 2188724,
      "postDate": "2023-03-19T23:16:49.727Z",
      "content": "<p>thinking of implementing a differentiable topk feature selector … </p>",
      "rawMarkdown": "thinking of implementing a differentiable topk feature selector ... "
    },
    {
      "id": 2187997,
      "postDate": "2023-03-19T08:22:03.700Z",
      "content": "<p>Thanks for your share！！ can I ask what do you mean of \"use triu\"</p>",
      "rawMarkdown": "Thanks for your share！！ can I ask what do you mean of \"use triu\"",
      "replies": [
        {
          "id": 2188014,
          "postDate": "2023-03-19T08:29:31.507Z",
          "content": "<p>upper triangle maxtrix index (you only need half of the pairwise distance matrix)</p>",
          "rawMarkdown": "upper triangle maxtrix index (you only need half of the pairwise distance matrix)",
          "votes": 1,
          "replies": [
            {
              "id": 2188043,
              "postDate": "2023-03-19T08:50:39.050Z",
              "content": "<p>Got it, thanks a lot!   I noticed that point_dim down to 960 from 1002 when apply \"fixed triu\", what's the meaning of \"fixed triu\"👀</p>",
              "rawMarkdown": "Got it, thanks a lot!   I noticed that point_dim down to 960 from 1002 when apply \"fixed triu\", what's the meaning of \"fixed triu\"👀"
            }
          ]
        }
      ]
    },
    {
      "id": 2183926,
      "postDate": "2023-03-16T03:29:44.980Z",
      "content": "<blockquote>\n  <p>one (or two) layer transformer is sufficient</p>\n</blockquote>\n<p>hmmm I have different results to this, deeper transformer lead to a higher CV:</p>\n<p>1 layer   0.50<br>\n2 layers 0.53<br>\n3 layers 0.55<br>\n4 layers 0.56</p>",
      "rawMarkdown": "> one (or two) layer transformer is sufficient\n\nhmmm I have different results to this, deeper transformer lead to a higher CV:\n\n1 layer   0.50\n2 layers 0.53\n3 layers 0.55\n4 layers 0.56",
      "replies": [
        {
          "id": 2183941,
          "postDate": "2023-03-16T03:51:01.303Z",
          "content": "<p>performanced is limited by data. model configuration and feature extraction are just means to ahieve it.</p>\n<p>there is better feature represenation with less transfomer layer</p>",
          "rawMarkdown": "performanced is limited by data. model configuration and feature extraction are just means to ahieve it.\n\nthere is better feature represenation with less transfomer layer\n"
        }
      ]
    },
    {
      "id": 2183228,
      "postDate": "2023-03-15T14:37:10.477Z",
      "content": "<p>Hi~ You said that <code>one (or two) layer transformer is sufficient</code>, is that means we only need 1 layer of MultiHeadAttention? In the paper \"Attention is all you need\", a complete transformer contains 6 duplicated encoders and decoders, and each encoder contains one MultiHeadAttention Layer and something else. I learned your code <a href=\"https://www.kaggle.com/code/hengck23/lb-0-67-one-pytorch-transformer-solution\" target=\"_blank\">lb-0-67-one-pytorch-transformer-solution</a>, and I think you mean we just need one layer of MultiHeadAttention but not just one layer of transformer?</p>\n<p>Best wishes<br>\nhappyOldMan</p>",
      "rawMarkdown": "Hi~ You said that `one (or two) layer transformer is sufficient`, is that means we only need 1 layer of MultiHeadAttention? In the paper \"Attention is all you need\", a complete transformer contains 6 duplicated encoders and decoders, and each encoder contains one MultiHeadAttention Layer and something else. I learned your code [lb-0-67-one-pytorch-transformer-solution](https://www.kaggle.com/code/hengck23/lb-0-67-one-pytorch-transformer-solution), and I think you mean we just need one layer of MultiHeadAttention but not just one layer of transformer?\n\nBest wishes\nhappyOldMan",
      "replies": [
        {
          "id": 2183236,
          "postDate": "2023-03-15T14:41:51.207Z",
          "content": "<p>one or two layer of MultiHeadAttention.<br>\nit also depends on the strength of your emebedding.</p>\n<p>it may work well if you have gru or/and  MultiHeadAttention</p>\n<hr>\n<p>anyway, the key is really how you normalize and extract/represent your feature from the 2d/3d points.</p>",
          "rawMarkdown": "one or two layer of MultiHeadAttention.\nit also depends on the strength of your emebedding.\n\nit may work well if you have gru or/and  MultiHeadAttention\n\n---\n\nanyway, the key is really how you normalize and extract/represent your feature from the 2d/3d points.",
          "replies": [
            {
              "id": 2183317,
              "postDate": "2023-03-15T15:26:18.740Z",
              "content": "<p>Got it, thanks!</p>",
              "rawMarkdown": "Got it, thanks!"
            }
          ]
        }
      ]
    },
    {
      "id": 2182925,
      "postDate": "2023-03-15T11:41:14.323Z",
      "content": "<p>[team forming advertisement]</p>\n<p>Looking for experienced Kagglers to form a team. We are interested in working with 8-bit quantization aware training with Keras or pytorch and implementation in tflite. </p>\n<p>It's important to note that our main objective is to learn from experts in deep learning technology, specifically embedded implementation. While we haven't yet decided if we will use an int8 model for our submission, as we are aware that only TFLite supports 8-bit computation for ARM CPU.</p>\n<p>We're committed to achieving the highest ranking on LB, but we also want to emphasize that Kaggle is about having fun and enjoying the competition.</p>\n<p>If this sounds like a team you'd like to be a part of, please submit your application and state your experiences in 8-bit quantization.<br>\nPlease note that only shortlisted candidates will be contacted.</p>\n<p>Thank you!</p>\n<p>(drafted by chatgpt)<br>\nlast update 15-mar</p>",
      "rawMarkdown": "[team forming advertisement]\n\nLooking for experienced Kagglers to form a team. We are interested in working with 8-bit quantization aware training with Keras or pytorch and implementation in tflite. \n\nIt's important to note that our main objective is to learn from experts in deep learning technology, specifically embedded implementation. While we haven't yet decided if we will use an int8 model for our submission, as we are aware that only TFLite supports 8-bit computation for ARM CPU.\n\nWe're committed to achieving the highest ranking on LB, but we also want to emphasize that Kaggle is about having fun and enjoying the competition.\n\nIf this sounds like a team you'd like to be a part of, please submit your application and state your experiences in 8-bit quantization.\nPlease note that only shortlisted candidates will be contacted.\n\nThank you!\n\n(drafted by chatgpt)\nlast update 15-mar",
      "replies": [
        {
          "id": 2207158,
          "postDate": "2023-04-03T06:34:50.307Z",
          "content": "<p>last update 03-apr</p>\n<p>the key to winning is to create more data.  Here, wre inviting external data creator to join our team.<br>\nTo apply, you should show your work for:</p>\n<pre><code>1) use mediapipe to create 543 landmarks on some open source public dataset. Show some samples using public dataset or notebook.\n2) apply some model (your model or public model) and the difference for top-1 accuracy : kaggle dataset - your dataset &lt; 15%\n</code></pre>\n<p>Please note that only shortlisted candidates will be contacted.</p>\n<p>Thank you!</p>",
          "rawMarkdown": "last update 03-apr\n\nthe key to winning is to create more data.  Here, wre inviting external data creator to join our team.\nTo apply, you should show your work for:\n\n```\n1) use mediapipe to create 543 landmarks on some open source public dataset. Show some samples using public dataset or notebook.\n2) apply some model (your model or public model) and the difference for top-1 accuracy : kaggle dataset - your dataset < 15%\n```\nPlease note that only shortlisted candidates will be contacted.\n\nThank you!"
        }
      ]
    },
    {
      "id": 2182591,
      "postDate": "2023-03-15T07:36:04.037Z",
      "content": "<p>Thanks for sharing your experiment result, I’m curious about the val acc you got. I’ve seen many got a pretty high val acc. (around 0.75), your top1 acc is lower that that, what may be the reason？<br>\nIn my transformer model, I tried many different modifications to the transformer block and it’s hard to break 0.7 on val acc, I got my highest acc at 0.69.</p>",
      "rawMarkdown": "Thanks for sharing your experiment result, I’m curious about the val acc you got. I’ve seen many got a pretty high val acc. (around 0.75), your top1 acc is lower that that, what may be the reason？\nIn my transformer model, I tried many different modifications to the transformer block and it’s hard to break 0.7 on val acc, I got my highest acc at 0.69.",
      "replies": [
        {
          "id": 2183260,
          "postDate": "2023-03-15T14:58:31.307Z",
          "content": "<p>there are two validation set i use:<br>\n[1] one is just random split (monitor the point when validation loss = training loss, and top-1 accuracy)<br>\n[2] one is straification based on participant id</p>\n<p>cv for [1] can be up to 0.79 (after SWA averaging) / LB 0.69? <br>\ncv for [2] can be up to 0.66 (after SWA averaging) / LB 0.67</p>",
          "rawMarkdown": "there are two validation set i use:\n[1] one is just random split (monitor the point when validation loss = training loss, and top-1 accuracy)\n[2] one is straification based on participant id\n\ncv for [1] can be up to 0.79 (after SWA averaging) / LB 0.69? \ncv for [2] can be up to 0.66 (after SWA averaging) / LB 0.67",
          "votes": 1
        }
      ]
    },
    {
      "id": 2167655,
      "postDate": "2023-03-03T16:25:25.970Z",
      "content": "<blockquote>\n  <p>time_taken = 35 sec (gpu, btach=32)</p>\n</blockquote>\n<p>May I ask if this is the time of one epoch? Or is it a time of 100 epoch? I'm training on tensorflow, but it feels slow.</p>",
      "rawMarkdown": ">time_taken = 35 sec (gpu, btach=32)\n\nMay I ask if this is the time of one epoch? Or is it a time of 100 epoch? I'm training on tensorflow, but it feels slow.",
      "replies": [
        {
          "id": 2168595,
          "postDate": "2023-03-04T11:07:50.940Z",
          "content": "<p>total validation time for 20_000 videos</p>",
          "rawMarkdown": "total validation time for 20_000 videos"
        }
      ]
    },
    {
      "id": 2187722,
      "postDate": "2023-03-19T01:56:45.423Z",
      "rawMarkdown": "",
      "votes": -5,
      "isDeleted": true,
      "replies": [
        {
          "id": 2187785,
          "postDate": "2023-03-19T03:45:59.167Z",
          "rawMarkdown": "",
          "votes": 3,
          "isDeleted": true
        }
      ]
    },
    {
      "id": 2179510,
      "postDate": "2023-03-13T08:14:48.807Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2206052,
      "postDate": "2023-04-02T08:05:45.377Z",
      "content": "<p>Thanks for sharing!!</p>",
      "rawMarkdown": "Thanks for sharing!!"
    }
  ],
  "comments": [
    {
      "id": 2179326,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-03-13T04:40:49.920000",
      "content": "<p>i have been playing with controlnet and i come across this:<br>\n<img src=\"https://i.ibb.co/0cdwBSL/Selection-999-1366.png\" alt=\"https://i.ibb.co/0cdwBSL/Selection-999-1366.png\"></p>\n<p>maybe can make some fun animation video while waiting for training</p>",
      "votes": 14,
      "replies": [
        {
          "id": 2180654,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2023-03-14T02:26:27.493000",
          "content": "<p>landmark --&gt; controlnet --&gt;stable diffusion --&gt; image+style+augment --&gt; media pose --&gt;landmark+augment ???</p>\n<p>note that you can have depthmap as control etc and you can change viewpoint, etc</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2202379,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2023-03-30T02:22:09.557000",
          "content": "<p>controlnet based on face landmark<br>\n<a href=\"https://huggingface.co/georgefen/Face-Landmark-ControlNet\" target=\"_blank\">https://huggingface.co/georgefen/Face-Landmark-ControlNet</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2483568,
          "author_name": "Das Koushik",
          "author_url": "",
          "post_date": "2023-10-15T19:40:02.877000",
          "content": "<p>Hi, that is really interesting. I am guessing this will create images frame by frame and we merge all the frames to create a video. Can you give me gist of it? I'm still learning, but seems like a fun learning project. Thank you!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2201909,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-03-29T16:19:45.363000",
      "content": "<p>flip augmentation magic:<br>\nCV 0.703 (single fold)<br>\nLB 0.72</p>\n<pre><code>LHAND = np.arange(468, 489).tolist() # 21\nRHAND = np.arange(522, 543).tolist() # 21\nPOSE  = np.arange(489, 522).tolist() # 33\nFACE  = np.arange(0,468).tolist()    #468\n\nREYE = [\n    33, 7, 163, 144, 145, 153, 154, 155, 133,\n    246, 161, 160, 159, 158, 157, 173,\n]\nLEYE = [\n    263, 249, 390, 373, 374, 380, 381, 382, 362,\n    466, 388, 387, 386, 385, 384, 398,\n]\nNOSE=[\n    1,2,98,327\n]\nSLIP = [\n    78, 95, 88, 178, 87, 14, 317, 402, 318, 324, 308,\n    191, 80, 81, 82, 13, 312, 311, 310, 415,\n]\nSPOSE = (np.array([\n    11,13,15,12,14,16,23,24,\n])+489).tolist()\n\n\n...\n## assume zero mean xyz. so flip can be implement by multuiplication of -1\ndef do_hflip_hand(lhand, rhand):\n    rhand[...,0] *= -1\n    lhand[...,0] *= -1\n    rhand, lhand = lhand,rhand\n    return lhand, rhand\n\ndef do_hflip_spose(spose):\n    spose[...,0] *= -1\n    spose = spose[:,[3,4,5,0,1,2,7,6]]\n    return spose\n\ndef do_hflip_slip(slip):\n    slip[...,0] *= -1\n    slip = slip[:,[10,9,8,7,6,5,4,3,2,1,0]+[19,18,17,16,15,14,13,12,11]]\n    return slip\n\n...\n...\n        xyz = load_relevant_data_subset(pq_file)\n        lhand = xyz[:,LHAND]\n        rhand = xyz[:,RHAND]\n        spose = xyz[:,SPOSE]\n        leye = xyz[:,LEYE]\n        reye = xyz[:,REYE]\n        slip = xyz[:,SLIP]\n        nose = xyz[:,NOSE]\n\n\n    if is_aug==1:\n        if np.random.rand()&lt;0.5:\n            lhand, rhand = do_hflip_hand(lhand, rhand)\n            spose = do_hflip_spose(spose)\n            leye, reye = do_hflip_eye(leye, reye)\n            slip = do_hflip_slip(slip)\n            nose = do_hflip_nose(nose)\n</code></pre>\n<p>train log:</p>\n<pre><code>fold_type = kaggle-part\nfold = 0\ntrain_dataset : \n    len = 71518\n    num_participant_id = 16\n\nvalid_dataset : \n    len = 22959\n    num_participant_id = 5\n...\n   batch_size = 64 \n   experiment = ['tx025-no-z-flip', 'run_train_fold0.py']\n                           |---------------- VALID---------|---- TRAIN/BATCH ----------------------\nrate      iter       epoch | loss   top1   top2    top5    | loss                 | time           \n---------------------------------------------------------------------------------------------------\n...\n1.00e-4   00050265*  45.00 | 2.025  0.698  0.7992  0.870   | 3.887  0.000  0.000  |  1 hr 48 min\n1.00e-4   00053616*  48.00 | 2.042  0.697  0.7969  0.870   | 3.879  0.000  0.000  |  1 hr 56 min\n</code></pre>",
      "votes": 8,
      "replies": [
        {
          "id": 2202177,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2023-03-29T19:49:19.067000",
          "content": "<p>one may want to try \"cut and paste augment\": assemble body parts from different signers</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2202195,
              "author_name": "Andrij",
              "author_url": "",
              "post_date": "2023-03-29T20:06:22.660000",
              "content": "<p>I've already tried something similar when I changed body parts. On my pipeline, the result is controversial. On the one hand, this gave an increase of 1 fold, but on the other hand, the ensemble did not improve. Of course I could have made a mistake somewhere in the code. And this reduced val_loss</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2202204,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-03-29T20:12:24.370000",
              "content": "<p>the parts have to be synchronised correctly in time. a better way is to use gan or pddm</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2202381,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-03-30T02:30:20.223000",
              "content": "<p>instead of cut and paste, imagine:</p>\n<ol>\n<li>there is a canoical signer</li>\n<li>you are use distance constrain and inverse kinemetics to transfer kaggle signer motion to this canoical signer</li>\n<li>check if we use only motion of a sinle canoical signer, classification is better</li>\n<li>think of a way to predict :  canoical signer motion given  transfer kaggle signer motion<br>\n(e.g. pre-processing or learn a simple net)</li>\n</ol>\n<p>basically, we are learning a better normalisation<br>\n<a href=\"https://www.youtube.com/watch?v=U9G2VlG9HbA\" target=\"_blank\">https://www.youtube.com/watch?v=U9G2VlG9HbA</a></p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2202388,
          "author_name": "Carno Zhao",
          "author_url": "",
          "post_date": "2023-03-30T02:48:43.213000",
          "content": "<p>🤔Do other landmarks really help? In my experiments, event simply add pose landmarks (arms position) decrease CV.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2202395,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-03-30T03:08:03.123000",
              "content": "<p>it depends on how you normalised i think.</p>\n<p>there is many noise in the data becuase each signer is of different size and their camera may have different frame rate. They also move differently.</p>\n<p>in theory:<br>\neye:  should help. for sign words  \"awake\", the eye is closed then opened.<br>\npose (arm): should help. this is global motion of hand.</p>\n<p>but if your nomralisation is poor or your model cannot filter out the noise, then there will be easy overfitting to noise</p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 2202398,
              "author_name": "Carno Zhao",
              "author_url": "",
              "post_date": "2023-03-30T03:12:11.127000",
              "content": "<p>yes, for me, normalization method also matters</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2202578,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2023-03-30T06:40:03.967000",
          "content": "<p>this is not too difficult if you think about it</p>\n<p><img src=\"https://i.ibb.co/5FLVDDw/Selection-999-1630.png\" alt=\"https://i.ibb.co/5FLVDDw/Selection-999-1630.png\"></p>",
          "votes": 1,
          "replies": [
            {
              "id": 2202731,
              "author_name": "hoyso48",
              "author_url": "",
              "post_date": "2023-03-30T09:06:33.303000",
              "content": "<p>I tried a very similar approach to this, but unfortunately, I didn't see any performance improvement compared to simple augmentation(maybe my approach was too naive). However, I think it's a fascinating idea, and I'm excited to hear about your progress if you decide to give it a go:).</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2229296,
          "author_name": "Araik Tamazian",
          "author_url": "",
          "post_date": "2023-04-21T08:27:47.767000",
          "content": "<p>According to LANDMARKS_REFINEMENT_LEFT_EYE_CONFIG and LANDMARKS_REFINEMENT_RIGHT_EYE_CONFIG (from <a href=\"https://github.com/tensorflow/tfjs-models/blob/master/face-landmarks-detection/src/tfjs/constants.ts\" target=\"_blank\">https://github.com/tensorflow/tfjs-models/blob/master/face-landmarks-detection/src/tfjs/constants.ts</a>) , your REYE if a left eye, and LEYE is a right eye. Are you sure that's correct?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2193356,
      "author_name": "KKY",
      "author_url": "",
      "post_date": "2023-03-23T08:28:19.040000",
      "content": "<p>I love your experiments records table! thanks for sharing this!</p>",
      "votes": 8,
      "replies": [
        {
          "id": 2193995,
          "author_name": "Dheeraj Pandey",
          "author_url": "",
          "post_date": "2023-03-23T16:50:13.230000",
          "content": "<p>+1 <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2178946,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-03-12T19:46:44.203000",
      "content": "<p>[unorganized] some proven intermediate results … will write in more detail later</p>\n<ul>\n<li><p>add motion (xyz at current frame - xyz at previous frame). i.e. each point now has 6 dim. need to handle NaN motion correctly<br>\nimprove CV by 0.015</p></li>\n<li><p>add framerate interpolation augmentation scipy.interpolate.interp1d()<br>\nimprove CV by 0.015</p></li>\n</ul>",
      "votes": 7,
      "replies": [
        {
          "id": 2180147,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2023-03-13T16:08:00.600000",
          "content": "<p>ideally, if there is lots and lots of training samples, you don't have to do anything, just add more parameters and the model would learn scale, rotation invariant features.</p>\n<p>if there is not much training samples, you can do augmentation, e.g. rotate/shift the hands</p>\n<p>you can also handcraft features, e.g. pairwise distance, pairwise angles  of the hand points (these are invariant to rotation)</p>\n<p>better still,  graph network CGN  or pointnet features are for this.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2180165,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-03-13T16:21:10.243000",
              "content": "<p>convergence is magical if you do things right !</p>\n<p>train/validation participant id does not overlap shown below:<br>\n<img src=\"https://i.ibb.co/g3C1Gq6/Selection-999-1379.png\" alt=\"https://i.ibb.co/g3C1Gq6/Selection-999-1379.png\"></p>\n<p>distance matrix is the poor man's graph network D x A x inv(D) …<br>\ni need to google for better joint kinematics features ….</p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 2180629,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-03-14T01:49:37.657000",
              "content": "<p>i wondered if i renderd the point into 48x48 images and use 3dCNN, what will be the speed and accuracy …</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2181090,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-03-14T10:11:12.587000",
              "content": "<p>possible of aux loss:</p>\n<ul>\n<li>predict if word is one-hand or two-hand word</li>\n<li>predict if the signer is using left, right or both hands</li>\n<li>predict signer identity</li>\n</ul>\n<p>….</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2182554,
              "author_name": "Lau2664",
              "author_url": "",
              "post_date": "2023-03-15T06:51:09.113000",
              "content": "<p>I didn’t try 3D CNN yet, but I have tried 2D CNN. I use a (batch_size, num_frame, num_landmark, 3) tensor as input. I didn’t do much experiments, the best val acc I got is around 0.53. I didn’t record the speed. There are many skeleton-based motion recognition papers that use CNN, but according to the paper, they all achieve a fantastic result, which is much higher than 0.53… I guess we need to process the data in some ways to make it suitable to CNN.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2182562,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-03-15T06:59:03.907000",
              "content": "<p><a href=\"https://ibb.co/85Qp2qv\"><img src=\"https://i.ibb.co/1q45K1y/Selection-999-1426.png\" alt=\"Selection-999-1426\"></a><br>\n<a href=\"https://ibb.co/RQ2vLxC\"><img src=\"https://i.ibb.co/SVNsqhK/Selection-999-1422.png\" alt=\"Selection-999-1422\"></a></p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2182671,
              "author_name": "Gregor Lied",
              "author_url": "",
              "post_date": "2023-03-15T08:41:20.820000",
              "content": "<p><a href=\"https://www.kaggle.com/jay2333\" target=\"_blank\">@jay2333</a> Papers in the skeleton-based action recognition domain rely on high-quality pose data. In most cases, they use HRNet pose estimator to extract the skeleton landmarks. Furthermore, there are only a few missing values. </p>\n<p>The landmarks in our dataset are accurate, but just not as good as landmarks from HRNet. Plus, we have a lot of missing values and we need to find a way to deal with them. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2183125,
              "author_name": "Lau2664",
              "author_url": "",
              "post_date": "2023-03-15T13:29:50.920000",
              "content": "<p><a href=\"https://www.kaggle.com/gregorlied\" target=\"_blank\">@gregorlied</a> Totally agree with that. According to the current public notebook, a simple fully connected network can perform well. So the main issue is the data process part.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2183239,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-03-15T14:42:51.547000",
              "content": "<p>\" fully connected network can perform well. \"</p>\n<p>the reason is that the seq are not really long</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2183267,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-03-15T15:03:07.460000",
              "content": "<p>there is a trick to speedup 3d CNN or 2d CNN</p>\n<p>basically it is like CLIP (aligned text and image)</p>\n<p>here you trained joint encoder for (pose, rendered image) pair.<br>\nrendered image has more information so they have better embedding.</p>\n<p>you can think of we are distilling image embedding to point embedding.<br>\nsince both embedding are \"aligned\", in testing we are only using point embedding (though we trained them jointly)</p>\n<hr>\n<p>try to use different color for part, shade for depth when rendering ….</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2185509,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-03-17T05:45:15.100000",
              "content": "<p>yet  another aux loss,</p>\n<ul>\n<li>if word are noun (object) or verb (action, motion), etc</li>\n</ul>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2203259,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-03-30T16:45:01.383000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2224756,
          "author_name": "Vadim Timakin",
          "author_url": "",
          "post_date": "2023-04-17T16:22:28.987000",
          "content": "<p>May I ask you how you actually apply interpolation? I'm trying to apply it, but even with soft parameters the score is lower</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2228313,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-04-20T12:56:37.840000",
      "content": "<p>watch this thread!</p>\n<p>receipe for LB0.76 one fold:</p>\n<hr>\n<p>i seems that there is somthing very strange about the dataset … maybe very noisy.  i have to use unconventional strong regularisation and training method ….</p>\n<p>hint: <br>\n[1] Dropout Reduces Underfitting<br>\n<a href=\"https://arxiv.org/abs/2303.01500\" target=\"_blank\">https://arxiv.org/abs/2303.01500</a></p>\n<p>i find that if i reduce learning rate, the model easily overfits. instead, according to paper[1], by setting dropout to different values at different training epoches, we can reduce validation loss (it has the same effects as reducing learning rate) without overfitting:</p>\n<ul>\n<li>early: train without dropout</li>\n<li>late: train with large dropout</li>\n<li>final: tune with small dropout</li>\n</ul>\n<p><img src=\"https://i.ibb.co/mz15RwQ/Selection-999-1847.png\" alt=\"https://i.ibb.co/mz15RwQ/Selection-999-1847.png\"></p>\n<p><img src=\"https://i.ibb.co/GvkZB5j/Selection-999-1848.png\" alt=\"https://i.ibb.co/GvkZB5j/Selection-999-1848.png\"></p>\n<pre><code>fold_type = kaggle-random (20% validation, 80% train. pure random)\nfold = 0\nvalid_dataset : \n    len = 18896\n    num_participant_id = 21\n    num_external_video = 0\n\n\ntime_taken =  0 min 27 sec\nloss = 1.7932500839233398\ntopk[0] = 0.8607112616426758 (LB 0.76)\ntopk[1] = 0.9180249788314987\ntopk[2] = 0.9374470787468248\ntopk[3] = 0.9457557154953429\ntopk[4] = 0.9513653683319221\ncheckpoint_file = /home/titanx/hengck/share1/kaggle/2022/hand-sign/result/run202/tx90-face-norm-480-0b/fold-0-kaggle-random/checkpoint/swa.model.pth\n</code></pre>",
      "votes": 7,
      "replies": [
        {
          "id": 2228775,
          "author_name": "Daniel Quintero",
          "author_url": "",
          "post_date": "2023-04-20T20:02:19.243000",
          "content": "<p>Have you tried using ViT or are you still using Conv1d + MHA?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2229582,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-04-21T13:40:39.923000",
              "content": "<p>dropout has the effect of flatten the valley of loss landscape.<br>\n(this is the same as SWA, deeply supervised loss,  etc)</p>\n<p>Of course, adeverisal training SAM (sharpness aware optimizer, aka sample perturbation via weight perturbation) also flattens the valley.<br>\nBelow is results of SAM, which is comparable to late dropout</p>\n<pre><code>fold_type = kaggle-part\nfold = 0\nvalid_dataset : \n    len = 22959\n    num_participant_id = 5\n    num_external_video = 0\n\n\n    22959 / 22959   0 min 30 sec\n\ntime_taken =  0 min 30 sec\nloss = 2.5410375595092773\ntopk[0] = 0.7265560346704996\ntopk[1] = 0.8181105448843591\ntopk[2] = 0.856265516790801\ntopk[3] = 0.8760398972080665\ntopk[4] = 0.8882791062328499\ncheckpoint_file = /home/user/share1/kaggle/2022/hand-sign/result/run202/tx90-face-norm-sam-fine-rho0.1/fold-0-kaggle-part/checkpoint/swa.model.pth\n----- end -----\n</code></pre>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2229794,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-04-21T17:30:00.240000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2229796,
              "author_name": "Daniel Quintero",
              "author_url": "",
              "post_date": "2023-04-21T17:30:20.373000",
              "content": "<p>It may be worthwhile to try additional parameters in conjunction with a 'late' configuration</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2163630,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-03-01T00:49:08.950000",
      "content": "<p>[reference]</p>\n<ul>\n<li>'Sign Pose-based Transformer for Word-level Sign Language Recognition' - Matyáš Boháček</li>\n<li>Combining Efficient and Precise Sign Language Recognition: Good pose estimation library is all you need- Matyáš Boháček</li>\n</ul>\n<p><a href=\"https://github.com/matyasbohacek/spoter\" target=\"_blank\">https://github.com/matyasbohacek/spoter</a><br>\n<a href=\"https://www.matyasbohacek.com/\" target=\"_blank\">https://www.matyasbohacek.com/</a><br>\n(uses MediaPipe library pose too)</p>\n<p><a href=\"https://huggingface.co/spaces/matyasbohacek/spoter-demo-test\" target=\"_blank\">https://huggingface.co/spaces/matyasbohacek/spoter-demo-test</a></p>\n<p><a href=\"https://www.youtube.com/watch?v=6i5ODpHMny8\" target=\"_blank\">https://www.youtube.com/watch?v=6i5ODpHMny8</a></p>\n<p><img src=\"https://i.ibb.co/r2NHnkh/Selection-999-1166.png\" alt=\"https://i.ibb.co/r2NHnkh/Selection-999-1166.png\"></p>\n<hr>\n<p>SLGTFORMER: AN ATTENTION-BASED APPROACH TO SIGN LANGUAGE RECOGNITION<br>\n<a href=\"https://github.com/neilsong/SLGTformer\" target=\"_blank\">https://github.com/neilsong/SLGTformer</a></p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 2163638,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-03-01T00:59:53.167000",
      "content": "<p>[external data]<br>\n… to be updated …</p>\n<p>HOW TO SIGN<br>\n<a href=\"https://how2sign.github.io/\" target=\"_blank\">https://how2sign.github.io/</a><br>\n<a href=\"https://openaccess.thecvf.com/content/CVPR2021/papers/Duarte_How2Sign_A_Large-Scale_Multimodal_Dataset_for_Continuous_American_Sign_Language_CVPR_2021_paper.pdf\" target=\"_blank\">https://openaccess.thecvf.com/content/CVPR2021/papers/Duarte_How2Sign_A_Large-Scale_Multimodal_Dataset_for_Continuous_American_Sign_Language_CVPR_2021_paper.pdf</a><br>\n(vocab=16k, time=79 hr, signer=11)</p>\n<p>MS-ASL American Sign Language Dataset<br>\n<a href=\"https://microsoft.github.io/data-for-society/dataset?d=MS-ASL-American-Sign-Language-Dataset\" target=\"_blank\">https://microsoft.github.io/data-for-society/dataset?d=MS-ASL-American-Sign-Language-Dataset</a><br>\n<a href=\"https://arxiv.org/pdf/1812.01053.pdf\" target=\"_blank\">https://arxiv.org/pdf/1812.01053.pdf</a><br>\n(vocab=1000, num video=25513, signer=222)</p>\n<p>BSL-1K: Scaling up co-articulated sign language recognition using mouthing cues<br>\n<a href=\"https://www.robots.ox.ac.uk/~vgg/research/bsl1k/\" target=\"_blank\">https://www.robots.ox.ac.uk/~vgg/research/bsl1k/</a><br>\n(vocab=1000, num video=1000)</p>\n<p>ChaLearn LAP Large Scale Signer Independent Isolated Sign Language<br>\nRecognition Challenge: Design, Results and Future Research</p>\n<p>ASL-Skeleton3D and ASL-Phono: Two Novel<br>\nDatasets for the American Sign Language</p>\n<hr>\n<ul>\n<li><p>how to read media pipline to extra keypoint for external data <br>\n… to be udpated …</p></li>\n<li><p>more about PopSign learning app<br>\n<a href=\"https://devpost.com/software/pop-sign-learning\" target=\"_blank\">https://devpost.com/software/pop-sign-learning</a><br>\n<a href=\"https://www.youtube.com/watch?v=nbQAfiJzp2g&amp;t=430s\" target=\"_blank\">https://www.youtube.com/watch?v=nbQAfiJzp2g&amp;t=430s</a><br>\n<a href=\"https://github.com/Accessible-Technology-in-Sign/PopSign\" target=\"_blank\">https://github.com/Accessible-Technology-in-Sign/PopSign</a></p></li>\n</ul>\n<p><a href=\"https://github.com/Accessible-Technology-in-Sign/ASLRT\" target=\"_blank\">https://github.com/Accessible-Technology-in-Sign/ASLRT</a><br>\n\"A pipeline to run tracking Kinect, MediaPipe and AlphaPose experiments on sign language feature data using either HMMs, SBHMMs, or Transformers.\"</p>\n<p>collect data you from youtube?</p>",
      "votes": 6,
      "replies": [
        {
          "id": 2176272,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2023-03-10T14:03:14.067000",
          "content": "<p>pytorch mediapipe … hmm<br>\n<a href=\"https://github.com/zmurez/MediaPipePyTorch\" target=\"_blank\">https://github.com/zmurez/MediaPipePyTorch</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2183164,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2023-03-15T13:53:17.117000",
          "content": "<p>this image explain why hand are missing.<br>\nyou probably can ignore the legs</p>\n<p><img src=\"https://i.ibb.co/237HvmF/original.png\" alt=\"https://i.ibb.co/237HvmF/original.png\"></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2183363,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2023-03-15T15:52:54.357000",
          "content": "<p>there is a related product called copycat:<br>\nCopyCat: Using Sign Language Recognition to Help Deaf Children Acquire Language Skills<br>\n<a href=\"https://gvu.gatech.edu/research/projects/copycat-using-sign-language-recognition-help-deaf-children-acquire-language-skills#:~:text=CopyCat%20is%20a%20game%20where,language%20skills%20and%20working%20memory\" target=\"_blank\">https://gvu.gatech.edu/research/projects/copycat-using-sign-language-recognition-help-deaf-children-acquire-language-skills#:~:text=CopyCat%20is%20a%20game%20where,language%20skills%20and%20working%20memory</a>.</p>\n<p><a href=\"https://www.youtube.com/watch?v=WMFZVcey8FU\" target=\"_blank\">https://www.youtube.com/watch?v=WMFZVcey8FU</a><br>\n<a href=\"https://github.com/Accessible-Technology-in-Sign/ASLRT\" target=\"_blank\">https://github.com/Accessible-Technology-in-Sign/ASLRT</a></p>\n<p>much information on HMM + handcraft feature here</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2226810,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-04-19T08:43:02.580000",
      "content": "<p>i forget an import detail at the data card. suggestions for augmentation:</p>\n<p>\"While the app provided a video example of the sign desired, signers routinely made variants of the sign based on their background and region.\"  </p>\n<ul>\n<li>you need external data or create your variant</li>\n</ul>\n<p>\"More rarely, signers might fingerspell a sign, miss it completely, or produce the wrong sign.\"  </p>\n<ul>\n<li>need to measure label noise and decide if you want to exclude some data</li>\n</ul>\n<p>\" Extraneous movements, such as scratching an itch, or the ending movement from the previous sign or the onset of the next sign, are sometimes included.\"  </p>\n<ul>\n<li>random add hand (from other video at start and end). actually CAM heat map is about to hightlight the useful frame.</li>\n</ul>\n<p>\" Conversely, some signers pressed the button late or released the button early, causing cropping in some sign examples. \"  </p>\n<ul>\n<li>cut and paste (mix up)</li>\n</ul>\n<p>\"Some signers sign with their left hand; others sign with their right. Some signers switch their signing hand. All of these situations must be handled by the game’s recognition system.\"  </p>\n<ul>\n<li>flip</li>\n</ul>\n<hr>\n<p>a also note that some signer make repeated signs in one video</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2229212,
          "author_name": "Araik Tamazian",
          "author_url": "",
          "post_date": "2023-04-21T06:49:13.603000",
          "content": "<p>\"need to measure label noise and decide if you want to exclude some data\"</p>\n<p>So we need to specify some kind of \"ideal\" signs, and compare train data with it?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2229217,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-04-21T06:51:29.157000",
              "content": "<p>those  that have validation/train error higher than normal  are supsicous outliers  </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2229256,
              "author_name": "Jackson You",
              "author_url": "",
              "post_date": "2023-04-21T07:37:43.810000",
              "content": "<p>If the hidden dataset is very cleanly drawn, I think data with a low frame count can be left out.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2229468,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-04-21T11:49:10.123000",
              "content": "<p>there must be some reason why some video are very long (much longer than average).<br>\nat the signer moving very slowly? or are they not cropping the video (press start and end) correctly? or are they signing wrongly?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2229484,
              "author_name": "Jackson You",
              "author_url": "",
              "post_date": "2023-04-21T12:07:25.227000",
              "content": "<p>You're right, long frame data has some problems. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2229604,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-04-21T14:07:24.983000",
              "content": "<p>this is explain why conv1d does not work but conv1d+max pool works (about LB 0.72, compared to transformer of 0.74)</p>\n<pre><code>        self.conv2 = nn.Sequential(\n            nn.Conv1d(1024, 768, kernel_size=3, padding=1, stride=1),\n            LayerNorm1D(768),\n            nn.ReLU(inplace=True),\n            nn.Dropout(p=0.1),\n            nn.MaxPool1d(kernel_size=3, padding=1, stride=2),\n        )\n</code></pre>\n<p>max pool is poor man's attention pool(softmax versus hard max)</p>\n<p>also, why adding noise (dropout, dropframe, random values, mix frame …) before mha works. it forces the model to select the correct signal.</p>\n<p>i am very surprise of the small num of paramaters  and performance of transformer compared to other model design (like conv1d, dense, etc)</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2224078,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-04-17T01:26:37.653000",
      "content": "<p>it is suprised for me to achieve 0.75 in one fold with participant split train data(num participant = 5 in validation data). This makes me think that 0.76 is possible. </p>\n<pre><code>fold_type = kaggle-part\nfold = 0\nvalid_dataset : \n    len = 22959\n    num_participant_id = 5 \n\n    22959 / 22959   0 min 42 sec\n\ntime_taken =  0 min 42 sec\nloss = 2.406057596206665\ntopk[0] = 0.7277320440785748\ntopk[1] = 0.8219870203406072\ntopk[2] = 0.858530423798946\ntopk[3] = 0.8757785617840498\ntopk[4] = 0.8888888888888888\ncheckpoint_file = /home/titanx/hengck/share1/kaggle/2022/hand-sign/result/run201/tx61a-1-frame-first-scale-2/fold-0-kaggle-part/checkpoint/swa.model.pth\n</code></pre>\n<p>and i hit the jackpot 888888888</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2224095,
          "author_name": "Rob Mulla",
          "author_url": "",
          "post_date": "2023-04-17T02:08:50.303000",
          "content": "<p>888888888!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2225136,
          "author_name": "chemdatafarmer",
          "author_url": "",
          "post_date": "2023-04-17T23:29:35.293000",
          "content": "<p>Are these accuracies shown for the 5 holdout participants?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2206909,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-04-03T02:01:33.027000",
      "content": "<p>a couple on conv1d (as spatial temporal encoder) + MHA attention seems to have good results.<br>\nsome good ideas to design efficient transformer.</p>\n<p>[1] FastViT: A Fast Hybrid Vision Transformer using Structural Reparameterization<br>\n<a href=\"https://arxiv.org/pdf/2303.14189.pdf\" target=\"_blank\">https://arxiv.org/pdf/2303.14189.pdf</a></p>\n<p>[2] Next-ViT: Next Generation Vision Transformer for Efficient Deployment in Realistic Industrial Scenarios<br>\n<a href=\"https://arxiv.org/abs/2207.05501\" target=\"_blank\">https://arxiv.org/abs/2207.05501</a></p>\n<p>to avoid uncessary permute reshape, i can write my q,v,k linear operation in terms of conv1d.<br>\ni will be removing absolute position embedding and use convolutional position embedding instead (convolutional FFN, i.e. position information is implicitly embedding on 1d conv)</p>\n<p>actually it is aslo possible to use 2d convolution, where (W,H) = (x,y,z of spatial, t of temporal).</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2206947,
          "author_name": "Carno Zhao",
          "author_url": "",
          "post_date": "2023-04-03T02:55:46.753000",
          "content": "<p>Won't conv bring much more parameters?</p>\n<p>I suppose what you mean is :</p>\n<pre><code>\ninputs_embeds = inputs_embeds + position_embeds \n\n\ninputs_embeds = inputs_embeds.permute(, , ) \ninputs_embeds = conv1d(inputs_embeds)          \ninputs_embeds = inputs_embeds.permute(, , ) \n</code></pre>\n<p>where the number of parameters in original is TxH, and the conv has HxHx3</p>\n<p>I also tried this idea, but found that the conv is much harder to train than MHSA, maybe because of more parameters.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2207339,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-04-03T10:49:08.177000",
              "content": "<p><a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">@carnozhao</a> </p>\n<p>1d convolution multi-head-attention:</p>\n<pre><code>class MyMultiHeadAttention(nn.Module):\n    def __init__(self,\n            embed_dim,\n            out_dim,\n            qk_dim,\n            v_dim,\n            num_head,\n        ):\n        super().__init__()\n        self.embed_dim = embed_dim\n        self.num_head  = num_head\n        self.qk_dim = qk_dim\n        self.v_dim  = v_dim\n\n        self.q = nn.Conv1d(embed_dim, qk_dim*num_head,1)\n        self.k = nn.Conv1d(embed_dim, qk_dim*num_head,1) # stride=2 for token reduction, kernel&gt;1 for mixing\n        self.v = nn.Conv1d(embed_dim, v_dim*num_head,1)\n\n        self.out = nn.Conv1d(v_dim*num_head, out_dim, 1)\n        self.scale = 1/(qk_dim**0.5)\n\n    #https://github.com/pytorch/pytorch/issues/40497\n    def forward(self, x, x_mask):\n        B,dim,L= x.shape\n\n        num_head = self.num_head\n        qk_dim = self.qk_dim\n        v_dim = self.v_dim\n\n        q = self.q(x) #B,qk_dim,L\n        k = self.k(x)\n        v = self.v(x)\n        q = q.reshape(B, num_head, qk_dim, L).permute(0,1,3,2).contiguous()\n        k = k.reshape(B, num_head, qk_dim, L)#.permute(0,1,2,3).contiguous()\n        v = v.reshape(B, num_head, v_dim,  L).permute(0,1,3,2).contiguous()\n\n        dot = torch.matmul(q, k) *self.scale  # H L L\n        x_mask = x_mask.reshape(B,1,1,L).expand(-1,num_head,L,-1)\n        #dot[x_mask]= -1e4\n        dot.masked_fill_(x_mask, -1e4)\n        attn = F.softmax(dot, -1)    # L L\n\n        v = torch.matmul(attn, v)  # L H dim\n        v = v.permute(0,1,3,2).reshape(B, v_dim*num_head,L).contiguous()\n        out = self.out(v)\n        return out\n\n#---\nembed_dim = 512\nout_dim   = 512\nqk_dim    = 512//4 #for one head\nv_dim     = 512//4\nnum_head  = 4\nmax_length = 96\nbatch_size = 4\n\nmha = MyMultiHeadAttention(\n    embed_dim,\n    out_dim,\n    qk_dim,\n    v_dim,\n    num_head,\n)\nx  = torch.from_numpy( np.random.uniform(-1,1,(batch_size, embed_dim, max_length))).float()\nx_mask  = torch.from_numpy( np.random.uniform(0,1,(batch_size, max_length)))\nx_mask = x_mask&gt;0.5\n\ny=mha(x,x_mask)\nprint(y.shape) #torch.Size([4, 512, 96])\n</code></pre>\n<p>as a comparison:</p>\n<p>original linear multi-head-attention:</p>\n<pre><code>class MyMultiHeadAttention(nn.Module):\n    def __init__(self,\n            embed_dim,\n            out_dim,\n            qk_dim,\n            v_dim,\n            num_head,\n        ):\n        super().__init__()\n        self.embed_dim = embed_dim\n        self.num_head  = num_head\n        self.qk_dim = qk_dim\n        self.v_dim  = v_dim\n\n        self.q = nn.Linear(embed_dim, qk_dim*num_head)\n        self.k = nn.Linear(embed_dim, qk_dim*num_head)\n        self.v = nn.Linear(embed_dim, v_dim*num_head)\n\n        self.out = nn.Linear(v_dim*num_head, out_dim)\n        self.scale = 1/(qk_dim**0.5)\n\n    #https://github.com/pytorch/pytorch/issues/40497\n    def forward(self, x, x_mask):\n        B,L,dim = x.shape\n        #out, _ = self.mha(x,x,x, key_padding_mask=x_mask)\n        num_head = self.num_head\n        qk_dim = self.qk_dim\n        v_dim = self.v_dim\n\n        q = self.q(x)\n        k = self.k(x)\n        v = self.v(x)\n        q = q.reshape(B, L, num_head, qk_dim).permute(0,2,1,3).contiguous()\n        k = k.reshape(B, L, num_head, qk_dim).permute(0,2,3,1).contiguous()\n        v = v.reshape(B, L, num_head, v_dim ).permute(0,2,1,3).contiguous()\n\n        dot = torch.matmul(q, k) *self.scale  # H L L\n        x_mask = x_mask.reshape(B,1,1,L).expand(-1,num_head,L,-1)\n        #dot[x_mask]= -1e4\n        dot.masked_fill_(x_mask, -1e4)\n        attn = F.softmax(dot, -1)    # L L\n\n        v = torch.matmul(attn, v)  # L H dim\n        v = v.permute(0,2,1,3).reshape(B,L, v_dim*num_head).contiguous()\n        out = self.out(v)\n        return out\n</code></pre>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 2208680,
          "author_name": "Maxim B",
          "author_url": "",
          "post_date": "2023-04-04T08:53:19.857000",
          "content": "<p>My conv1d layers are not getting through the pytorch to tflite conversion. Do you still do pytorch -&gt; onnx -&gt;tflite conversion or do you have a tf/keras twin which you convert to tflite?</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2204321,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-03-31T13:31:22.793000",
      "content": "<p>this is how much the face (upper triangle = eyes and lip) and body(lower quadliteral = shoulders and hips)  have moved for one hand sign.</p>\n<p>I wonder if there will be improvement in accuracy if i remove these noise motion?</p>\n<p><img src=\"https://i.ibb.co/xgPNq9L/Selection-999-1647.png\" alt=\"https://i.ibb.co/xgPNq9L/Selection-999-1647.png\"></p>",
      "votes": 3,
      "replies": [
        {
          "id": 2204337,
          "author_name": "A.P.",
          "author_url": "",
          "post_date": "2023-03-31T13:42:42.513000",
          "content": "<p>Tried to apply moving average across frames but it doesn't work for me </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2206111,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-04-02T09:49:09.440000",
      "content": "<p>simplest normalisation code</p>\n<pre><code>REF = [500, 501, 512, 513, 159,  386, 13,]\n\ndef do_normalise_by_ref(xyz, ref):  \n    K = xyz.shape[-1]\n    xyz_flat = ref.reshape(-1,K)\n    m = np.nanmean(xyz_flat,0).reshape(1,1,K)\n    s = np.nanstd(xyz_flat, 0).mean() \n    xyz = xyz - m\n    xyz = xyz / s\n    return xyz\n\nxyz = load_relevant_data_subset(pq_file)[...,:2]\nxyz = do_normalise_by_ref(xyz,xyz[:,REF])\n</code></pre>",
      "votes": 4,
      "replies": [
        {
          "id": 2207859,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-04-03T16:59:36.753000",
          "content": "",
          "votes": 0,
          "replies": [
            {
              "id": 2207929,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-04-03T17:40:00.700000",
              "content": "<pre><code>REF = [500, 501, 512, 513, 159,  386, 13,]\nref = xyz[:,REF]\nxyz_flat = ref.reshape(-1,K)\n</code></pre>\n<p>using ref=xyz will be more accurate<br>\nyou can conduct experiments to find the subset of most stable normalising points</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2208007,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-04-03T18:33:32.503000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2178029,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-03-12T02:36:03.543000",
      "content": "<p>[self-study]</p>\n<ul>\n<li>Sign Language Translation with Transformers<br>\n<a href=\"https://www.youtube.com/watch?v=E5nKeEvoAK0\" target=\"_blank\">https://www.youtube.com/watch?v=E5nKeEvoAK0</a></li>\n</ul>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2163502,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-02-28T21:37:38.583000",
      "content": "<p>[ensemble trick]</p>\n<ul>\n<li>since only one model can be submitted and there is a time limit, we can distilled \"ensemble knowledge\" into one model.</li>\n<li>a better solution is distilled encoder (or multiple encoders) and multiple heads. heads are using different \"view\" of encoded features, final prediction is averaged over the heads</li>\n</ul>\n<p>more on this later</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2163494,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-02-28T21:20:38.053000",
      "content": "<p>[augmentation and tricks]</p>\n<ul>\n<li>simulate different frame length </li>\n<li>simulate different NaN (as in train data)</li>\n<li>simulate different missing frame or differnt frame rate? </li>\n<li>3d guassian noise</li>\n<li>type of landmark. use ['face', 'left_hand', 'pose', 'right_hand'] for normalisation and augmentation<br>\n(e.g. affine transform to group of point --&gt; bigger mouth, longer arm …, rotate hand)</li>\n</ul>\n<p>since landmark is xyz, we are having a 3d motion capture. would be fun to use GAN to create more data. distangle motion and  shape</p>\n<p>possible to synthetic animate motion for new person via motion transfer. But we need to find some media pipe human xyz models. </p>",
      "votes": 3,
      "replies": [
        {
          "id": 2178001,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2023-03-12T02:16:37.907000",
          "content": "<p>augmentation paper:<br>\n[1] PoseAug: A Differentiable Pose Augmentation Framework for 3D Human Pose Estimation<br>\n<a href=\"https://arxiv.org/pdf/2105.02465.pdf\" target=\"_blank\">https://arxiv.org/pdf/2105.02465.pdf</a></p>\n<p>[2] DH-AUG: DH Forward Kinematics Model Driven Augmentation for 3D Human Pose Estimation<br>\n<a href=\"https://arxiv.org/pdf/2207.09303.pdf\" target=\"_blank\">https://arxiv.org/pdf/2207.09303.pdf</a></p>",
          "votes": 2,
          "replies": [
            {
              "id": 2178020,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-03-12T02:29:06.657000",
              "content": "<p>i check github for ASL generator<br>\n\"3D Avatar ASL github \"</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2178774,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-03-12T17:46:15.033000",
              "content": "<p>i wonder if motion transfer would work?<br>\ni.e. transfer from one signer to another</p>\n<p><a href=\"https://deepmotionediting.github.io/style_transfer\" target=\"_blank\">https://deepmotionediting.github.io/style_transfer</a></p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2170203,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-03-05T18:52:03.680000",
      "content": "<p>[normalization]</p>\n<p>normalization is all you need!</p>\n<pre><code>def pre_process(xyz):\n    lip = xyz[:, LIP]\n    lhand = xyz[:, LHAND]\n    rhand = xyz[:, RHAND]\n    xyz = torch.cat([ #(none, 82, 3)\n        lip,\n        lhand,\n        rhand,\n    ],1)\n    xyz = xyz - xyz[~torch.isnan(xyz)].mean(0,keepdims=True) #noramlisation to common maen\n    xyz = xyz / xyz[~torch.isnan(xyz)].std(0, keepdims=True)\n    xyz[torch.isnan(xyz)] = 0\n    xyz = xyz[:max_length]\n    return xyz\n</code></pre>",
      "votes": 4,
      "replies": [
        {
          "id": 2170328,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2023-03-05T21:44:12.203000",
          "content": "<pre><code>** start training here! **\n   batch_size = 64 \n   experiment = ['transformer-pool-03-hand_lip', 'run_train_fold0.py']\n                           |---------------- VALID---------|---- TRAIN/BATCH ----------------------\nrate      iter       epoch | loss   top1   top2    top5    | loss                 | time           \n---------------------------------------------------------------------------------------------------\n0.00e+0   00000000*   0.00 | 5.544  0.005  0.0078  0.019   | 0.000  0.000  0.000  |  0 hr 00 min\n1.00e-4   00003600*   3.00 | 3.183  0.293  0.4145  0.572   | 2.815  0.000  0.000  |  0 hr 04 min\n1.00e-4   00007200*   6.00 | 2.637  0.406  0.5310  0.678   | 2.038  0.000  0.000  |  0 hr 08 min\n1.00e-4   00010800*   9.00 | 2.404  0.451  0.5784  0.719   | 1.671  0.000  0.000  |  0 hr 11 min\n1.00e-4   00014400*  12.00 | 2.363  0.475  0.6005  0.729   | 1.540  0.000  0.000  |  0 hr 15 min\n1.00e-4   00018000*  15.00 | 2.198  0.505  0.6321  0.758   | 1.296  0.000  0.000  |  0 hr 19 min\n1.00e-4   00021600*  18.00 | 2.146  0.519  0.6440  0.765   | 1.225  0.000  0.000  |  0 hr 23 min\n1.00e-4   00025200*  21.00 | 2.141  0.524  0.6510  0.765   | 1.090  0.000  0.000  |  0 hr 27 min\n1.00e-4   00028800*  24.00 | 2.101  0.529  0.6570  0.775   | 0.985  0.000  0.000  |  0 hr 31 min\n1.00e-4   00032400*  27.00 | 2.071  0.540  0.6621  0.777   | 0.934  0.000  0.000  |  0 hr 35 min\n1.00e-4   00036000*  30.00 | 2.131  0.538  0.6573  0.772   | 0.911  0.000  0.000  |  0 hr 39 min\n1.00e-4   00039600*  33.00 | 2.070  0.544  0.6674  0.778   | 0.823  0.000  0.000  |  0 hr 43 min\n1.00e-4   00043200*  36.00 | 2.104  0.543  0.6666  0.779   | 0.736  0.000  0.000  |  0 hr 47 min\n1.00e-4   00046800*  39.00 | 2.114  0.544  0.6701  0.779   | 0.738  0.000  0.000  |  0 hr 50 min\n</code></pre>\n<p>fast convergence after normalisation for one layer transformer encoder with cls pooling<br>\nvalidation split is by participant id, i.e  participant id does not overlap in test and train</p>",
          "votes": 5,
          "replies": [
            {
              "id": 2172332,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-03-07T13:08:18.303000",
              "content": "<p>learning based shape normalisation:</p>\n<ul>\n<li><p>Spatial Transformer Networks<br>\n<a href=\"https://arxiv.org/pdf/1506.02025.pdf\" target=\"_blank\">https://arxiv.org/pdf/1506.02025.pdf</a></p></li>\n<li><p>PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation<br>\n<a href=\"https://openaccess.thecvf.com/content_cvpr_2017/papers/Qi_PointNet_Deep_Learning_CVPR_2017_paper.pdf\" target=\"_blank\">https://openaccess.thecvf.com/content_cvpr_2017/papers/Qi_PointNet_Deep_Learning_CVPR_2017_paper.pdf</a></p></li>\n</ul>\n<p>\"We predict an affine transformation matrix by a mini-network (T-net in Fig 2) and directly apply this transformation to the coordinates of input points\", This is basically a 3d spatial transformer network (not the modern seq-to-seq transformer)</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2174824,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-03-09T12:24:09.203000",
              "content": "<p>effects of normalization<br>\n<img src=\"https://i.ibb.co/bLYJG27/Selection-999-1295.png\" alt=\"https://i.ibb.co/bLYJG27/Selection-999-1295.png\"></p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2176110,
              "author_name": "Darien Schettler",
              "author_url": "",
              "post_date": "2023-03-10T12:08:13.087000",
              "content": "<p>I'm fascinated by this. </p>\n<ul>\n<li>Did you use the methods you linked to do the normalization? </li>\n<li>It looks like there is an upward swinging arm (right side of left image) in the original but this looks to be shifted into the mass in the right image. Do you think that's the correct transformation?</li>\n<li>How did you visualize this? (X,Y) coordinates scatter plotted clearly, but I like the soft faded way it looks. Any tips?</li>\n</ul>\n<p>As an aside, what are your thoughts about leveraging angles between keypoint-keypoint vectors within a feature set (primarily hands and pose maybe -- face wouldn't work) instead of the points themselves? I believe, although I may be wrong, that this might make the feature set more agnostic to drift, scale, etc.</p>\n<p>I plan to try a simple model shortly where I replace the hand features with the representative keypoint-keypoint vector angles.</p>\n<hr>\n<p>Anyway, definitely following, thanks for sharing !</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2176238,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-03-10T13:27:31.560000",
              "content": "<p>it is just simple mean and std normalization.<br>\nthere is probably not enough resource (100 msec) to do complicated normalisation</p>\n<pre><code>    fig, ax = plt.subplots()\n    ax.invert_yaxis()\n\n    kaggle_df = pd.read_csv(f'{root_dir}/data/asl-signs/train.ver01.csv')\n    for t,d in kaggle_df.iterrows():\n        print('\\r', f'{t}/{len(kaggle_df)}', end='', flush=True)\n        pq_file = f'{root_dir}/data/asl-signs/{d.path}'\n        xyz = load_relevant_data_subset(pq_file)\n\n\n        pose = xyz[:,POSE]\n        pose = pose.reshape(-1,3)\n\n        non_nan = np.all(~np.isnan(pose), axis=-1)\n        pose = pose[non_nan]\n\n        pose = pose-pose.mean(0,keepdims=True) \n        pose = pose/pose.std(0,keepdims=True)\n\n        ax.scatter(pose[:,0], pose[:,1], alpha=0.01)\n\n        #plt.show()\n        plt.waitforbuttonpress()\n</code></pre>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2176319,
              "author_name": "Darien Schettler",
              "author_url": "",
              "post_date": "2023-03-10T14:56:49.513000",
              "content": "<p>Awesome. Makes sense!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2177309,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-03-11T11:10:47.910000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2179776,
              "author_name": "Vladimir Bogachev",
              "author_url": "",
              "post_date": "2023-03-13T11:55:35.670000",
              "content": "<p>Hi, I'm having trouble reproducing your results. I'm using one-layer transformer with participants 26734, 28656, 16069 as validation, and my top1 val caps at 0.35 using the following hyperparams:</p>\n<ul>\n<li>batch_size 64</li>\n<li>optimizer RAdam+Lookahead (k=5)</li>\n<li>landmarks lips, left_hand, right_hand</li>\n<li>constant learning rate 1e-4 </li>\n<li>embed_dim 1024</li>\n<li>num_heads 8</li>\n<li>cls dropout 0.4</li>\n<li>max_len 512</li>\n<li>cls pooling</li>\n<li>40 epoch<br>\nDo you have some idea why this doesn't work?</li>\n</ul>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2179965,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-03-13T13:56:05.057000",
              "content": "<p><a href=\"https://www.kaggle.com/vbogach\" target=\"_blank\">@vbogach</a> </p>\n<p>most likely there is some bug in your training code</p>\n<p>the training loss is shown at:<br>\n<a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/391265#2170328\" target=\"_blank\">https://www.kaggle.com/competitions/asl-signs/discussion/391265#2170328</a></p>\n<p>possible bug ares:</p>\n<ul>\n<li>forget to normalize shape</li>\n<li>NaN values are not handled</li>\n<li>forget to set model .eval() and .train()</li>\n<li>use probability to ce loss function (instead of logit)</li>\n<li>wrong loss function</li>\n<li>wrong augmentation </li>\n</ul>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2180400,
              "author_name": "Vladimir Bogachev",
              "author_url": "",
              "post_date": "2023-03-13T19:20:24.993000",
              "content": "<p>It looks like I was normalizing each landmark separately and that's why it didn't work. Thanks for your help!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2240056,
      "author_name": "Bob_Fan91",
      "author_url": "",
      "post_date": "2023-04-30T07:22:23.977000",
      "content": "<p>Excellent sharing！I learned a lot from this.Thank you very much!Wish you can get the top prize!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2224702,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-04-17T15:03:30.517000",
      "content": "<p><a href=\"https://github.com/facebookresearch/dropout\" target=\"_blank\">https://github.com/facebookresearch/dropout</a><br>\nDropout Reduces Underfitting</p>\n<p><img src=\"https://user-images.githubusercontent.com/8370623/222586143-3500fa5b-c294-48c9-a5cf-5fac2659e519.png\" alt=\"https://user-images.githubusercontent.com/8370623/222586143-3500fa5b-c294-48c9-a5cf-5fac2659e519.png\"></p>\n<p>late dropout works for me</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2224196,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-04-17T04:45:42.443000",
      "content": "<p>replace the word \"academics\" with \"kaggle\"</p>\n<p><img src=\"https://i.ibb.co/rxL6Thm/Selection-999-1831.png\" alt=\"https://i.ibb.co/rxL6Thm/Selection-999-1831.png\"></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2210622,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-04-05T14:08:28.510000",
      "content": "<p>i wonder if (inverse) DTW can be used as augmentation?<br>\n<a href=\"https://ai.googleblog.com/2021/01/recognizing-pose-similarity-in-images.html\" target=\"_blank\">https://ai.googleblog.com/2021/01/recognizing-pose-similarity-in-images.html</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2204573,
      "author_name": "mp",
      "author_url": "",
      "post_date": "2023-03-31T17:21:01.103000",
      "content": "<p>Hi, I've noticed that you use <code>xyz = xyz - xyz[~torch.isnan(xyz)].mean(0,keepdim=True)</code> to normalize the input. However, <code>xyz[~torch.isnan(xyz)]</code> (or <code>np.isnan</code>) returns a <em>flattened</em> tensor of size <code>(num_frames * num_landmarks * 3 - num_nans)</code>, which means that calling <code>mean</code> will return a scalar. In other words, what you're really doing is subtracting the mean of everything in the sequence instead of normalizing each landmark <strong>across frames</strong> (passing <code>dim=0</code> and <code>keepdim=True</code>). Am I getting it correctly?</p>\n<p>Nevertheless, I believe the correct mean normalization here is to subtract the mean of all landmark positions <strong>for each frame</strong> by doing something like <code>xyz -= xyz.nanmean(dim=1, keepdim=True)</code>. It might be even better to normalize x, y, and z separately but I haven't figured out how to do it yet.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2204581,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2023-03-31T17:26:31.320000",
          "content": "<p>do a few print statement:</p>\n<pre><code>print(xyz.shape)\nprint((torch.isnan(xyz)).shape)\nprint((xyz[~torch.isnan(xyz)]).shape)\nprint((xyz[~torch.isnan(xyz)].mean(0,keepdim=True)).shape)\n</code></pre>",
          "votes": 0,
          "replies": [
            {
              "id": 2204811,
              "author_name": "mp",
              "author_url": "",
              "post_date": "2023-04-01T01:44:46.860000",
              "content": "<p>Here are the sizes of everything, which are literally the same as what I said above. Was it your original intention to normalize by the single mean of all landmarks in the sequence? Just curious about this as <code>dim=0, keepdim=True</code> means the opposite.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6782247%2F733adc2538a07c8859e8a7b41bcb027a%2FScreenshot%20from%202023-04-01%2008-36-47.png?generation=1680313038017988&amp;alt=media\" alt=\"\"></p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2204855,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-04-01T03:27:25.140000",
              "content": "<p><br>\n</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2204858,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-04-01T03:29:00.130000",
              "content": "<p><br>\n</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2204886,
              "author_name": "mp",
              "author_url": "",
              "post_date": "2023-04-01T04:17:58.660000",
              "content": "<p>Glad to help. I'm not sure how to do the same for unit variance as there is no such thing as <code>nanstd</code>. Please share if you know.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2204946,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-04-01T05:07:46.860000",
              "content": "<p>i make the mistake becuase of <br>\n<a href=\"https://pytorch.org/docs/stable/generated/torch.nanmean.html\" target=\"_blank\">https://pytorch.org/docs/stable/generated/torch.nanmean.html</a></p>\n<p>quote; \"(torch.nanmean(a) is equivalent to torch.mean(a[~a.isnan()])).\"</p>\n<p>i suggest you use np for preparaing pytorch train dataset.<br>\nfor tflite, you will be using tf functions, so it is not a problem<br>\n(tf.experimental.numpy.std)</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2205036,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-04-01T07:33:15.507000",
              "content": "<p><a href=\"https://www.kaggle.com/notnitsuj\" target=\"_blank\">@notnitsuj</a> </p>\n<p>sorry for the confusion. the original implementation is \"almost correct\"<br>\n<img src=\"https://i.ibb.co/095CCVD/Selection-999-1671.png\" alt=\"https://i.ibb.co/095CCVD/Selection-999-1671.png\"></p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2205057,
              "author_name": "mp",
              "author_url": "",
              "post_date": "2023-04-01T07:52:39.387000",
              "content": "<p>I wonder if it's better to normalize per frame, though. I'll test that out. <br>\nSince you're normalizing everything at once, I suggest you can just delete the arguments for <code>mean</code> as they have no effects on a single scalar and only create confusion.<br>\n[Edit] Wait, what do you mean by \"almost correct\"? Can you clarify what kind of normalization are we using here?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2205212,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-04-01T10:44:11.257000",
              "content": "<p>if your xyz is already zero mean, then the average distance of points from center(0,0,0) should be</p>\n<pre><code>var  =  np.nanstd(xyz.reshape(-1,3))**2 # mean of x**2, y**2, zz*2\ndistance_sq  = var.sum()\ndistance = distance_sq**0.5\n</code></pre>\n<p>this should be the normalising factor</p>",
              "votes": 2,
              "replies": []
            }
          ]
        },
        {
          "id": 2206355,
          "author_name": "Denys Shushpanov",
          "author_url": "",
          "post_date": "2023-04-02T13:54:54.773000",
          "content": "<p>By normalising x, y, z separately you meant sth like this? <br>\n<code>xyz.nanmean(dim=(0, 1), keepdim=True)</code>'</p>\n<p>It results in normalisation of each of dimensions separately, so that mean and std for x, y, and z are the same (notice that means are around 0):<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2423410%2F20c87321212b07a3e5d3c1be4eb9ff3f%2FScreenshot%202023-04-02%20at%2015.59.16.png?generation=1680443969118291&amp;alt=media\" alt=\"\"></p>",
          "votes": 0,
          "replies": [
            {
              "id": 2206439,
              "author_name": "mp",
              "author_url": "",
              "post_date": "2023-04-02T14:49:32.373000",
              "content": "<p>Yes, but I also wanted to normalize the spatial coordinates for each frame, which means to only reduce dim=1. But it's just my intuition and I'm not sure if it will benefit the model. You can check out hengck23's replies above.</p>\n<p>We're getting the means close to 0 and the stds close to 1 as the original data are already normalized. This step is just for those <a href=\"https://www.kaggle.com/competitions/asl-signs/discussion/392286#2171542\" target=\"_blank\">artifacts</a> that are outside of [0,1].</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2197663,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-03-26T09:48:45.067000",
      "content": "<p>[previous results]</p>\n<p>15-mar<br>\n<img src=\"https://i.ibb.co/Yf52njT/Selection-999-1544.png\" alt=\"https://i.ibb.co/Yf52njT/Selection-999-1544.png\"></p>\n<p>03-apr<br>\n<img src=\"https://i.ibb.co/TgWJk32/Selection-999-1677.png\" alt=\"https://i.ibb.co/TgWJk32/Selection-999-1677.png\"></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2185116,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-03-16T20:24:55.600000",
      "content": "<p>i just realise my work flow is:</p>\n<p>1.click edit notebook  <br>\n2.click open data in new tab  </p>\n<ul>\n<li>upload new tflite file and wait  </li>\n<li>back to notebook page click check for update </li>\n</ul>\n<p>3.click submit  </p>\n<p>i wish to hvae one button just update dataset directly from the notebook page.<br>\nno need all the \"click check for update\"  </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2185307,
          "author_name": "arutema47",
          "author_url": "",
          "post_date": "2023-03-16T23:47:50.637000",
          "content": "<p>I wish you could just submit the model.tflite directly like csv submissions. But it's handy enough😗<br>\nI like this new kind of kaggle competitions focused on edge computing. Kudos to the kaggle team making it happen!</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2185324,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-03-17T00:43:51.023000",
              "content": "<p><a href=\"https://www.kaggle.com/discussions/general/395353\" target=\"_blank\">https://www.kaggle.com/discussions/general/395353</a></p>\n<p>maybe one day we can have a chatgpt AI bot to help us with that:<br>\nprompt 'upload new tflitemodel and submit'</p>\n<p>and i think this is useful<br>\nprompt 'check license issue for this external dataset at <a href=\"http://www.xxx.xxx\" target=\"_blank\">www.xxx.xxx</a>. if there is an issue, send email to moderator and ask'</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2182036,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-03-14T22:33:06.623000",
      "content": "<p>did another try mirror augmentation?<br>\nif yes, can i have your feedbacks?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2184692,
          "author_name": "Muai",
          "author_url": "",
          "post_date": "2023-03-16T14:48:04.883000",
          "content": "<p>I've tried mirror augmentation with a simple MLP model, but didn't get better results for now, I'm still stuck at 0.63LB with and without this augmentation.<br>\nI'll experiment further this weekend and I'm interested in other feedbacks on this topic as I supposed it would mitigate the right-handed/left-handed issue.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2192090,
              "author_name": "Chinartist",
              "author_url": "",
              "post_date": "2023-03-22T12:21:33.837000",
              "content": "<p>How do you implement the mirror augmentation?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2210891,
              "author_name": "Muai",
              "author_url": "",
              "post_date": "2023-04-05T17:13:06.600000",
              "content": "<p>I simply multiply the x coordinates by -1 after normalization</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2176454,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-03-10T17:03:18.613000",
      "content": "<p>[pretrain model]</p>\n<p><a href=\"https://github.com/AI4Bharat/OpenHands\" target=\"_blank\">https://github.com/AI4Bharat/OpenHands</a><br>\npretrain model based on mediapipe !!!!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2176596,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2023-03-10T19:06:17.637000",
          "content": "<p>this is one thing what i want to do:<br>\ngenerative pretraining<br>\n[1]SignBERT: Pre-Training of Hand-Model-Aware Representation for Sign Language Recognition<br>\n<a href=\"https://arxiv.org/pdf/2110.05382.pdf\" target=\"_blank\">https://arxiv.org/pdf/2110.05382.pdf</a></p>\n<p><img src=\"https://i.ibb.co/bd6YCGS/Selection-999-1327.png\" alt=\"https://i.ibb.co/bd6YCGS/Selection-999-1327.png\"></p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 2177220,
          "author_name": "Jackson You",
          "author_url": "",
          "post_date": "2023-03-11T09:34:14.987000",
          "content": "<p>It is interesting! Do you have some thoughts about creating baseline using that pretrained model?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2178093,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-03-12T05:19:08.527000",
              "content": "<p>related: tokenization as pretraining<br>\n<img src=\"https://i.ibb.co/ZYrHjCv/Selection-999-1357.png\" alt=\"https://i.ibb.co/ZYrHjCv/Selection-999-1357.png\"></p>\n<p>BEST: BERT Pre-Training for Sign Language Recognition with Coupling Tokenization<br>\n<a href=\"https://arxiv.org/pdf/2302.05075.pdf\" target=\"_blank\">https://arxiv.org/pdf/2302.05075.pdf</a></p>\n<hr>\n<p>diffusion-based autoencoder (for augmentation?)<br>\nVector Quantized Diffusion Model with CodeUnet for Text-to-Sign Pose Sequences Generation<br>\n<a href=\"https://arxiv.org/pdf/2208.09141.pdf\" target=\"_blank\">https://arxiv.org/pdf/2208.09141.pdf</a></p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2178130,
              "author_name": "Just A game on your lips",
              "author_url": "",
              "post_date": "2023-03-12T06:24:05.213000",
              "content": "<p>It's too large architecture, I think.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2178214,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-03-12T08:16:53.277000",
              "content": "<p>you can distill later</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2181076,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-03-14T09:48:18.840000",
              "content": "<p>in theory, the max numbers of hand pattern is fixed(e.g. limited by rotation angle of finger joints, etc). hand you can actually build a \"tokenizer\" (aka lookup table)</p>\n<p>it is interesting to make a tsne plot of the the hands</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2165058,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-03-02T00:46:44.770000",
      "content": "<p>[framework conversion]</p>\n<ul>\n<li><p>keras, pytorch:<br>\n<a href=\"https://github.com/leondgarse/keras_cv_attention_models\" target=\"_blank\">https://github.com/leondgarse/keras_cv_attention_models</a></p></li>\n<li><p>pytorch, onnx, tf<br>\n<a href=\"https://github.com/PINTO0309/onnx2tf\" target=\"_blank\">https://github.com/PINTO0309/onnx2tf</a></p></li>\n</ul>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2229717,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-04-21T16:15:47.363000",
      "content": "<p>i check some of the confusing words …. i wonder if a second classifier for specific pairs (or groups) of confusing words would help (e.g. using more features, number of frames, etc)</p>\n<pre><code>167 pen (99)\n     43, 0.43434, 167 pen                     \n     28, 0.28283, 168 pencil                  \n      9, 0.09091,  40 chair                   \n      7, 0.07071, 176 potty                   \n      2, 0.02020,   0 TV        \n\n168 pencil (96)\n     66, 0.68750, 168 pencil                  \n     15, 0.15625, 167 pen                     \n      2, 0.02083,   6 another                 \n      2, 0.02083, 176 potty       \n\n129 kitty (103)\n     77, 0.74757, 129 kitty                   \n     10, 0.09709,  38 cat                     \n      3, 0.02913,  19 bee                     \n      3, 0.02913, 173 please     \n\n\n11 awake (96)\n     46, 0.47917,  11 awake                   \n     34, 0.35417, 233 wake                    \n      7, 0.07292, 131 later            \n</code></pre>",
      "votes": 2,
      "replies": [
        {
          "id": 2230055,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2023-04-22T00:56:43.143000",
          "content": "<p><a href=\"https://arxiv.org/pdf/2303.12080.pdf\" target=\"_blank\">https://arxiv.org/pdf/2303.12080.pdf</a><br>\nNatural Language-Assisted Sign Language Recognition</p>\n<p>you can see that the confused words has similary meaning (e.g. kitty/cat, pencil/pen)<br>\ncheck the paper on how to improve on this.<br>\n(this may the the reason why label smooth work for me, but according to the paper word-based label smooth works better)</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2230115,
              "author_name": "Just A game on your lips",
              "author_url": "",
              "post_date": "2023-04-22T03:23:30.280000",
              "content": "<p>Can I ask the simple question, which classes you didn't apply flip-agumentations. Thankss </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2211075,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-04-05T19:24:21.707000",
      "content": "<p>distangle hand shape and hand motion:<br>\nPhonologically-meaningful Subunits for Deep Learning-based Sign Language Recognition<br>\n<a href=\"https://www.slrtp.com/papers/full_papers/SLRTP.FP.02.012.paper.pdf\" target=\"_blank\">https://www.slrtp.com/papers/full_papers/SLRTP.FP.02.012.paper.pdf</a></p>\n<p>motion token?</p>\n<p>\"The large majority of sign language recognition systems based on deep learning adopt a word model approach. Here we present a system that works with subunits, rather than word models. We propose a pipelined approach to deep learning that uses a factorisation algorithm to derive hand motion features, embedded within a low-rank trajectory space.\"</p>\n<p>\"To separate the real hand motions from the camera and whole body movements of the signer, we propose to use a non-rigid structure from motion (NRSfM) technique based on the factorisation method [54].\"</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2200211,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-03-28T11:23:42.713000",
      "content": "<p>i probably got it all wrong.<br>\nthe correct way to shape normalise and do feature extraction is given by:</p>\n<p><a href=\"https://google.github.io/mediapipe/solutions/pose_classification.html\" target=\"_blank\">https://google.github.io/mediapipe/solutions/pose_classification.html</a><br>\n<a href=\"https://www.tensorflow.org/lite/tutorials/pose_classification\" target=\"_blank\">https://www.tensorflow.org/lite/tutorials/pose_classification</a></p>",
      "votes": 2,
      "replies": [
        {
          "id": 2200260,
          "author_name": "Mykola",
          "author_url": "",
          "post_date": "2023-03-28T12:22:35.167000",
          "content": "<pre><code>Next, convert the landmark coordinates to a feature vector by:\n\n- Moving the pose center to the origin.\n- Scaling the pose so that the pose size becomes 1\n- Flattening these coordinates into a feature vector\n</code></pre>\n<p>I am worried about \"scaling the pose so that the pose size becomes 1\". Let's say we are focused on the hand skeleton. If we have flat palm and fist - both of them would be scaled to the size of 1. Would it be beneficial for the model? I guess without per sample scaling it would be possible to extract additional information right from a scale of a given sign. What do you think about it?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2202120,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-03-29T18:46:10.453000",
              "content": "<p>\"i probably got it all wrong.\"</p>\n<p>becuase if your feature is pairwise distance, then your normalisation is also pairwise distance.<br>\nobjective of normalisation is to make feature values the same for similar class</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2205289,
              "author_name": "Vladimir Bogachev",
              "author_url": "",
              "post_date": "2023-04-01T12:07:27.503000",
              "content": "<p>Hi, my pairwise distance calculation code reduces training speed dramatically from 20 batches/s to 4 batches/s. Do you precalculate pairwise distances or do you calculate it live? Do you have similar speed reduction?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2205319,
              "author_name": "Mykola",
              "author_url": "",
              "post_date": "2023-04-01T12:32:13.977000",
              "content": "<p>If you are using tf.data.Dataset I would recommend you to use .cache() after feature calculation. It will precalculate all features and cache them. All other epochs would use cache instead of calculation all from scatch</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2189046,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-03-20T07:37:41.097000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 2193611,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-03-23T11:37:21.013000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2195003,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-03-24T10:34:29.210000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2178024,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-03-12T02:33:03.533000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2163489,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-02-28T21:14:11.653000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2230043,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-04-22T00:40:49.027000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2229823,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-04-21T18:14:11.380000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2222116,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-04-14T22:59:25.597000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2211093,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-04-05T19:39:46.310000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2209841,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-04-05T01:01:04.810000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2216136,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-04-09T19:52:14.493000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2209724,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-04-04T21:08:42.600000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2209702,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-04-04T20:46:35.190000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2204403,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-03-31T14:34:11.150000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2204423,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-03-31T14:59:33.130000",
          "content": "",
          "votes": 1,
          "replies": [
            {
              "id": 2204436,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-03-31T15:03:47.210000",
              "content": "",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2204457,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-03-31T15:28:02.627000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2204464,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-03-31T15:32:47.537000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2205038,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-04-01T07:36:22.143000",
              "content": "",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2203287,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-03-30T17:05:00.597000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2203309,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-03-30T17:33:06.200000",
          "content": "",
          "votes": 0,
          "replies": [
            {
              "id": 2205331,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-04-01T12:38:25.987000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2202934,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-03-30T12:07:27.730000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2202945,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-03-30T12:15:06.343000",
          "content": "",
          "votes": 2,
          "replies": [
            {
              "id": 2202963,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-03-30T12:25:20.450000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2202972,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-03-30T12:33:43.527000",
              "content": "",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2202977,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-03-30T12:40:41.453000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2202980,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-03-30T12:43:23.397000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2203191,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-03-30T15:44:17.403000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2203306,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-03-30T17:28:17.947000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2205366,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-04-01T13:35:07.010000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2205487,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-04-01T15:54:15.407000",
              "content": "",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2202855,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-03-30T11:24:00.570000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2202383,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-03-30T02:40:57.767000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2194498,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-03-24T01:46:16.067000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2189929,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-03-20T22:31:03.683000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2189299,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-03-20T11:43:27.953000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2188739,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-03-20T00:06:49.330000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2188788,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-03-20T01:40:20.897000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2188797,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-03-20T01:59:21.083000",
          "content": "",
          "votes": 1,
          "replies": [
            {
              "id": 2188822,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-03-20T02:50:25.633000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2189403,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-03-20T13:19:55.457000",
              "content": "",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2188724,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-03-19T23:16:49.727000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2187997,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-03-19T08:22:03.700000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2188014,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-03-19T08:29:31.507000",
          "content": "",
          "votes": 1,
          "replies": [
            {
              "id": 2188043,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-03-19T08:50:39.050000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2183926,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-03-16T03:29:44.980000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2183941,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-03-16T03:51:01.303000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2183228,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-03-15T14:37:10.477000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2183236,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-03-15T14:41:51.207000",
          "content": "",
          "votes": 0,
          "replies": [
            {
              "id": 2183317,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-03-15T15:26:18.740000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2182925,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-03-15T11:41:14.323000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2207158,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-04-03T06:34:50.307000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2182591,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-03-15T07:36:04.037000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2183260,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-03-15T14:58:31.307000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2167655,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-03-03T16:25:25.970000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2168595,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-03-04T11:07:50.940000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2187722,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-03-19T01:56:45.423000",
      "content": "",
      "votes": -5,
      "replies": [
        {
          "id": 2187785,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-03-19T03:45:59.167000",
          "content": "",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2179510,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-03-13T08:14:48.807000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2206052,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-04-02T08:05:45.377000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2163485": "experiments in progress ...\n\n![https://i.ibb.co/9TvfrhW/Selection-999-1855.png](https://i.ibb.co/9TvfrhW/Selection-999-1855.png)\n\n---\n\nprevious results:\nhttps://www.kaggle.com/competitions/asl-signs/discussion/391265#2197663\n\n\n---\n\nafter some experiments for the first week, here is the master plan:\n\n[1] one (or two) layer transformer is sufficient. for one fold, memory is only 2mb ( tf.lite.Optimize.DEFAULT, i.e. int8 in storage) and 25min for all 40\\_000 test videos (or 30 msec per video). Future because of missing frame, long term attention has its advantage\n\nin fact two layer transformer would lead to over fitting.\n\n[2] next step is to spend more time on data representation, especially normalization\n\n[3] use of external data, maybe need to use mediapipe to extract landmarks on external ASL video. \nOther ways to use increase data increases, adversarial/diffusion augmentation, masked token prediction, etc\n\nI feel that SSL with masked token prediction (e.g. generative pretraining) fits this problem well as there are high proportion of missing frames.\n\n[4], There no much of algorithm design because network is quite shallow. if there is, it would be more on increasing parameters for accuracy and speeds. There are many VIT tricks that can be used here, e.g. window attention, poolforming, convolutional positional encoding, grouped conv, ...\n\n\n----\nFirst results \n- one fold, max video length=60 (crop first) : LB 0.65 \n- 1 fold, max video length=512 (crop center) :  LB  0.65  25 min\n- 4 fold, max video length=512 (crop center) :  LB  0.68  32 min\n- 5 fold, max video length=512 (crop center) :  LB  (better) 0.68  35 min\n\nCV one fold using train/validation random split:  cv 0.72~0.74 (probably the limit if good generalisation)\nCV one fold using train/validation participant non-overlap split:  cv 0.61 (compared with above, show the severity of lack of train data) \n\ninference:\nhttps://www.kaggle.com/code/hengck23/lb-0-65-one-pytorch-transformer-solution\n\ntraining code:\nhttps://www.kaggle.com/datasets/hengck23/asl-demo\n\n\n",
    "2179326": "i have been playing with controlnet and i come across this:\n![https://i.ibb.co/0cdwBSL/Selection-999-1366.png](https://i.ibb.co/0cdwBSL/Selection-999-1366.png)\n\nmaybe can make some fun animation video while waiting for training",
    "2201909": "flip augmentation magic:\nCV 0.703 (single fold)\nLB 0.72\n\n```\nLHAND = np.arange(468, 489).tolist() # 21\nRHAND = np.arange(522, 543).tolist() # 21\nPOSE  = np.arange(489, 522).tolist() # 33\nFACE  = np.arange(0,468).tolist()    #468\n\nREYE = [\n\t33, 7, 163, 144, 145, 153, 154, 155, 133,\n\t246, 161, 160, 159, 158, 157, 173,\n]\nLEYE = [\n\t263, 249, 390, 373, 374, 380, 381, 382, 362,\n\t466, 388, 387, 386, 385, 384, 398,\n]\nNOSE=[\n\t1,2,98,327\n]\nSLIP = [\n\t78, 95, 88, 178, 87, 14, 317, 402, 318, 324, 308,\n\t191, 80, 81, 82, 13, 312, 311, 310, 415,\n]\nSPOSE = (np.array([\n\t11,13,15,12,14,16,23,24,\n])+489).tolist()\n\n\n...\n## assume zero mean xyz. so flip can be implement by multuiplication of -1\ndef do_hflip_hand(lhand, rhand):\n\trhand[...,0] *= -1\n\tlhand[...,0] *= -1\n\trhand, lhand = lhand,rhand\n\treturn lhand, rhand\n\ndef do_hflip_spose(spose):\n\tspose[...,0] *= -1\n\tspose = spose[:,[3,4,5,0,1,2,7,6]]\n\treturn spose\n\ndef do_hflip_slip(slip):\n\tslip[...,0] *= -1\n\tslip = slip[:,[10,9,8,7,6,5,4,3,2,1,0]+[19,18,17,16,15,14,13,12,11]]\n\treturn slip\n\n...\n...\n\t\txyz = load_relevant_data_subset(pq_file)\n\t\tlhand = xyz[:,LHAND]\n\t\trhand = xyz[:,RHAND]\n\t\tspose = xyz[:,SPOSE]\n\t\tleye = xyz[:,LEYE]\n\t\treye = xyz[:,REYE]\n\t\tslip = xyz[:,SLIP]\n\t\tnose = xyz[:,NOSE]\n\n\n\tif is_aug==1:\n\t\tif np.random.rand()<0.5:\n\t\t\tlhand, rhand = do_hflip_hand(lhand, rhand)\n\t\t\tspose = do_hflip_spose(spose)\n\t\t\tleye, reye = do_hflip_eye(leye, reye)\n\t\t\tslip = do_hflip_slip(slip)\n\t\t\tnose = do_hflip_nose(nose)\n\n```\n\ntrain log:\n\n```\nfold_type = kaggle-part\nfold = 0\ntrain_dataset : \n\tlen = 71518\n\tnum_participant_id = 16\n\nvalid_dataset : \n\tlen = 22959\n\tnum_participant_id = 5\n...\n   batch_size = 64 \n   experiment = ['tx025-no-z-flip', 'run_train_fold0.py']\n                           |---------------- VALID---------|---- TRAIN/BATCH ----------------------\nrate      iter       epoch | loss   top1   top2    top5    | loss                 | time           \n---------------------------------------------------------------------------------------------------\n...\n1.00e-4   00050265*  45.00 | 2.025  0.698  0.7992  0.870   | 3.887  0.000  0.000  |  1 hr 48 min\n1.00e-4   00053616*  48.00 | 2.042  0.697  0.7969  0.870   | 3.879  0.000  0.000  |  1 hr 56 min\n\n```\n\n\n",
    "2193356": "I love your experiments records table! thanks for sharing this!",
    "2178946": "[unorganized] some proven intermediate results ... will write in more detail later\n\n- add motion (xyz at current frame - xyz at previous frame). i.e. each point now has 6 dim. need to handle NaN motion correctly\nimprove CV by 0.015\n\n\n- add framerate interpolation augmentation scipy.interpolate.interp1d()\nimprove CV by 0.015\n",
    "2228313": "watch this thread!\n\nreceipe for LB0.76 one fold:\n\n---\n\ni seems that there is somthing very strange about the dataset ... maybe very noisy.  i have to use unconventional strong regularisation and training method ....\n \n\nhint: \n[1] Dropout Reduces Underfitting\nhttps://arxiv.org/abs/2303.01500\n\ni find that if i reduce learning rate, the model easily overfits. instead, according to paper[1], by setting dropout to different values at different training epoches, we can reduce validation loss (it has the same effects as reducing learning rate) without overfitting:\n- early: train without dropout\n- late: train with large dropout\n- final: tune with small dropout\n\n![https://i.ibb.co/mz15RwQ/Selection-999-1847.png](https://i.ibb.co/mz15RwQ/Selection-999-1847.png)\n\n![https://i.ibb.co/GvkZB5j/Selection-999-1848.png](https://i.ibb.co/GvkZB5j/Selection-999-1848.png)\n\n```\nfold_type = kaggle-random (20% validation, 80% train. pure random)\nfold = 0\nvalid_dataset : \n\tlen = 18896\n\tnum_participant_id = 21\n\tnum_external_video = 0\n\n\ntime_taken =  0 min 27 sec\nloss = 1.7932500839233398\ntopk[0] = 0.8607112616426758 (LB 0.76)\ntopk[1] = 0.9180249788314987\ntopk[2] = 0.9374470787468248\ntopk[3] = 0.9457557154953429\ntopk[4] = 0.9513653683319221\ncheckpoint_file = /home/titanx/hengck/share1/kaggle/2022/hand-sign/result/run202/tx90-face-norm-480-0b/fold-0-kaggle-random/checkpoint/swa.model.pth\n\n```",
    "2163630": "[reference]\n- 'Sign Pose-based Transformer for Word-level Sign Language Recognition' - Matyáš Boháček\n- Combining Efficient and Precise Sign Language Recognition: Good pose estimation library is all you need- Matyáš Boháček\n\nhttps://github.com/matyasbohacek/spoter\nhttps://www.matyasbohacek.com/\n(uses MediaPipe library pose too)\n\nhttps://huggingface.co/spaces/matyasbohacek/spoter-demo-test\n\nhttps://www.youtube.com/watch?v=6i5ODpHMny8\n\n![https://i.ibb.co/r2NHnkh/Selection-999-1166.png](https://i.ibb.co/r2NHnkh/Selection-999-1166.png)\n\n\n---\nSLGTFORMER: AN ATTENTION-BASED APPROACH TO SIGN LANGUAGE RECOGNITION\nhttps://github.com/neilsong/SLGTformer\n",
    "2163638": "[external data]\n... to be updated ...\n\nHOW TO SIGN\nhttps://how2sign.github.io/\nhttps://openaccess.thecvf.com/content/CVPR2021/papers/Duarte_How2Sign_A_Large-Scale_Multimodal_Dataset_for_Continuous_American_Sign_Language_CVPR_2021_paper.pdf\n(vocab=16k, time=79 hr, signer=11)\n\nMS-ASL American Sign Language Dataset\nhttps://microsoft.github.io/data-for-society/dataset?d=MS-ASL-American-Sign-Language-Dataset\nhttps://arxiv.org/pdf/1812.01053.pdf\n(vocab=1000, num video=25513, signer=222)\n\nBSL-1K: Scaling up co-articulated sign language recognition using mouthing cues\nhttps://www.robots.ox.ac.uk/~vgg/research/bsl1k/\n(vocab=1000, num video=1000)\n\n\nChaLearn LAP Large Scale Signer Independent Isolated Sign Language\nRecognition Challenge: Design, Results and Future Research\n\nASL-Skeleton3D and ASL-Phono: Two Novel\nDatasets for the American Sign Language\n\n---\n\n- how to read media pipline to extra keypoint for external data \n... to be udpated ...\n\n- more about PopSign learning app\nhttps://devpost.com/software/pop-sign-learning\nhttps://www.youtube.com/watch?v=nbQAfiJzp2g&t=430s\nhttps://github.com/Accessible-Technology-in-Sign/PopSign\n\nhttps://github.com/Accessible-Technology-in-Sign/ASLRT\n\"A pipeline to run tracking Kinect, MediaPipe and AlphaPose experiments on sign language feature data using either HMMs, SBHMMs, or Transformers.\"\n\ncollect data you from youtube?\n\n\n\n\n\n\n\n",
    "2226810": "i forget an import detail at the data card. suggestions for augmentation:\n\n\"While the app provided a video example of the sign desired, signers routinely made variants of the sign based on their background and region.\"  \n- you need external data or create your variant\n\n\"More rarely, signers might fingerspell a sign, miss it completely, or produce the wrong sign.\"  \n- need to measure label noise and decide if you want to exclude some data\n\n\" Extraneous movements, such as scratching an itch, or the ending movement from the previous sign or the onset of the next sign, are sometimes included.\"  \n- random add hand (from other video at start and end). actually CAM heat map is about to hightlight the useful frame.\n\n\" Conversely, some signers pressed the button late or released the button early, causing cropping in some sign examples. \"  \n- cut and paste (mix up)\n\n\"Some signers sign with their left hand; others sign with their right. Some signers switch their signing hand. All of these situations must be handled by the game’s recognition system.\"  \n- flip\n\n\n---\n\na also note that some signer make repeated signs in one video\n\n",
    "2224078": "it is suprised for me to achieve 0.75 in one fold with participant split train data(num participant = 5 in validation data). This makes me think that 0.76 is possible. \n\n```\n\nfold_type = kaggle-part\nfold = 0\nvalid_dataset : \n\tlen = 22959\n\tnum_participant_id = 5 \n \n    22959 / 22959   0 min 42 sec\n\ntime_taken =  0 min 42 sec\nloss = 2.406057596206665\ntopk[0] = 0.7277320440785748\ntopk[1] = 0.8219870203406072\ntopk[2] = 0.858530423798946\ntopk[3] = 0.8757785617840498\ntopk[4] = 0.8888888888888888\ncheckpoint_file = /home/titanx/hengck/share1/kaggle/2022/hand-sign/result/run201/tx61a-1-frame-first-scale-2/fold-0-kaggle-part/checkpoint/swa.model.pth\n \n\n```\n\nand i hit the jackpot 888888888",
    "2206909": "a couple on conv1d (as spatial temporal encoder) + MHA attention seems to have good results.\nsome good ideas to design efficient transformer.\n\n[1] FastViT: A Fast Hybrid Vision Transformer using Structural Reparameterization\nhttps://arxiv.org/pdf/2303.14189.pdf\n\n[2] Next-ViT: Next Generation Vision Transformer for Efficient Deployment in Realistic Industrial Scenarios\nhttps://arxiv.org/abs/2207.05501\n\nto avoid uncessary permute reshape, i can write my q,v,k linear operation in terms of conv1d.\ni will be removing absolute position embedding and use convolutional position embedding instead (convolutional FFN, i.e. position information is implicitly embedding on 1d conv)\n\n\nactually it is aslo possible to use 2d convolution, where (W,H) = (x,y,z of spatial, t of temporal).",
    "2204321": "this is how much the face (upper triangle = eyes and lip) and body(lower quadliteral = shoulders and hips)  have moved for one hand sign.\n\nI wonder if there will be improvement in accuracy if i remove these noise motion?\n\n![https://i.ibb.co/xgPNq9L/Selection-999-1647.png](https://i.ibb.co/xgPNq9L/Selection-999-1647.png)",
    "2206111": "simplest normalisation code\n\n```\n\nREF = [500, 501, 512, 513, 159,  386, 13,]\n\ndef do_normalise_by_ref(xyz, ref):  \n\tK = xyz.shape[-1]\n\txyz_flat = ref.reshape(-1,K)\n\tm = np.nanmean(xyz_flat,0).reshape(1,1,K)\n\ts = np.nanstd(xyz_flat, 0).mean() \n\txyz = xyz - m\n\txyz = xyz / s\n\treturn xyz\n\nxyz = load_relevant_data_subset(pq_file)[...,:2]\nxyz = do_normalise_by_ref(xyz,xyz[:,REF])\n\n```\n",
    "2178029": "[self-study]\n- Sign Language Translation with Transformers\nhttps://www.youtube.com/watch?v=E5nKeEvoAK0\n\n",
    "2163502": "[ensemble trick]\n- since only one model can be submitted and there is a time limit, we can distilled \"ensemble knowledge\" into one model.\n- a better solution is distilled encoder (or multiple encoders) and multiple heads. heads are using different \"view\" of encoded features, final prediction is averaged over the heads\n  \n\nmore on this later",
    "2163494": "[augmentation and tricks]\n- simulate different frame length \n- simulate different NaN (as in train data)\n- simulate different missing frame or differnt frame rate? \n- 3d guassian noise\n- type of landmark. use ['face', 'left_hand', 'pose', 'right_hand'] for normalisation and augmentation\n  (e.g. affine transform to group of point --> bigger mouth, longer arm ..., rotate hand)\n\nsince landmark is xyz, we are having a 3d motion capture. would be fun to use GAN to create more data. distangle motion and  shape\n\npossible to synthetic animate motion for new person via motion transfer. But we need to find some media pipe human xyz models. ",
    "2170203": "[normalization]\n\nnormalization is all you need!\n```\n\n\ndef pre_process(xyz):\n\tlip = xyz[:, LIP]\n\tlhand = xyz[:, LHAND]\n\trhand = xyz[:, RHAND]\n\txyz = torch.cat([ #(none, 82, 3)\n\t\tlip,\n\t\tlhand,\n\t\trhand,\n\t],1)\n\txyz = xyz - xyz[~torch.isnan(xyz)].mean(0,keepdims=True) #noramlisation to common maen\n\txyz = xyz / xyz[~torch.isnan(xyz)].std(0, keepdims=True)\n\txyz[torch.isnan(xyz)] = 0\n\txyz = xyz[:max_length]\n\treturn xyz\n``` ",
    "2240056": "Excellent sharing！I learned a lot from this.Thank you very much!Wish you can get the top prize!",
    "2224702": "https://github.com/facebookresearch/dropout\nDropout Reduces Underfitting\n\n![https://user-images.githubusercontent.com/8370623/222586143-3500fa5b-c294-48c9-a5cf-5fac2659e519.png](https://user-images.githubusercontent.com/8370623/222586143-3500fa5b-c294-48c9-a5cf-5fac2659e519.png)\n\nlate dropout works for me",
    "2224196": "replace the word \"academics\" with \"kaggle\"\n\n![https://i.ibb.co/rxL6Thm/Selection-999-1831.png](https://i.ibb.co/rxL6Thm/Selection-999-1831.png)",
    "2210622": "i wonder if (inverse) DTW can be used as augmentation?\nhttps://ai.googleblog.com/2021/01/recognizing-pose-similarity-in-images.html",
    "2204573": "Hi, I've noticed that you use `xyz = xyz - xyz[~torch.isnan(xyz)].mean(0,keepdim=True)` to normalize the input. However, `xyz[~torch.isnan(xyz)]` (or `np.isnan`) returns a *flattened* tensor of size `(num_frames * num_landmarks * 3 - num_nans)`, which means that calling `mean` will return a scalar. In other words, what you're really doing is subtracting the mean of everything in the sequence instead of normalizing each landmark **across frames** (passing `dim=0` and `keepdim=True`). Am I getting it correctly?\n\nNevertheless, I believe the correct mean normalization here is to subtract the mean of all landmark positions **for each frame** by doing something like `xyz -= xyz.nanmean(dim=1, keepdim=True)`. It might be even better to normalize x, y, and z separately but I haven't figured out how to do it yet.",
    "2197663": "[previous results]\n\n15-mar\n![https://i.ibb.co/Yf52njT/Selection-999-1544.png](https://i.ibb.co/Yf52njT/Selection-999-1544.png)\n\n03-apr\n![https://i.ibb.co/TgWJk32/Selection-999-1677.png](https://i.ibb.co/TgWJk32/Selection-999-1677.png)\n",
    "2185116": "i just realise my work flow is:\n\n1.click edit notebook  \n2.click open data in new tab  \n - upload new tflite file and wait  \n - back to notebook page click check for update \n \n3.click submit  \n\ni wish to hvae one button just update dataset directly from the notebook page.\nno need all the \"click check for update\"  ",
    "2182036": "did another try mirror augmentation?\nif yes, can i have your feedbacks?",
    "2176454": "[pretrain model]\n\nhttps://github.com/AI4Bharat/OpenHands\npretrain model based on mediapipe !!!!",
    "2165058": "[framework conversion]\n\n- keras, pytorch:\nhttps://github.com/leondgarse/keras_cv_attention_models\n\n- pytorch, onnx, tf\nhttps://github.com/PINTO0309/onnx2tf",
    "2229717": "i check some of the confusing words .... i wonder if a second classifier for specific pairs (or groups) of confusing words would help (e.g. using more features, number of frames, etc)\n\n```\n\n167 pen (99)\n\t 43, 0.43434, 167 pen                     \n\t 28, 0.28283, 168 pencil                  \n\t  9, 0.09091,  40 chair                   \n\t  7, 0.07071, 176 potty                   \n\t  2, 0.02020,   0 TV        \n\n168 pencil (96)\n\t 66, 0.68750, 168 pencil                  \n\t 15, 0.15625, 167 pen                     \n\t  2, 0.02083,   6 another                 \n\t  2, 0.02083, 176 potty       \n\n129 kitty (103)\n\t 77, 0.74757, 129 kitty                   \n\t 10, 0.09709,  38 cat                     \n\t  3, 0.02913,  19 bee                     \n\t  3, 0.02913, 173 please     \n\n\n11 awake (96)\n\t 46, 0.47917,  11 awake                   \n\t 34, 0.35417, 233 wake                    \n\t  7, 0.07292, 131 later            \n```\n\n\n\n\n",
    "2211075": "distangle hand shape and hand motion:\nPhonologically-meaningful Subunits for Deep Learning-based Sign Language Recognition\nhttps://www.slrtp.com/papers/full_papers/SLRTP.FP.02.012.paper.pdf\n\nmotion token?\n\n\"The large majority of sign language recognition systems based on deep learning adopt a word model approach. Here we present a system that works with subunits, rather than word models. We propose a pipelined approach to deep learning that uses a factorisation algorithm to derive hand motion features, embedded within a low-rank trajectory space.\"\n\n\n\"To separate the real hand motions from the camera and whole body movements of the signer, we propose to use a non-rigid structure from motion (NRSfM) technique based on the factorisation method [54].\"\n\n",
    "2200211": "i probably got it all wrong.\nthe correct way to shape normalise and do feature extraction is given by:\n\nhttps://google.github.io/mediapipe/solutions/pose_classification.html\nhttps://www.tensorflow.org/lite/tutorials/pose_classification\n",
    "2189046": "we have an interesting results of LB0.71 with single fold using all train data",
    "2178024": "[opensource implementation]\n\n[0] ST-GCN \nhttps://github.com/yysijie/st-gcn\n\n[1] SL-GCN \nhttps://github.com/jackyjsy/SAM-SLR-v2\n\n[2] ECCV 2022 Sign Spotting Challenge 3rd ranked solution\nhttps://chalearnlap.cvc.uab.cat/media/results/None/top3-track1_H1KGRhQ.pdf \nhttps://github.com/ycmin95/Chalearn_2022_Sign_Spotting_MSSL_track/tree/master/dataset\n\ninclude how to use mediapipe to generate data\n",
    "2163489": "[tflite conversion]\n\n-  https://github.com/onnx/onnx-tensorflow\n-  fallback plan (aka write my own converter):\n    - re-code in keras. need to perform one to one check for each nn.Module(), tf.keras.layers().\n    - copy weights from pytorch to keras (trick is to give good names to layer for easier code)\n    - keras to tflite",
    "2230043": "extra attribute for training (aux loss). if you have external training set, you can now use non kaggle gloss (word labels) for training too!\n\n\nTowards Zero-shot Sign Language Recognition\nhttps://arxiv.org/pdf/2201.05914.pdf\nhttps://bmvc2019.org/wp-content/uploads/papers/0122-paper.pdf\n\n\"We further annotate the datasets with high-level attributes that are gathered from American Sign Language Hand Shape Dictionary [87].\"\n\na page from American Sign Language Hand Shape Dictionary:\n![https://i.ibb.co/Zzdt317/6e980c3584164c851f4b2d6788bb46626712a27c.jpg](https://i.ibb.co/Zzdt317/6e980c3584164c851f4b2d6788bb46626712a27c.jpg)\n\nso you have attribute like:\n- palm orientation (in, out, up, down, left, or right)\n- movement (up, down, left, right, inward, outward, circular, wrist movement, finger movement)\n- hand shape (A-hand, S-hand, 5-hand, etc.)\n\n\n![https://i.ibb.co/3TdzLM5/Selection-999-1856.png](https://i.ibb.co/3TdzLM5/Selection-999-1856.png)\n\n",
    "2229823": "how do you get the point_dim 1414 in the lastest update?",
    "2222116": "there is another way (maybe better?) to approach the problem:\nsegmentation : \n- classify every frame  (instead of one video), i.e. making frame features\n- you can combine several videos during training\n- if you want you can train an aggregator at the end  to combine frame features into video features to do video prediction\n\n---\n\nan extension will be classify every N-frame intervals, etc to learn videolet features \n\n---\n\nyou can using classification to perform vector quantisation (VQ) ...  kinda fake AE",
    "2211093": "joint angle as a feature:\nhttps://github.com/google/mediapipe/issues/2999\nhttps://github.com/TemugeB/joint_angles_calculate",
    "2209841": "Could you say which of participants are in fold 0 in your table?",
    "2209724": "how to augment:\n\nconceptually:\n![https://i.ibb.co/FsskMh7/Selection-999-1690.png](https://i.ibb.co/FsskMh7/Selection-999-1690.png)\n\nneed to constraint the augmentation correctly.\nmore operations will be stretch shear\n\nyou can observe a few coordinate images of the same word of the same/different signers\n\n---\n\nactually one can perform classification on these coordinate images (i.e. treat it as an image problem) if you can render perform 2d convolution fast enough ... ",
    "2209702": "a possible way to select pairwise distance feature:\n\n\nSequential Attention for Feature Selection\nhttps://arxiv.org/abs/2209.14881",
    "2204403": "Thank you for sharing. Did you use entire dataset as training data? How can you check the accuracy of a model at each checkpoint when using all train samples?",
    "2203287": "I am getting out of memory error, even when taking emb_dim = 256, head = 4, transformer_layer = 1. any optimization i am missing out on?",
    "2202934": "May I ask are you still using 1-layer transformer? I can't believe that one attention pooling can achieve 73 single-fold score. \n\nOr maybe your deeper embedding layer (fc-norm-act-fc-norm-act) can do the same feature processing as more transformer layers?",
    "2202855": "Label smoothing  and flip didn't work for me.",
    "2202383": "Hi @hengck23 , what is cls and what is all token mean pooling?",
    "2194498": "i find it easier to write your own attention.\nyou can control q,k,v dim differently etc.\nyou can control local attention, e.g. which query is attent to which values (according hand, lip parts or joints)\n\nyou can learned the affinity matrix of graph CGN if you treat affinity=attention weights",
    "2189929": "Many thanks for you work @hengck23.\nDo you ensemble your models or do you just take the best model through CV?\nWhen I ensemble, it seems that inference exceeds an hour...",
    "2189299": "can you do validation without validation data?\ne.g. if i used all the train images for training, how do i know if there is over fitting?\n\nin theory, yes ....\nyou can still use the following to judge if there is overfitting:\n1. we usually have a training set and validation set. then we monitor the training and validation loss.\n2. But there are other indicators that is correlated to degree of overfitting. Validation loss is only one of them\n3. say if you have only train data, you can measure \n- degree of regularization and effects on train loss (for different models of different design)\n- rate of change of loss/prediction as data is perturbed\n- rate of change of loss/prediction as parameters is perturbed\n- observe attention, heatmap, etc ....\n\nyou should think of how these indicators change for an optimum, over-fitted, under-fitted trained models.\n\nBut of course having another set of data is still the best way to check over-fitting. so instead of finding methods and indicators, you may just find another set of data (external or synthetically created)\n\n\n",
    "2188739": "Thank you so much for sharing your approach @hengck23  I was able to train the model from [here](https://www.kaggle.com/datasets/hengck23/asl-demo) \nI pretty much changed no settings except epochs set to 20, because I want the pipeline to work first. \nI used the [inference](https://www.kaggle.com/code/hengck23/lb-0-65-one-pytorch-transformer-solution) however I'm getting OOM even with `max_length=40`  \nAny suggestions ? ",
    "2188724": "thinking of implementing a differentiable topk feature selector ... ",
    "2187997": "Thanks for your share！！ can I ask what do you mean of \"use triu\"",
    "2183926": "> one (or two) layer transformer is sufficient\n\nhmmm I have different results to this, deeper transformer lead to a higher CV:\n\n1 layer   0.50\n2 layers 0.53\n3 layers 0.55\n4 layers 0.56",
    "2183228": "Hi~ You said that `one (or two) layer transformer is sufficient`, is that means we only need 1 layer of MultiHeadAttention? In the paper \"Attention is all you need\", a complete transformer contains 6 duplicated encoders and decoders, and each encoder contains one MultiHeadAttention Layer and something else. I learned your code [lb-0-67-one-pytorch-transformer-solution](https://www.kaggle.com/code/hengck23/lb-0-67-one-pytorch-transformer-solution), and I think you mean we just need one layer of MultiHeadAttention but not just one layer of transformer?\n\nBest wishes\nhappyOldMan",
    "2182925": "[team forming advertisement]\n\nLooking for experienced Kagglers to form a team. We are interested in working with 8-bit quantization aware training with Keras or pytorch and implementation in tflite. \n\nIt's important to note that our main objective is to learn from experts in deep learning technology, specifically embedded implementation. While we haven't yet decided if we will use an int8 model for our submission, as we are aware that only TFLite supports 8-bit computation for ARM CPU.\n\nWe're committed to achieving the highest ranking on LB, but we also want to emphasize that Kaggle is about having fun and enjoying the competition.\n\nIf this sounds like a team you'd like to be a part of, please submit your application and state your experiences in 8-bit quantization.\nPlease note that only shortlisted candidates will be contacted.\n\nThank you!\n\n(drafted by chatgpt)\nlast update 15-mar",
    "2182591": "Thanks for sharing your experiment result, I’m curious about the val acc you got. I’ve seen many got a pretty high val acc. (around 0.75), your top1 acc is lower that that, what may be the reason？\nIn my transformer model, I tried many different modifications to the transformer block and it’s hard to break 0.7 on val acc, I got my highest acc at 0.69.",
    "2167655": ">time_taken = 35 sec (gpu, btach=32)\n\nMay I ask if this is the time of one epoch? Or is it a time of 100 epoch? I'm training on tensorflow, but it feels slow.",
    "2187722": "",
    "2179510": "",
    "2206052": "Thanks for sharing!!"
  }
}