{
  "id": 401034,
  "title": "First public notebook using GCN archetecture",
  "url": "/competitions/asl-signs/discussion/401034",
  "author_name": "",
  "post_date": "2023-04-11T12:36:16.176156Z",
  "votes": 4,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I made a <a href=\"https://www.kaggle.com/code/timxinxinpeng/isolated-sign-language-recognition-with-stgcn\" target=\"_blank\">notebook</a> which implements the ST-GCN architecture, but the performance is bizarely low. Actually it has been common practice to use GCN-based model for ISLR tasks, like <a href=\"https://speechsigns.uvigo.es/wp-content/uploads/2022/04/Vazquez-Enriquez_Isolated_Sign_Language_Recognition_With_Multi-Scale_Spatial-Temporal_Graph_Convolutional_Networks_CVPRW_2021_paper.pdf\" target=\"_blank\">https://speechsigns.uvigo.es/wp-content/uploads/2022/04/Vazquez-Enriquez_Isolated_Sign_Language_Recognition_With_Multi-Scale_Spatial-Temporal_Graph_Convolutional_Networks_CVPRW_2021_paper.pdf</a> and <a href=\"https://www.semanticscholar.org/paper/Spatial-Temporal-Graph-Convolutional-Networks-for-Amorim-Mac%C3%AAdo/8b9e818a72e73c6730e618eaf8389a672748c212\" target=\"_blank\">https://www.semanticscholar.org/paper/Spatial-Temporal-Graph-Convolutional-Networks-for-Amorim-Mac%C3%AAdo/8b9e818a72e73c6730e618eaf8389a672748c212</a> . But I don't know why my attempt to use ST-GCN for this competition is a total failure and almost no ppl else mentioned they used GCN for this task.</p>",
  "messages": [
    {
      "id": "2218097",
      "postDate": "04/11/2023 12:36:16",
      "content": "<p>I made a <a href=\"https://www.kaggle.com/code/timxinxinpeng/isolated-sign-language-recognition-with-stgcn\" target=\"_blank\">notebook</a> which implements the ST-GCN architecture, but the performance is bizarely low. Actually it has been common practice to use GCN-based model for ISLR tasks, like <a href=\"https://speechsigns.uvigo.es/wp-content/uploads/2022/04/Vazquez-Enriquez_Isolated_Sign_Language_Recognition_With_Multi-Scale_Spatial-Temporal_Graph_Convolutional_Networks_CVPRW_2021_paper.pdf\" target=\"_blank\">https://speechsigns.uvigo.es/wp-content/uploads/2022/04/Vazquez-Enriquez_Isolated_Sign_Language_Recognition_With_Multi-Scale_Spatial-Temporal_Graph_Convolutional_Networks_CVPRW_2021_paper.pdf</a> and <a href=\"https://www.semanticscholar.org/paper/Spatial-Temporal-Graph-Convolutional-Networks-for-Amorim-Mac%C3%AAdo/8b9e818a72e73c6730e618eaf8389a672748c212\" target=\"_blank\">https://www.semanticscholar.org/paper/Spatial-Temporal-Graph-Convolutional-Networks-for-Amorim-Mac%C3%AAdo/8b9e818a72e73c6730e618eaf8389a672748c212</a> . But I don't know why my attempt to use ST-GCN for this competition is a total failure and almost no ppl else mentioned they used GCN for this task.</p>",
      "rawMarkdown": "I made a [notebook](https://www.kaggle.com/code/timxinxinpeng/isolated-sign-language-recognition-with-stgcn) which implements the ST-GCN architecture, but the performance is bizarely low. Actually it has been common practice to use GCN-based model for ISLR tasks, like https://speechsigns.uvigo.es/wp-content/uploads/2022/04/Vazquez-Enriquez_Isolated_Sign_Language_Recognition_With_Multi-Scale_Spatial-Temporal_Graph_Convolutional_Networks_CVPRW_2021_paper.pdf and https://www.semanticscholar.org/paper/Spatial-Temporal-Graph-Convolutional-Networks-for-Amorim-Mac%C3%AAdo/8b9e818a72e73c6730e618eaf8389a672748c212 . But I don't know why my attempt to use ST-GCN for this competition is a total failure and almost no ppl else mentioned they used GCN for this task.",
      "votes": null
    },
    {
      "id": "2225672",
      "postDate": "04/18/2023 10:44:16",
      "content": "<p>I think you could be on to something ! You have missed out the preprocessing normalisation of the skeleton data which seems important. check this out <a href=\"https://github.com/kdkalvik/ST-GCN/blob/master/data_gen/preprocess.py\" target=\"_blank\">https://github.com/kdkalvik/ST-GCN/blob/master/data_gen/preprocess.py</a>. Also since the model you have created only looks at the hands it would be worth checking that the hands actually appear in the 12 frames you have chosen by using the current TFRecord setup - likely only 1-6frames actually contains a hand so lots of training on NAN values.<br>\nAlso worth remembering that one hand is always empty - the participants only sign with one hand - so one is always full of NAN values !<br>\nGood luck !!</p>",
      "rawMarkdown": "I think you could be on to something ! You have missed out the preprocessing normalisation of the skeleton data which seems important. check this out https://github.com/kdkalvik/ST-GCN/blob/master/data_gen/preprocess.py. Also since the model you have created only looks at the hands it would be worth checking that the hands actually appear in the 12 frames you have chosen by using the current TFRecord setup - likely only 1-6frames actually contains a hand so lots of training on NAN values.\nAlso worth remembering that one hand is always empty - the participants only sign with one hand - so one is always full of NAN values !\nGood luck !!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2225672,
      "author_name": "blankfish",
      "author_url": "",
      "post_date": "04/18/2023 10:44:16",
      "content": "<p>I think you could be on to something ! You have missed out the preprocessing normalisation of the skeleton data which seems important. check this out <a href=\"https://github.com/kdkalvik/ST-GCN/blob/master/data_gen/preprocess.py\" target=\"_blank\">https://github.com/kdkalvik/ST-GCN/blob/master/data_gen/preprocess.py</a>. Also since the model you have created only looks at the hands it would be worth checking that the hands actually appear in the 12 frames you have chosen by using the current TFRecord setup - likely only 1-6frames actually contains a hand so lots of training on NAN values.<br>\nAlso worth remembering that one hand is always empty - the participants only sign with one hand - so one is always full of NAN values !<br>\nGood luck !!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2218097": "I made a [notebook](https://www.kaggle.com/code/timxinxinpeng/isolated-sign-language-recognition-with-stgcn) which implements the ST-GCN architecture, but the performance is bizarely low. Actually it has been common practice to use GCN-based model for ISLR tasks, like https://speechsigns.uvigo.es/wp-content/uploads/2022/04/Vazquez-Enriquez_Isolated_Sign_Language_Recognition_With_Multi-Scale_Spatial-Temporal_Graph_Convolutional_Networks_CVPRW_2021_paper.pdf and https://www.semanticscholar.org/paper/Spatial-Temporal-Graph-Convolutional-Networks-for-Amorim-Mac%C3%AAdo/8b9e818a72e73c6730e618eaf8389a672748c212 . But I don't know why my attempt to use ST-GCN for this competition is a total failure and almost no ppl else mentioned they used GCN for this task.",
    "2225672": "I think you could be on to something ! You have missed out the preprocessing normalisation of the skeleton data which seems important. check this out https://github.com/kdkalvik/ST-GCN/blob/master/data_gen/preprocess.py. Also since the model you have created only looks at the hands it would be worth checking that the hands actually appear in the 12 frames you have chosen by using the current TFRecord setup - likely only 1-6frames actually contains a hand so lots of training on NAN values.\nAlso worth remembering that one hand is always empty - the participants only sign with one hand - so one is always full of NAN values !\nGood luck !!"
  },
  "source": "meta"
}