{
  "id": 395704,
  "title": "RNN-based vs algorithm-based approach",
  "url": "/competitions/asl-signs/discussion/395704",
  "author_name": "",
  "post_date": "2023-03-18T12:02:51.851303300Z",
  "votes": 4,
  "comment_count": 6,
  "views": 0,
  "content": "<p>What do you think about using algorithm-based methods for landmarks such as graph models? Can it be useful? Will LSTM, GRU, Transformer solutions work better?</p>",
  "messages": [
    {
      "id": "2187082",
      "postDate": "03/18/2023 12:02:51",
      "content": "<p>What do you think about using algorithm-based methods for landmarks such as graph models? Can it be useful? Will LSTM, GRU, Transformer solutions work better?</p>",
      "rawMarkdown": "What do you think about using algorithm-based methods for landmarks such as graph models? Can it be useful? Will LSTM, GRU, Transformer solutions work better?",
      "votes": null
    },
    {
      "id": "2187635",
      "postDate": "03/18/2023 22:13:22",
      "content": "<p>I've been using graph models in this competition, mostly as a playground to learn more about them and to see how they do with this sort of data. Thus far, I haven't seen any particular improvements over other approaches I've seen/tried, but I haven't had time to test everything I've wanted to yet. I've mostly been working on applying basic GCNs (Kipf and Welling style) in various ways, so there's a lot more to do. That said, I'm starting to think simple GCNs don't learn enough from the immediate node neighbors to make them stand out vs other approaches. If anyone has found the opposite, I'd love to learn about the approach they're using.</p>\n<p>I know <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> implemented a nice transformer solution that's been doing well on the leaderboard and their approach has been well documented in discussions and notebooks (bonus points for handling the pytorch -&gt; tflite conversion). <a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> has a pretty interesting notebook that uses feature engineering to train a feed forward nn that's also done pretty well particularly given it's simplicity. There are a couple good notebooks out there for LSTMs and GRUs as well.</p>",
      "rawMarkdown": "I've been using graph models in this competition, mostly as a playground to learn more about them and to see how they do with this sort of data. Thus far, I haven't seen any particular improvements over other approaches I've seen/tried, but I haven't had time to test everything I've wanted to yet. I've mostly been working on applying basic GCNs (Kipf and Welling style) in various ways, so there's a lot more to do. That said, I'm starting to think simple GCNs don't learn enough from the immediate node neighbors to make them stand out vs other approaches. If anyone has found the opposite, I'd love to learn about the approach they're using.\n\nI know @hengck23 implemented a nice transformer solution that's been doing well on the leaderboard and their approach has been well documented in discussions and notebooks (bonus points for handling the pytorch -> tflite conversion). @roberthatch has a pretty interesting notebook that uses feature engineering to train a feed forward nn that's also done pretty well particularly given it's simplicity. There are a couple good notebooks out there for LSTMs and GRUs as well.",
      "votes": null
    },
    {
      "id": "2188275",
      "postDate": "03/19/2023 12:59:09",
      "content": "<p>Got it. I think that makes sense. Maybe we can use graph model as an additional branch or for generating additional features. Big tnx for your answer.</p>",
      "rawMarkdown": "Got it. I think that makes sense. Maybe we can use graph model as an additional branch or for generating additional features. Big tnx for your answer.",
      "votes": null
    },
    {
      "id": "2188529",
      "postDate": "03/19/2023 18:10:50",
      "content": "<p>No problem :). I've been trying to use GAT_V2 today from the paper \"How Attentive Are Graph Neural Networks?\", but it seems like its major contribution (at least for how I've attempted to employ it) has just been to slow things down XD. </p>\n<p>I was hoping that the attention mechanism in this paper would help capture longer range dependencies between nodes, but so far I've seen no evidence of that.</p>",
      "rawMarkdown": "No problem :). I've been trying to use GAT_V2 today from the paper \"How Attentive Are Graph Neural Networks?\", but it seems like its major contribution (at least for how I've attempted to employ it) has just been to slow things down XD. \n\nI was hoping that the attention mechanism in this paper would help capture longer range dependencies between nodes, but so far I've seen no evidence of that.",
      "votes": null
    },
    {
      "id": "2188805",
      "postDate": "03/20/2023 02:18:27",
      "content": "<p>if you want to use GCN, you should follow the work of TC-GCN and  SL-GCN<br>\nthey have pretrained model of mediapiple as well</p>\n<p><a href=\"https://arxiv.org/pdf/2103.08833.pdf\" target=\"_blank\">https://arxiv.org/pdf/2103.08833.pdf</a><br>\n<a href=\"https://github.com/neilsong/SLGTformer\" target=\"_blank\">https://github.com/neilsong/SLGTformer</a></p>",
      "rawMarkdown": "if you want to use GCN, you should follow the work of TC-GCN and  SL-GCN\nthey have pretrained model of mediapiple as well\n\nhttps://arxiv.org/pdf/2103.08833.pdf\nhttps://github.com/neilsong/SLGTformer",
      "votes": null
    },
    {
      "id": "2191355",
      "postDate": "03/21/2023 23:04:54",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> ! I will have to check it out </p>",
      "rawMarkdown": "Thanks @hengck23 ! I will have to check it out",
      "votes": null
    },
    {
      "id": "2204796",
      "postDate": "04/01/2023 00:38:18",
      "content": "<p>These papers, and others, have been really useful. Starting to find graph-based methods that are getting closer to the 0.7 mark :).</p>",
      "rawMarkdown": "These papers, and others, have been really useful. Starting to find graph-based methods that are getting closer to the 0.7 mark :).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2187635,
      "author_name": "chemdatafarmer",
      "author_url": "",
      "post_date": "03/18/2023 22:13:22",
      "content": "<p>I've been using graph models in this competition, mostly as a playground to learn more about them and to see how they do with this sort of data. Thus far, I haven't seen any particular improvements over other approaches I've seen/tried, but I haven't had time to test everything I've wanted to yet. I've mostly been working on applying basic GCNs (Kipf and Welling style) in various ways, so there's a lot more to do. That said, I'm starting to think simple GCNs don't learn enough from the immediate node neighbors to make them stand out vs other approaches. If anyone has found the opposite, I'd love to learn about the approach they're using.</p>\n<p>I know <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> implemented a nice transformer solution that's been doing well on the leaderboard and their approach has been well documented in discussions and notebooks (bonus points for handling the pytorch -&gt; tflite conversion). <a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> has a pretty interesting notebook that uses feature engineering to train a feed forward nn that's also done pretty well particularly given it's simplicity. There are a couple good notebooks out there for LSTMs and GRUs as well.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2188275,
          "author_name": "pe4eniks",
          "author_url": "",
          "post_date": "03/19/2023 12:59:09",
          "content": "<p>Got it. I think that makes sense. Maybe we can use graph model as an additional branch or for generating additional features. Big tnx for your answer.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2188529,
              "author_name": "chemdatafarmer",
              "author_url": "",
              "post_date": "03/19/2023 18:10:50",
              "content": "<p>No problem :). I've been trying to use GAT_V2 today from the paper \"How Attentive Are Graph Neural Networks?\", but it seems like its major contribution (at least for how I've attempted to employ it) has just been to slow things down XD. </p>\n<p>I was hoping that the attention mechanism in this paper would help capture longer range dependencies between nodes, but so far I've seen no evidence of that.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2188805,
                  "author_name": "hengck23",
                  "author_url": "",
                  "post_date": "03/20/2023 02:18:27",
                  "content": "<p>if you want to use GCN, you should follow the work of TC-GCN and  SL-GCN<br>\nthey have pretrained model of mediapiple as well</p>\n<p><a href=\"https://arxiv.org/pdf/2103.08833.pdf\" target=\"_blank\">https://arxiv.org/pdf/2103.08833.pdf</a><br>\n<a href=\"https://github.com/neilsong/SLGTformer\" target=\"_blank\">https://github.com/neilsong/SLGTformer</a></p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2191355,
                      "author_name": "chemdatafarmer",
                      "author_url": "",
                      "post_date": "03/21/2023 23:04:54",
                      "content": "<p>Thanks <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> ! I will have to check it out </p>",
                      "votes": null,
                      "replies": []
                    },
                    {
                      "id": 2204796,
                      "author_name": "chemdatafarmer",
                      "author_url": "",
                      "post_date": "04/01/2023 00:38:18",
                      "content": "<p>These papers, and others, have been really useful. Starting to find graph-based methods that are getting closer to the 0.7 mark :).</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2187082": "What do you think about using algorithm-based methods for landmarks such as graph models? Can it be useful? Will LSTM, GRU, Transformer solutions work better?",
    "2187635": "I've been using graph models in this competition, mostly as a playground to learn more about them and to see how they do with this sort of data. Thus far, I haven't seen any particular improvements over other approaches I've seen/tried, but I haven't had time to test everything I've wanted to yet. I've mostly been working on applying basic GCNs (Kipf and Welling style) in various ways, so there's a lot more to do. That said, I'm starting to think simple GCNs don't learn enough from the immediate node neighbors to make them stand out vs other approaches. If anyone has found the opposite, I'd love to learn about the approach they're using.\n\nI know @hengck23 implemented a nice transformer solution that's been doing well on the leaderboard and their approach has been well documented in discussions and notebooks (bonus points for handling the pytorch -> tflite conversion). @roberthatch has a pretty interesting notebook that uses feature engineering to train a feed forward nn that's also done pretty well particularly given it's simplicity. There are a couple good notebooks out there for LSTMs and GRUs as well.",
    "2188275": "Got it. I think that makes sense. Maybe we can use graph model as an additional branch or for generating additional features. Big tnx for your answer.",
    "2188529": "No problem :). I've been trying to use GAT_V2 today from the paper \"How Attentive Are Graph Neural Networks?\", but it seems like its major contribution (at least for how I've attempted to employ it) has just been to slow things down XD. \n\nI was hoping that the attention mechanism in this paper would help capture longer range dependencies between nodes, but so far I've seen no evidence of that.",
    "2188805": "if you want to use GCN, you should follow the work of TC-GCN and  SL-GCN\nthey have pretrained model of mediapiple as well\n\nhttps://arxiv.org/pdf/2103.08833.pdf\nhttps://github.com/neilsong/SLGTformer",
    "2191355": "Thanks @hengck23 ! I will have to check it out",
    "2204796": "These papers, and others, have been really useful. Starting to find graph-based methods that are getting closer to the 0.7 mark :)."
  },
  "source": "meta"
}