{
  "id": 397433,
  "title": "Transformers vs Others",
  "url": "/competitions/asl-signs/discussion/397433",
  "author_name": "chemdatafarmer",
  "post_date": "2023-03-25T16:02:00.972000",
  "votes": 5,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Hi Everyone,</p>\n<p>I got into this competition to learn a bit more about time series data, as well as to try out some graph-based neural networks. So far my experiments with GNNs have been edifying, but the resulting models haven't been competitive with other models.</p>\n<p>I see there are a good number of LB scores &gt;= 0.7 now. Has anyone had luck achieving this kind of score without using a transformer?</p>\n<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> and <a href=\"https://www.kaggle.com/markwijkhuizen\" target=\"_blank\">@markwijkhuizen</a> have some fantastic notebooks on transformer solutions (great educational resource) and I'm simply interested if anyone's had luck with other architectures being able to obtain a similar level of performance.</p>",
  "messages": [
    {
      "id": 2196963,
      "postDate": "2023-03-25T18:49:25.417Z",
      "content": "<p>We don't use transformers. Single model, single fold 0.73</p>",
      "rawMarkdown": "We don't use transformers. Single model, single fold 0.73",
      "votes": 6,
      "replies": [
        {
          "id": 2196976,
          "postDate": "2023-03-25T19:03:54.823Z",
          "content": "<p>I can’t reach better than LB 0.7 on single fold. How do you preprocess input? </p>",
          "rawMarkdown": "I can’t reach better than LB 0.7 on single fold. How do you preprocess input? ",
          "votes": 1
        },
        {
          "id": 2197017,
          "postDate": "2023-03-25T19:55:55.953Z",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/kolyaforrat\" target=\"_blank\">@kolyaforrat</a> thanks for sharing! I'm glad to see there will be different approaches that reach competitive scores. It's also nice to hear this is the result from a single model with a single fold validation. I'm really enjoying this competition and learning from the other folks participating. </p>",
          "rawMarkdown": "Hey @kolyaforrat thanks for sharing! I'm glad to see there will be different approaches that reach competitive scores. It's also nice to hear this is the result from a single model with a single fold validation. I'm really enjoying this competition and learning from the other folks participating. ",
          "replies": [
            {
              "id": 2202672,
              "postDate": "2023-03-30T08:18:40.620Z",
              "content": "<p>use all train sample<br>\nflip augmentation<br>\nLB 0.73 for one \"fold\", one \"transformer\" model<br>\n<a href=\"https://www.kaggle.com/code/hengck23/lb-0-73-single-fold-transformer-architecture\" target=\"_blank\">https://www.kaggle.com/code/hengck23/lb-0-73-single-fold-transformer-architecture</a></p>",
              "rawMarkdown": "use all train sample\nflip augmentation\nLB 0.73 for one \"fold\", one \"transformer\" model\nhttps://www.kaggle.com/code/hengck23/lb-0-73-single-fold-transformer-architecture"
            },
            {
              "id": 2204781,
              "postDate": "2023-03-31T23:32:46.737Z",
              "content": "<p>Thanks <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> !</p>",
              "rawMarkdown": "Thanks @hengck23 !"
            }
          ]
        }
      ]
    },
    {
      "id": 2196708,
      "postDate": "2023-03-25T16:02:00.973Z",
      "content": "<p>Hi Everyone,</p>\n<p>I got into this competition to learn a bit more about time series data, as well as to try out some graph-based neural networks. So far my experiments with GNNs have been edifying, but the resulting models haven't been competitive with other models.</p>\n<p>I see there are a good number of LB scores &gt;= 0.7 now. Has anyone had luck achieving this kind of score without using a transformer?</p>\n<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> and <a href=\"https://www.kaggle.com/markwijkhuizen\" target=\"_blank\">@markwijkhuizen</a> have some fantastic notebooks on transformer solutions (great educational resource) and I'm simply interested if anyone's had luck with other architectures being able to obtain a similar level of performance.</p>",
      "rawMarkdown": "Hi Everyone,\n\nI got into this competition to learn a bit more about time series data, as well as to try out some graph-based neural networks. So far my experiments with GNNs have been edifying, but the resulting models haven't been competitive with other models.\n\nI see there are a good number of LB scores >= 0.7 now. Has anyone had luck achieving this kind of score without using a transformer?\n\n@hengck23 and @markwijkhuizen have some fantastic notebooks on transformer solutions (great educational resource) and I'm simply interested if anyone's had luck with other architectures being able to obtain a similar level of performance.",
      "votes": 5
    },
    {
      "id": 2196825,
      "postDate": "2023-03-25T17:28:56.760Z",
      "content": "<p>I am also interested in this. Especially, if there is a single model, not an ensemble of different approaches.</p>",
      "rawMarkdown": "I am also interested in this. Especially, if there is a single model, not an ensemble of different approaches.",
      "votes": 1
    },
    {
      "id": 2217996,
      "postDate": "2023-04-11T10:25:59.953Z",
      "content": "<p>I also tried to use GCN based models, but haven't got any luck yet. Can I know whether you are using all the landmarks or just the hand?</p>",
      "rawMarkdown": "I also tried to use GCN based models, but haven't got any luck yet. Can I know whether you are using all the landmarks or just the hand?",
      "replies": [
        {
          "id": 2218674,
          "postDate": "2023-04-11T23:03:17.310Z",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/timxinxinpeng\" target=\"_blank\">@timxinxinpeng</a> I have tried all sorts of variations here. I've tried GCNs, GATs, things I found in papers…Eventually I got some things to sort of work, but they were much slower than other models like transformers or conv nets and didn't seem to provide any advantage. I tried just hands, hands+pose, hands+pose+facial_features, etc with different linking strategies to break regularities in the graph and to make the most important nodes pass messages between eachother…Perhaps someone else has had better luck.</p>",
          "rawMarkdown": "Hey @timxinxinpeng I have tried all sorts of variations here. I've tried GCNs, GATs, things I found in papers...Eventually I got some things to sort of work, but they were much slower than other models like transformers or conv nets and didn't seem to provide any advantage. I tried just hands, hands+pose, hands+pose+facial_features, etc with different linking strategies to break regularities in the graph and to make the most important nodes pass messages between eachother...Perhaps someone else has had better luck."
        }
      ]
    },
    {
      "id": 2200173,
      "postDate": "2023-03-28T10:46:29.200Z",
      "content": "<p>Is transformer possible to be combined with PopSign?</p>",
      "rawMarkdown": "Is transformer possible to be combined with PopSign?"
    },
    {
      "id": 2197313,
      "postDate": "2023-03-26T04:34:00.230Z",
      "content": "<p>even it is a transformer, it is only one layer. it is more like attention pooling</p>",
      "rawMarkdown": "even it is a transformer, it is only one layer. it is more like attention pooling",
      "replies": [
        {
          "id": 2197798,
          "postDate": "2023-03-26T12:46:46.857Z",
          "content": "<p>Great point <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> thank you. I've been spending the morning trying to better understand self attention mechanisms and reading papers that use it in different ways for similar tasks. Your comment helped solidify the idea for me :).</p>",
          "rawMarkdown": "Great point @hengck23 thank you. I've been spending the morning trying to better understand self attention mechanisms and reading papers that use it in different ways for similar tasks. Your comment helped solidify the idea for me :).",
          "replies": [
            {
              "id": 2197814,
              "postDate": "2023-03-26T13:02:54.460Z",
              "content": "<p>For the record I have nothing against transformers, in fact I think they are rather interesting. I was just curious if they were clearly prevailing, or if there were other ways to cross the 0.7 threshold.</p>",
              "rawMarkdown": "For the record I have nothing against transformers, in fact I think they are rather interesting. I was just curious if they were clearly prevailing, or if there were other ways to cross the 0.7 threshold."
            }
          ]
        }
      ]
    },
    {
      "id": 2198926,
      "postDate": "2023-03-27T11:21:07.093Z",
      "rawMarkdown": "",
      "votes": -1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2196963,
      "author_name": "Kolya Forrat",
      "author_url": "",
      "post_date": "2023-03-25T18:49:25.417000",
      "content": "<p>We don't use transformers. Single model, single fold 0.73</p>",
      "votes": 6,
      "replies": [
        {
          "id": 2196976,
          "author_name": "A.P.",
          "author_url": "",
          "post_date": "2023-03-25T19:03:54.823000",
          "content": "<p>I can’t reach better than LB 0.7 on single fold. How do you preprocess input? </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2197017,
          "author_name": "chemdatafarmer",
          "author_url": "",
          "post_date": "2023-03-25T19:55:55.953000",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/kolyaforrat\" target=\"_blank\">@kolyaforrat</a> thanks for sharing! I'm glad to see there will be different approaches that reach competitive scores. It's also nice to hear this is the result from a single model with a single fold validation. I'm really enjoying this competition and learning from the other folks participating. </p>",
          "votes": 0,
          "replies": [
            {
              "id": 2202672,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-03-30T08:18:40.620000",
              "content": "<p>use all train sample<br>\nflip augmentation<br>\nLB 0.73 for one \"fold\", one \"transformer\" model<br>\n<a href=\"https://www.kaggle.com/code/hengck23/lb-0-73-single-fold-transformer-architecture\" target=\"_blank\">https://www.kaggle.com/code/hengck23/lb-0-73-single-fold-transformer-architecture</a></p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2204781,
              "author_name": "chemdatafarmer",
              "author_url": "",
              "post_date": "2023-03-31T23:32:46.737000",
              "content": "<p>Thanks <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> !</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2196825,
      "author_name": "Mykola",
      "author_url": "",
      "post_date": "2023-03-25T17:28:56.760000",
      "content": "<p>I am also interested in this. Especially, if there is a single model, not an ensemble of different approaches.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2217996,
      "author_name": "Tim Xinxin Peng",
      "author_url": "",
      "post_date": "2023-04-11T10:25:59.953000",
      "content": "<p>I also tried to use GCN based models, but haven't got any luck yet. Can I know whether you are using all the landmarks or just the hand?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2218674,
          "author_name": "chemdatafarmer",
          "author_url": "",
          "post_date": "2023-04-11T23:03:17.310000",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/timxinxinpeng\" target=\"_blank\">@timxinxinpeng</a> I have tried all sorts of variations here. I've tried GCNs, GATs, things I found in papers…Eventually I got some things to sort of work, but they were much slower than other models like transformers or conv nets and didn't seem to provide any advantage. I tried just hands, hands+pose, hands+pose+facial_features, etc with different linking strategies to break regularities in the graph and to make the most important nodes pass messages between eachother…Perhaps someone else has had better luck.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2200173,
      "author_name": "HongCheng",
      "author_url": "",
      "post_date": "2023-03-28T10:46:29.200000",
      "content": "<p>Is transformer possible to be combined with PopSign?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2197313,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-03-26T04:34:00.230000",
      "content": "<p>even it is a transformer, it is only one layer. it is more like attention pooling</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2197798,
          "author_name": "chemdatafarmer",
          "author_url": "",
          "post_date": "2023-03-26T12:46:46.857000",
          "content": "<p>Great point <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> thank you. I've been spending the morning trying to better understand self attention mechanisms and reading papers that use it in different ways for similar tasks. Your comment helped solidify the idea for me :).</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2197814,
              "author_name": "chemdatafarmer",
              "author_url": "",
              "post_date": "2023-03-26T13:02:54.460000",
              "content": "<p>For the record I have nothing against transformers, in fact I think they are rather interesting. I was just curious if they were clearly prevailing, or if there were other ways to cross the 0.7 threshold.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2198926,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-03-27T11:21:07.093000",
      "content": "",
      "votes": -1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2196963": "We don't use transformers. Single model, single fold 0.73",
    "2196708": "Hi Everyone,\n\nI got into this competition to learn a bit more about time series data, as well as to try out some graph-based neural networks. So far my experiments with GNNs have been edifying, but the resulting models haven't been competitive with other models.\n\nI see there are a good number of LB scores >= 0.7 now. Has anyone had luck achieving this kind of score without using a transformer?\n\n@hengck23 and @markwijkhuizen have some fantastic notebooks on transformer solutions (great educational resource) and I'm simply interested if anyone's had luck with other architectures being able to obtain a similar level of performance.",
    "2196825": "I am also interested in this. Especially, if there is a single model, not an ensemble of different approaches.",
    "2217996": "I also tried to use GCN based models, but haven't got any luck yet. Can I know whether you are using all the landmarks or just the hand?",
    "2200173": "Is transformer possible to be combined with PopSign?",
    "2197313": "even it is a transformer, it is only one layer. it is more like attention pooling",
    "2198926": ""
  }
}