{
  "id": 500639,
  "title": "Transformers doesn't perform??",
  "url": "/competitions/leash-BELKA/discussion/500639",
  "author_name": "",
  "post_date": "2024-05-06T12:04:05.543466100Z",
  "votes": 5,
  "comment_count": 2,
  "views": 0,
  "content": "<p>After reviewing past competitions focused on protein-related tasks, I noted that the most successful models employed Transformer architectures. So, I decided to give them a try, However, the performance was disappointingly poor. Has anyone else experienced similar outcomes with Transformers?</p>\n<p>Given these results, I am considering GNNs but it read it was extremely computationally expensive.</p>",
  "messages": [
    {
      "id": "2796787",
      "postDate": "05/06/2024 12:04:05",
      "content": "<p>After reviewing past competitions focused on protein-related tasks, I noted that the most successful models employed Transformer architectures. So, I decided to give them a try, However, the performance was disappointingly poor. Has anyone else experienced similar outcomes with Transformers?</p>\n<p>Given these results, I am considering GNNs but it read it was extremely computationally expensive.</p>",
      "rawMarkdown": "After reviewing past competitions focused on protein-related tasks, I noted that the most successful models employed Transformer architectures. So, I decided to give them a try, However, the performance was disappointingly poor. Has anyone else experienced similar outcomes with Transformers?\n\nGiven these results, I am considering GNNs but it read it was extremely computationally expensive.",
      "votes": null
    },
    {
      "id": "2799135",
      "postDate": "05/07/2024 15:39:16",
      "content": "<p>transformer from scratch work for me.</p>\n<p>check that you have use the correct learning rate. transformer needs smaller learning rate.</p>\n<p>my suggestion:</p>\n<ol>\n<li>build a simple 3 layer conv1d as baseline</li>\n<li>modify the above to hierarchical scaled version to reduce computation (e.g. length L --&gt; avg pooled to 1/2 L --&gt;1/4 L)</li>\n<li>you can simply add some encoder transformer layer on top of conv1d. </li>\n</ol>\n<p>For each of the model1,2,3 change parameters to keep performance better or the same as baseline. you should easily get LB near 0.600 (cv 0.700)</p>",
      "rawMarkdown": "transformer from scratch work for me.\n\ncheck that you have use the correct learning rate. transformer needs smaller learning rate.\n\nmy suggestion:\n1. build a simple 3 layer conv1d as baseline\n2. modify the above to hierarchical scaled version to reduce computation (e.g. length L --> avg pooled to 1/2 L -->1/4 L)\n3. you can simply add some encoder transformer layer on top of conv1d. \n\nFor each of the model1,2,3 change parameters to keep performance better or the same as baseline. you should easily get LB near 0.600 (cv 0.700)",
      "votes": null
    },
    {
      "id": "2800275",
      "postDate": "05/08/2024 06:25:11",
      "content": "<p>Okey i will try different learning rates and some dimensionality reduction. Thanks for the help!</p>",
      "rawMarkdown": "Okey i will try different learning rates and some dimensionality reduction. Thanks for the help!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2799135,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "05/07/2024 15:39:16",
      "content": "<p>transformer from scratch work for me.</p>\n<p>check that you have use the correct learning rate. transformer needs smaller learning rate.</p>\n<p>my suggestion:</p>\n<ol>\n<li>build a simple 3 layer conv1d as baseline</li>\n<li>modify the above to hierarchical scaled version to reduce computation (e.g. length L --&gt; avg pooled to 1/2 L --&gt;1/4 L)</li>\n<li>you can simply add some encoder transformer layer on top of conv1d. </li>\n</ol>\n<p>For each of the model1,2,3 change parameters to keep performance better or the same as baseline. you should easily get LB near 0.600 (cv 0.700)</p>",
      "votes": null,
      "replies": [
        {
          "id": 2800275,
          "author_name": "alrjandrorivasespa",
          "author_url": "",
          "post_date": "05/08/2024 06:25:11",
          "content": "<p>Okey i will try different learning rates and some dimensionality reduction. Thanks for the help!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2796787": "After reviewing past competitions focused on protein-related tasks, I noted that the most successful models employed Transformer architectures. So, I decided to give them a try, However, the performance was disappointingly poor. Has anyone else experienced similar outcomes with Transformers?\n\nGiven these results, I am considering GNNs but it read it was extremely computationally expensive.",
    "2799135": "transformer from scratch work for me.\n\ncheck that you have use the correct learning rate. transformer needs smaller learning rate.\n\nmy suggestion:\n1. build a simple 3 layer conv1d as baseline\n2. modify the above to hierarchical scaled version to reduce computation (e.g. length L --> avg pooled to 1/2 L -->1/4 L)\n3. you can simply add some encoder transformer layer on top of conv1d. \n\nFor each of the model1,2,3 change parameters to keep performance better or the same as baseline. you should easily get LB near 0.600 (cv 0.700)",
    "2800275": "Okey i will try different learning rates and some dimensionality reduction. Thanks for the help!"
  },
  "source": "meta"
}