{
  "id": 71105,
  "title": "Let's talk about Transformer?",
  "url": "/competitions/quora-insincere-questions-classification/discussion/71105",
  "author_name": "",
  "post_date": "2018-11-10T09:10:54.374173800Z",
  "votes": 23,
  "comment_count": 8,
  "views": 0,
  "content": "<p>If you haven't tried it yet, I can confirm that it fits nicely into a kernel.</p>\n\n<p>The version I've been using is the <a href=\"https://arxiv.org/abs/1801.10198\">Transformer-Decoder</a>, which in theory can be nearly twice as fast as the original version. The problem is its accuracy, which is terrible atm when it could barely get 0.6 F1 local.</p>\n\n<p>I will release my stuff when it's ready but I'm not sure if I'm going in the right direction. What are your experiments with it, does it really work?</p>",
  "messages": [
    {
      "id": "418612",
      "postDate": "11/10/2018 09:10:54",
      "content": "<p>If you haven't tried it yet, I can confirm that it fits nicely into a kernel.</p>\n\n<p>The version I've been using is the <a href=\"https://arxiv.org/abs/1801.10198\">Transformer-Decoder</a>, which in theory can be nearly twice as fast as the original version. The problem is its accuracy, which is terrible atm when it could barely get 0.6 F1 local.</p>\n\n<p>I will release my stuff when it's ready but I'm not sure if I'm going in the right direction. What are your experiments with it, does it really work?</p>",
      "rawMarkdown": "If you haven't tried it yet, I can confirm that it fits nicely into a kernel.\n\nThe version I've been using is the [Transformer-Decoder][1], which in theory can be nearly twice as fast as the original version. The problem is its accuracy, which is terrible atm when it could barely get 0.6 F1 local.\n\nI will release my stuff when it's ready but I'm not sure if I'm going in the right direction. What are your experiments with it, does it really work?\n\n  [1]: https://arxiv.org/abs/1801.10198",
      "votes": null
    },
    {
      "id": "419269",
      "postDate": "11/11/2018 16:03:11",
      "content": "<p>I am trying hard to implement transformer on kernels too ... Not getting very good results either for now. </p>\n\n<p>BTW, I am just wondering why people downvote this kind of post ?</p>",
      "rawMarkdown": "I am trying hard to implement transformer on kernels too ... Not getting very good results either for now. \n\nBTW, I am just wondering why people downvote this kind of post ?",
      "votes": null
    },
    {
      "id": "419372",
      "postDate": "11/11/2018 20:26:19",
      "content": "<p>What does your classifier look like? In the original paper they said that they only took the last row of the final attention matrix and add a dense layer there, it didn't work well for me.  Flattening the whole matrix helped a bit but the score is still poor. I'm thinking about using a CNN.</p>",
      "rawMarkdown": "What does your classifier look like? In the original paper they said that they only took the last row of the final attention matrix and add a dense layer there, it didn't work well for me.  Flattening the whole matrix helped a bit but the score is still poor. I'm thinking about using a CNN.",
      "votes": null
    },
    {
      "id": "420036",
      "postDate": "11/13/2018 00:48:36",
      "content": "<p>This paper from EMNLP 2018 may help, though not exactly this topic: <a href=\"https://arxiv.org/abs/1808.08946\">https://arxiv.org/abs/1808.08946</a></p>\n\n<blockquote>\n  <p>Our experimental results show that: 1) self-attentional networks and CNNs do not outperform RNNs in modeling subject-verb agreement over long distances; 2) self-attentional networks perform distinctly better than RNNs and CNNs on word sense disambiguation.</p>\n</blockquote>",
      "rawMarkdown": "This paper from EMNLP 2018 may help, though not exactly this topic: https://arxiv.org/abs/1808.08946\n\n&gt; Our experimental results show that: 1) self-attentional networks and CNNs do not outperform RNNs in modeling subject-verb agreement over long distances; 2) self-attentional networks perform distinctly better than RNNs and CNNs on word sense disambiguation.",
      "votes": null
    },
    {
      "id": "420692",
      "postDate": "11/14/2018 02:02:42",
      "content": "<p>My initial attempts to apply the transformer encoder to classification: <a href=\"https://www.kaggle.com/shujian/transformer-initial-attempt\">https://www.kaggle.com/shujian/transformer-initial-attempt</a> The result is not good. I will keep working on it this weekend.</p>",
      "rawMarkdown": "My initial attempts to apply the transformer encoder to classification: https://www.kaggle.com/shujian/transformer-initial-attempt The result is not good. I will keep working on it this weekend.",
      "votes": null
    },
    {
      "id": "420716",
      "postDate": "11/14/2018 02:55:14",
      "content": "<p>Nice job. I'll be looking forward to your results.</p>",
      "rawMarkdown": "Nice job. I'll be looking forward to your results.",
      "votes": null
    },
    {
      "id": "440948",
      "postDate": "12/18/2018 05:43:57",
      "content": "<p>Does BERT model work well on the task?</p>",
      "rawMarkdown": "Does BERT model work well on the task?",
      "votes": null
    },
    {
      "id": "449379",
      "postDate": "01/03/2019 03:48:38",
      "content": "<p>I only use Transformer encoder like bert, and I get the same F1 score as you. May be it takes more epochs to converge.</p>",
      "rawMarkdown": "I only use Transformer encoder like bert, and I get the same F1 score as you. May be it takes more epochs to converge.",
      "votes": null
    },
    {
      "id": "470245",
      "postDate": "02/12/2019 16:38:37",
      "content": "<p>I have tried to implement the Transformer for a trading problem (predicting stock price), since it is sequential data I thought it might work. My initial attempt wasn't so good since the results are nearly around 50-51% accuracy. There's a lot of parameters to change. But I'll try harder, it could be fun. </p>",
      "rawMarkdown": "I have tried to implement the Transformer for a trading problem (predicting stock price), since it is sequential data I thought it might work. My initial attempt wasn't so good since the results are nearly around 50-51% accuracy. There's a lot of parameters to change. But I'll try harder, it could be fun.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 419269,
      "author_name": "serigne",
      "author_url": "",
      "post_date": "11/11/2018 16:03:11",
      "content": "<p>I am trying hard to implement transformer on kernels too ... Not getting very good results either for now. </p>\n\n<p>BTW, I am just wondering why people downvote this kind of post ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 419372,
          "author_name": "suicaokhoailang",
          "author_url": "",
          "post_date": "11/11/2018 20:26:19",
          "content": "<p>What does your classifier look like? In the original paper they said that they only took the last row of the final attention matrix and add a dense layer there, it didn't work well for me.  Flattening the whole matrix helped a bit but the score is still poor. I'm thinking about using a CNN.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 420036,
      "author_name": "shujian",
      "author_url": "",
      "post_date": "11/13/2018 00:48:36",
      "content": "<p>This paper from EMNLP 2018 may help, though not exactly this topic: <a href=\"https://arxiv.org/abs/1808.08946\">https://arxiv.org/abs/1808.08946</a></p>\n\n<blockquote>\n  <p>Our experimental results show that: 1) self-attentional networks and CNNs do not outperform RNNs in modeling subject-verb agreement over long distances; 2) self-attentional networks perform distinctly better than RNNs and CNNs on word sense disambiguation.</p>\n</blockquote>",
      "votes": null,
      "replies": []
    },
    {
      "id": 420692,
      "author_name": "shujian",
      "author_url": "",
      "post_date": "11/14/2018 02:02:42",
      "content": "<p>My initial attempts to apply the transformer encoder to classification: <a href=\"https://www.kaggle.com/shujian/transformer-initial-attempt\">https://www.kaggle.com/shujian/transformer-initial-attempt</a> The result is not good. I will keep working on it this weekend.</p>",
      "votes": null,
      "replies": [
        {
          "id": 420716,
          "author_name": "suicaokhoailang",
          "author_url": "",
          "post_date": "11/14/2018 02:55:14",
          "content": "<p>Nice job. I'll be looking forward to your results.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 440948,
      "author_name": "lagkanacl",
      "author_url": "",
      "post_date": "12/18/2018 05:43:57",
      "content": "<p>Does BERT model work well on the task?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 449379,
      "author_name": "zengpingchen",
      "author_url": "",
      "post_date": "01/03/2019 03:48:38",
      "content": "<p>I only use Transformer encoder like bert, and I get the same F1 score as you. May be it takes more epochs to converge.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 470245,
      "author_name": "miljan",
      "author_url": "",
      "post_date": "02/12/2019 16:38:37",
      "content": "<p>I have tried to implement the Transformer for a trading problem (predicting stock price), since it is sequential data I thought it might work. My initial attempt wasn't so good since the results are nearly around 50-51% accuracy. There's a lot of parameters to change. But I'll try harder, it could be fun. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "418612": "If you haven't tried it yet, I can confirm that it fits nicely into a kernel.\n\nThe version I've been using is the [Transformer-Decoder][1], which in theory can be nearly twice as fast as the original version. The problem is its accuracy, which is terrible atm when it could barely get 0.6 F1 local.\n\nI will release my stuff when it's ready but I'm not sure if I'm going in the right direction. What are your experiments with it, does it really work?\n\n  [1]: https://arxiv.org/abs/1801.10198",
    "419269": "I am trying hard to implement transformer on kernels too ... Not getting very good results either for now. \n\nBTW, I am just wondering why people downvote this kind of post ?",
    "419372": "What does your classifier look like? In the original paper they said that they only took the last row of the final attention matrix and add a dense layer there, it didn't work well for me.  Flattening the whole matrix helped a bit but the score is still poor. I'm thinking about using a CNN.",
    "420036": "This paper from EMNLP 2018 may help, though not exactly this topic: https://arxiv.org/abs/1808.08946\n\n&gt; Our experimental results show that: 1) self-attentional networks and CNNs do not outperform RNNs in modeling subject-verb agreement over long distances; 2) self-attentional networks perform distinctly better than RNNs and CNNs on word sense disambiguation.",
    "420692": "My initial attempts to apply the transformer encoder to classification: https://www.kaggle.com/shujian/transformer-initial-attempt The result is not good. I will keep working on it this weekend.",
    "420716": "Nice job. I'll be looking forward to your results.",
    "440948": "Does BERT model work well on the task?",
    "449379": "I only use Transformer encoder like bert, and I get the same F1 score as you. May be it takes more epochs to converge.",
    "470245": "I have tried to implement the Transformer for a trading problem (predicting stock price), since it is sequential data I thought it might work. My initial attempt wasn't so good since the results are nearly around 50-51% accuracy. There's a lot of parameters to change. But I'll try harder, it could be fun."
  },
  "source": "meta"
}