{
  "id": 229640,
  "title": "non autoregressive decoder?",
  "url": "/competitions/bms-molecular-translation/discussion/229640",
  "author_name": "",
  "post_date": "2021-03-31T04:07:04.274709400Z",
  "votes": 9,
  "comment_count": 13,
  "views": 0,
  "content": "<p>as my specialization is in computer vision and i am not so familiar with NLP and seq2seq methods.</p>\n<p>i do a quick search on google for non-autoregressive decoder approaches but cannot find good papers.</p>\n<p>does anyone has recommendations on non-autoregressive decoder that has shown the same accuracy as autoregressive decoder?</p>\n<p>e.g: <a href=\"https://blog.einstein.ai/fully-parallel-text-generation-for-neural-machine-translation/\" target=\"_blank\">https://blog.einstein.ai/fully-parallel-text-generation-for-neural-machine-translation/</a></p>",
  "messages": [
    {
      "id": "1257728",
      "postDate": "03/31/2021 04:07:04",
      "content": "<p>as my specialization is in computer vision and i am not so familiar with NLP and seq2seq methods.</p>\n<p>i do a quick search on google for non-autoregressive decoder approaches but cannot find good papers.</p>\n<p>does anyone has recommendations on non-autoregressive decoder that has shown the same accuracy as autoregressive decoder?</p>\n<p>e.g: <a href=\"https://blog.einstein.ai/fully-parallel-text-generation-for-neural-machine-translation/\" target=\"_blank\">https://blog.einstein.ai/fully-parallel-text-generation-for-neural-machine-translation/</a></p>",
      "rawMarkdown": "as my specialization is in computer vision and i am not so familiar with NLP and seq2seq methods.\n\ni do a quick search on google for non-autoregressive decoder approaches but cannot find good papers.\n\ndoes anyone has recommendations on non-autoregressive decoder that has shown the same accuracy as autoregressive decoder?\n\ne.g: https://blog.einstein.ai/fully-parallel-text-generation-for-neural-machine-translation/",
      "votes": null
    },
    {
      "id": "1258250",
      "postDate": "03/31/2021 13:36:14",
      "content": "<p>Here is a recent paper on improved non-autoregressive machine translation: <a href=\"https://arxiv.org/abs/2012.15833v1\" target=\"_blank\">https://arxiv.org/abs/2012.15833v1</a><br>\nI don't know more, as it isn't my specialization either.</p>",
      "rawMarkdown": "Here is a recent paper on improved non-autoregressive machine translation: https://arxiv.org/abs/2012.15833v1\nI don't know more, as it isn't my specialization either.",
      "votes": null
    },
    {
      "id": "1258351",
      "postDate": "03/31/2021 14:49:06",
      "content": "<p>Just echoing heng here, I am also trying to perform non-autoregressive decoding, so far with dismal success. I'll bump this if I find anything useful.</p>",
      "rawMarkdown": "Just echoing heng here, I am also trying to perform non-autoregressive decoding, so far with dismal success. I'll bump this if I find anything useful.",
      "votes": null
    },
    {
      "id": "1259269",
      "postDate": "04/01/2021 09:18:03",
      "content": "<p>This would be cool to me too, +20h decoding process every time I want to submit.</p>",
      "rawMarkdown": "This would be cool to me too, +20h decoding process every time I want to submit.",
      "votes": null
    },
    {
      "id": "1260745",
      "postDate": "04/02/2021 11:42:13",
      "content": "<p>Blockwise Parallel Decoding for Deep Autoregressive Models<br>\n<a href=\"https://www.youtube.com/watch?v=3Tqp_B2G6u0\" target=\"_blank\">https://www.youtube.com/watch?v=3Tqp_B2G6u0</a></p>\n<p><img src=\"https://user-images.githubusercontent.com/7529838/48524307-388f9200-e8c3-11e8-8f3b-0fa19cc947b4.png\" alt=\"\"><br>\n<a href=\"https://github.com/kweonwooj/papers/issues/116\" target=\"_blank\">https://github.com/kweonwooj/papers/issues/116</a></p>",
      "rawMarkdown": "Blockwise Parallel Decoding for Deep Autoregressive Models\nhttps://www.youtube.com/watch?v=3Tqp_B2G6u0\n\n![](https://user-images.githubusercontent.com/7529838/48524307-388f9200-e8c3-11e8-8f3b-0fa19cc947b4.png)\nhttps://github.com/kweonwooj/papers/issues/116",
      "votes": null
    },
    {
      "id": "1260769",
      "postDate": "04/02/2021 11:57:38",
      "content": "<p>i figure out what to do:</p>\n<ul>\n<li>since non-auto aggressive perform worse in accuracy, i would not use it for submission</li>\n<li>but i can use it for hyperparameter searches and experiments. they provide lower bound for autoagressive models</li>\n</ul>",
      "rawMarkdown": "i figure out what to do:\n- since non-auto aggressive perform worse in accuracy, i would not use it for submission\n- but i can use it for hyperparameter searches and experiments. they provide lower bound for autoagressive models",
      "votes": null
    },
    {
      "id": "1260794",
      "postDate": "04/02/2021 12:16:46",
      "content": "<p><img src=\"https://i.ibb.co/w6HChT8/Selection-025.png\" alt=\"\"></p>",
      "rawMarkdown": "![](https://i.ibb.co/w6HChT8/Selection-025.png)",
      "votes": null
    },
    {
      "id": "1260806",
      "postDate": "04/02/2021 12:22:17",
      "content": "<p>does this affect the training process?</p>",
      "rawMarkdown": "does this affect the training process?",
      "votes": null
    },
    {
      "id": "1260811",
      "postDate": "04/02/2021 12:26:46",
      "content": "<p>the paper and youtube describe the training process in details.</p>\n<p>but basically</p>\n<ol>\n<li>you train your transformer as usual (i.e. the one by one decoder)</li>\n<li>then you add layers and build on top of the decoder (freezing it) so that you can decode one to multiple outputs.</li>\n<li>if you want, you can finetune end 2 end later</li>\n</ol>",
      "rawMarkdown": "the paper and youtube describe the training process in details.\n\nbut basically\n1. you train your transformer as usual (i.e. the one by one decoder)\n2. then you add layers and build on top of the decoder (freezing it) so that you can decode one to multiple outputs.\n3. if you want, you can finetune end 2 end later",
      "votes": null
    },
    {
      "id": "1260863",
      "postDate": "04/02/2021 13:20:49",
      "content": "<p>Thank you, I'll have a look!</p>",
      "rawMarkdown": "Thank you, I'll have a look!",
      "votes": null
    },
    {
      "id": "1280115",
      "postDate": "04/21/2021 15:20:58",
      "content": "<p>unlike language translation or image captioning, there this only one inchi string given an input image.<br>\nHence i think we can directly predict the whole inchi string instead of decoding one by one.</p>\n<p>i wonder if anyone has done this experiment?</p>",
      "rawMarkdown": "unlike language translation or image captioning, there this only one inchi string given an input image.\nHence i think we can directly predict the whole inchi string instead of decoding one by one.\n\ni wonder if anyone has done this experiment?",
      "votes": null
    },
    {
      "id": "1280285",
      "postDate": "04/21/2021 19:40:51",
      "content": "<p>Holy moly. </p>\n<p>That’s all.</p>",
      "rawMarkdown": "Holy moly. \n\nThat’s all.",
      "votes": null
    },
    {
      "id": "1280369",
      "postDate": "04/21/2021 20:59:34",
      "content": "<p>You can check TTS papers, newer approaches predict length first and then fillout i.e. FastSpeech</p>",
      "rawMarkdown": "You can check TTS papers, newer approaches predict length first and then fillout i.e. FastSpeech",
      "votes": null
    },
    {
      "id": "1283969",
      "postDate": "04/25/2021 12:59:29",
      "content": "<p>this is what i long suspected and i think it is very much applicable to this bms compeition</p>\n<p>DEEP ENCODER, SHALLOW DECODER: REEVALUATING NON-AUTOREGRESSIVE MACHINE TRANSLATION -ICRL 2021</p>\n<p>\"Our extensive experiments show that given a sufficiently deep encoder, a single-layer autoregressive decoder can substantially outperform strong non-autoregressive models with comparable inference speed\"</p>\n<p>take home message:<br>\nmake encoder deep<br>\nmake decoder shallow and fast</p>",
      "rawMarkdown": "this is what i long suspected and i think it is very much applicable to this bms compeition\n\nDEEP ENCODER, SHALLOW DECODER: REEVALUATING NON-AUTOREGRESSIVE MACHINE TRANSLATION -ICRL 2021\n\n\"Our extensive experiments show that given a sufficiently deep encoder, a single-layer autoregressive decoder can substantially outperform strong non-autoregressive models with comparable inference speed\"\n\ntake home message:\nmake encoder deep\nmake decoder shallow and fast",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1258250,
      "author_name": "cepheidq",
      "author_url": "",
      "post_date": "03/31/2021 13:36:14",
      "content": "<p>Here is a recent paper on improved non-autoregressive machine translation: <a href=\"https://arxiv.org/abs/2012.15833v1\" target=\"_blank\">https://arxiv.org/abs/2012.15833v1</a><br>\nI don't know more, as it isn't my specialization either.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1258351,
      "author_name": "interneuron",
      "author_url": "",
      "post_date": "03/31/2021 14:49:06",
      "content": "<p>Just echoing heng here, I am also trying to perform non-autoregressive decoding, so far with dismal success. I'll bump this if I find anything useful.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1259269,
      "author_name": "claverru",
      "author_url": "",
      "post_date": "04/01/2021 09:18:03",
      "content": "<p>This would be cool to me too, +20h decoding process every time I want to submit.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1280285,
          "author_name": "dschettler8845",
          "author_url": "",
          "post_date": "04/21/2021 19:40:51",
          "content": "<p>Holy moly. </p>\n<p>That’s all.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1260745,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "04/02/2021 11:42:13",
      "content": "<p>Blockwise Parallel Decoding for Deep Autoregressive Models<br>\n<a href=\"https://www.youtube.com/watch?v=3Tqp_B2G6u0\" target=\"_blank\">https://www.youtube.com/watch?v=3Tqp_B2G6u0</a></p>\n<p><img src=\"https://user-images.githubusercontent.com/7529838/48524307-388f9200-e8c3-11e8-8f3b-0fa19cc947b4.png\" alt=\"\"><br>\n<a href=\"https://github.com/kweonwooj/papers/issues/116\" target=\"_blank\">https://github.com/kweonwooj/papers/issues/116</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1260794,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "04/02/2021 12:16:46",
          "content": "<p><img src=\"https://i.ibb.co/w6HChT8/Selection-025.png\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1260806,
          "author_name": "claverru",
          "author_url": "",
          "post_date": "04/02/2021 12:22:17",
          "content": "<p>does this affect the training process?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1260811,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "04/02/2021 12:26:46",
          "content": "<p>the paper and youtube describe the training process in details.</p>\n<p>but basically</p>\n<ol>\n<li>you train your transformer as usual (i.e. the one by one decoder)</li>\n<li>then you add layers and build on top of the decoder (freezing it) so that you can decode one to multiple outputs.</li>\n<li>if you want, you can finetune end 2 end later</li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1260863,
          "author_name": "claverru",
          "author_url": "",
          "post_date": "04/02/2021 13:20:49",
          "content": "<p>Thank you, I'll have a look!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1260769,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "04/02/2021 11:57:38",
      "content": "<p>i figure out what to do:</p>\n<ul>\n<li>since non-auto aggressive perform worse in accuracy, i would not use it for submission</li>\n<li>but i can use it for hyperparameter searches and experiments. they provide lower bound for autoagressive models</li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1280115,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "04/21/2021 15:20:58",
      "content": "<p>unlike language translation or image captioning, there this only one inchi string given an input image.<br>\nHence i think we can directly predict the whole inchi string instead of decoding one by one.</p>\n<p>i wonder if anyone has done this experiment?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1280369,
          "author_name": "tugstugi",
          "author_url": "",
          "post_date": "04/21/2021 20:59:34",
          "content": "<p>You can check TTS papers, newer approaches predict length first and then fillout i.e. FastSpeech</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1283969,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "04/25/2021 12:59:29",
      "content": "<p>this is what i long suspected and i think it is very much applicable to this bms compeition</p>\n<p>DEEP ENCODER, SHALLOW DECODER: REEVALUATING NON-AUTOREGRESSIVE MACHINE TRANSLATION -ICRL 2021</p>\n<p>\"Our extensive experiments show that given a sufficiently deep encoder, a single-layer autoregressive decoder can substantially outperform strong non-autoregressive models with comparable inference speed\"</p>\n<p>take home message:<br>\nmake encoder deep<br>\nmake decoder shallow and fast</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1257728": "as my specialization is in computer vision and i am not so familiar with NLP and seq2seq methods.\n\ni do a quick search on google for non-autoregressive decoder approaches but cannot find good papers.\n\ndoes anyone has recommendations on non-autoregressive decoder that has shown the same accuracy as autoregressive decoder?\n\ne.g: https://blog.einstein.ai/fully-parallel-text-generation-for-neural-machine-translation/",
    "1258250": "Here is a recent paper on improved non-autoregressive machine translation: https://arxiv.org/abs/2012.15833v1\nI don't know more, as it isn't my specialization either.",
    "1258351": "Just echoing heng here, I am also trying to perform non-autoregressive decoding, so far with dismal success. I'll bump this if I find anything useful.",
    "1259269": "This would be cool to me too, +20h decoding process every time I want to submit.",
    "1260745": "Blockwise Parallel Decoding for Deep Autoregressive Models\nhttps://www.youtube.com/watch?v=3Tqp_B2G6u0\n\n![](https://user-images.githubusercontent.com/7529838/48524307-388f9200-e8c3-11e8-8f3b-0fa19cc947b4.png)\nhttps://github.com/kweonwooj/papers/issues/116",
    "1260769": "i figure out what to do:\n- since non-auto aggressive perform worse in accuracy, i would not use it for submission\n- but i can use it for hyperparameter searches and experiments. they provide lower bound for autoagressive models",
    "1260794": "![](https://i.ibb.co/w6HChT8/Selection-025.png)",
    "1260806": "does this affect the training process?",
    "1260811": "the paper and youtube describe the training process in details.\n\nbut basically\n1. you train your transformer as usual (i.e. the one by one decoder)\n2. then you add layers and build on top of the decoder (freezing it) so that you can decode one to multiple outputs.\n3. if you want, you can finetune end 2 end later",
    "1260863": "Thank you, I'll have a look!",
    "1280115": "unlike language translation or image captioning, there this only one inchi string given an input image.\nHence i think we can directly predict the whole inchi string instead of decoding one by one.\n\ni wonder if anyone has done this experiment?",
    "1280285": "Holy moly. \n\nThat’s all.",
    "1280369": "You can check TTS papers, newer approaches predict length first and then fillout i.e. FastSpeech",
    "1283969": "this is what i long suspected and i think it is very much applicable to this bms compeition\n\nDEEP ENCODER, SHALLOW DECODER: REEVALUATING NON-AUTOREGRESSIVE MACHINE TRANSLATION -ICRL 2021\n\n\"Our extensive experiments show that given a sufficiently deep encoder, a single-layer autoregressive decoder can substantially outperform strong non-autoregressive models with comparable inference speed\"\n\ntake home message:\nmake encoder deep\nmake decoder shallow and fast"
  },
  "source": "meta"
}