{
  "id": 231065,
  "title": "Batched beam search implementation",
  "url": "/competitions/bms-molecular-translation/discussion/231065",
  "author_name": "",
  "post_date": "2021-04-06T20:06:49.927582500Z",
  "votes": 76,
  "comment_count": 21,
  "views": 0,
  "content": "<p>I have updated the Y.Nakamas public notebook with a batched beam search decoder: <a href=\"https://www.kaggle.com/tugstugi/batched-beam-search-inference\" target=\"_blank\">https://www.kaggle.com/tugstugi/batched-beam-search-inference</a></p>\n<p>Depending on the model, it can bring up to 1.X improvement on LB. The original code is taken from <a href=\"https://github.com/IBM/pytorch-seq2seq/blob/master/seq2seq/models/TopKDecoder.py\" target=\"_blank\">https://github.com/IBM/pytorch-seq2seq/blob/master/seq2seq/models/TopKDecoder.py</a> (The _inflate function seems to be not working, so I had to replace it with torch.repeat_interleave)</p>",
  "messages": [
    {
      "id": "1265339",
      "postDate": "04/06/2021 20:06:49",
      "content": "<p>I have updated the Y.Nakamas public notebook with a batched beam search decoder: <a href=\"https://www.kaggle.com/tugstugi/batched-beam-search-inference\" target=\"_blank\">https://www.kaggle.com/tugstugi/batched-beam-search-inference</a></p>\n<p>Depending on the model, it can bring up to 1.X improvement on LB. The original code is taken from <a href=\"https://github.com/IBM/pytorch-seq2seq/blob/master/seq2seq/models/TopKDecoder.py\" target=\"_blank\">https://github.com/IBM/pytorch-seq2seq/blob/master/seq2seq/models/TopKDecoder.py</a> (The _inflate function seems to be not working, so I had to replace it with torch.repeat_interleave)</p>",
      "rawMarkdown": "I have updated the Y.Nakamas public notebook with a batched beam search decoder: https://www.kaggle.com/tugstugi/batched-beam-search-inference\n\nDepending on the model, it can bring up to 1.X improvement on LB. The original code is taken from https://github.com/IBM/pytorch-seq2seq/blob/master/seq2seq/models/TopKDecoder.py (The _inflate function seems to be not working, so I had to replace it with torch.repeat_interleave)",
      "votes": null
    },
    {
      "id": "1265341",
      "postDate": "04/06/2021 20:10:20",
      "content": "<p><a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> what is the LB score of your provided trained weight? Is LB20.3 from a model from a private data input source?</p>",
      "rawMarkdown": "yasufuminakama what is the LB score of your provided trained weight? Is LB20.3 from a model from a private data input source?",
      "votes": null
    },
    {
      "id": "1265462",
      "postDate": "04/06/2021 23:00:06",
      "content": "<p>I think it was 2 epoch training with instead of 1</p>",
      "rawMarkdown": "I think it was 2 epoch training with instead of 1",
      "votes": null
    },
    {
      "id": "1265477",
      "postDate": "04/06/2021 23:36:06",
      "content": "<p>Thanks!</p>\n<p>beam search is one of the must todo task in this competition</p>",
      "rawMarkdown": "Thanks!\n\nbeam search is one of the must todo task in this competition",
      "votes": null
    },
    {
      "id": "1265539",
      "postDate": "04/07/2021 02:11:09",
      "content": "<p>Thanks, great work!</p>",
      "rawMarkdown": "Thanks, great work!",
      "votes": null
    },
    {
      "id": "1265587",
      "postDate": "04/07/2021 03:42:31",
      "content": "<p>LB20.34 is using 2 epoch trained weight, so I haven't checked LB of 1 epoch trained weight.</p>",
      "rawMarkdown": "LB20.34 is using 2 epoch trained weight, so I haven't checked LB of 1 epoch trained weight.",
      "votes": null
    },
    {
      "id": "1265917",
      "postDate": "04/07/2021 10:27:55",
      "content": "<p>Thanks <br>\nMy beam search implementation gave me +1.xx boost on CV . But it's terribly  slow (not batch wise)</p>\n<p>That's why I can't submit. </p>",
      "rawMarkdown": "Thanks \nMy beam search implementation gave me +1.xx boost on CV . But it's terribly  slow (not batch wise)\n\nThat's why I can't submit.",
      "votes": null
    },
    {
      "id": "1265976",
      "postDate": "04/07/2021 11:39:28",
      "content": "<p>same problem for me too!</p>\n<p>how i wish that a group of us can pay a professional cuda SW engineer to code for us. there are 419 participiants … each of us fork out $20 …</p>",
      "rawMarkdown": "same problem for me too!\n\nhow i wish that a group of us can pay a professional cuda SW engineer to code for us. there are 419 participiants ... each of us fork out $20 ...",
      "votes": null
    },
    {
      "id": "1266468",
      "postDate": "04/07/2021 19:23:08",
      "content": "<p>I applied beam search to our best model .. score went from <code>1.34 --&gt; 1.35</code> =) .. I will update if we could improve our score using beam search and some analysis =) Thank you for sharing. </p>",
      "rawMarkdown": "I applied beam search to our best model .. score went from `1.34 --> 1.35` =) .. I will update if we could improve our score using beam search and some analysis =) Thank you for sharing.",
      "votes": null
    },
    {
      "id": "1266478",
      "postDate": "04/07/2021 19:44:46",
      "content": "<p>TPU top-k beam search: <a href=\"https://github.com/tensorflow/tensor2tensor/blob/master/tensor2tensor/utils/beam_search.py\" target=\"_blank\">https://github.com/tensorflow/tensor2tensor/blob/master/tensor2tensor/utils/beam_search.py</a><br>\nfor tf user only</p>\n<p>the ability to watch the live search is amazing</p>\n<pre><code>def beam_search(symbols_to_logits_fn,\n                initial_ids,\n                beam_size,\n                decode_length,\n                vocab_size,\n                alpha,\n                states=None,\n                eos_id=EOS_ID,\n                stop_early=True,\n                use_tpu=False,\n                use_top_k_with_unique=True):\n  \"\"\"Beam search with length penalties.\n  Requires a function that can take the currently decoded symbols and return\n  the logits for the next symbol. The implementation is inspired by\n  https://arxiv.org/abs/1609.08144.\n  When running, the beam search steps can be visualized by using tfdbg to watch\n  the operations generating the output ids for each beam step.  These operations\n  have the pattern:\n    (alive|finished)_topk_(seq,scores)\n</code></pre>",
      "rawMarkdown": "TPU top-k beam search: https://github.com/tensorflow/tensor2tensor/blob/master/tensor2tensor/utils/beam_search.py\nfor tf user only\n\nthe ability to watch the live search is amazing\n\n```\ndef beam_search(symbols_to_logits_fn,\n                initial_ids,\n                beam_size,\n                decode_length,\n                vocab_size,\n                alpha,\n                states=None,\n                eos_id=EOS_ID,\n                stop_early=True,\n                use_tpu=False,\n                use_top_k_with_unique=True):\n  \"\"\"Beam search with length penalties.\n  Requires a function that can take the currently decoded symbols and return\n  the logits for the next symbol. The implementation is inspired by\n  https://arxiv.org/abs/1609.08144.\n  When running, the beam search steps can be visualized by using tfdbg to watch\n  the operations generating the output ids for each beam step.  These operations\n  have the pattern:\n    (alive|finished)_topk_(seq,scores)\n```",
      "votes": null
    },
    {
      "id": "1266487",
      "postDate": "04/07/2021 19:50:52",
      "content": "<p>you can measure the topK accuracy for each token prediction (like imagenet) during training.<br>\nthis gives you an idea of how far the truth label is from argmax.</p>\n<p>with lb score &lt;2, i think your cross-entropy loss is in the range of 0.010 to 0.015, this is about 99% accurate. Hence the effect of top-k may diminish</p>",
      "rawMarkdown": "you can measure the topK accuracy for each token prediction (like imagenet) during training.\nthis gives you an idea of how far the truth label is from argmax.\n\n\nwith lb score <2, i think your cross-entropy loss is in the range of 0.010 to 0.015, this is about 99% accurate. Hence the effect of top-k may diminish",
      "votes": null
    },
    {
      "id": "1266492",
      "postDate": "04/07/2021 19:54:19",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> You are correct, with my latest model the gain is now only around 0.2… So better the model, the effect is far less.</p>",
      "rawMarkdown": "hengck23 You are correct, with my latest model the gain is now only around 0.2... So better the model, the effect is far less.",
      "votes": null
    },
    {
      "id": "1266497",
      "postDate": "04/07/2021 19:59:46",
      "content": "<p><a href=\"https://www.kaggle.com/drhabib\" target=\"_blank\">@drhabib</a> Have you measured the metric for your train fold?</p>",
      "rawMarkdown": "drhabib Have you measured the metric for your train fold?",
      "votes": null
    },
    {
      "id": "1266502",
      "postDate": "04/07/2021 20:03:00",
      "content": "<p>Nope .. only for <code>valid</code> and its <code>1.29</code>. </p>",
      "rawMarkdown": "Nope .. only for `valid` and its `1.29`.",
      "votes": null
    },
    {
      "id": "1266503",
      "postDate": "04/07/2021 20:03:27",
      "content": "<p>you can try softening the softmax output during training. e.g. using label smoothing.<br>\ntop-K works better if the argmax don't over dominate the other values.</p>\n<p>since we are training very long epochs, the argmax values grow very big, resulting in only one dominant path at beam search. you can verity this by:</p>\n<ol>\n<li>check the beam path values for strong models (little improvement)</li>\n<li>check the beam path values for weak models (more improvement)</li>\n</ol>\n<p>it is like ensemble. we can use p**0.5 in average. </p>\n<p>for max path, there may exist a good probability reshaping function.</p>",
      "rawMarkdown": "you can try softening the softmax output during training. e.g. using label smoothing.\ntop-K works better if the argmax don't over dominate the other values.\n\nsince we are training very long epochs, the argmax values grow very big, resulting in only one dominant path at beam search. you can verity this by:\n\n1.  check the beam path values for strong models (little improvement)\n2. check the beam path values for weak models (more improvement)\n\nit is like ensemble. we can use p**0.5 in average. \n\nfor max path, there may exist a good probability reshaping function.",
      "votes": null
    },
    {
      "id": "1266508",
      "postDate": "04/07/2021 20:07:18",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> yes, my next things to try are label smoothing and teacher forcing scheduling. But every experiment takes too long time :D</p>",
      "rawMarkdown": "hengck23 yes, my next things to try are label smoothing and teacher forcing scheduling. But every experiment takes too long time :D",
      "votes": null
    },
    {
      "id": "1266509",
      "postDate": "04/07/2021 20:08:13",
      "content": "<p>tensorRT beam search</p>\n<p><img src=\"https://developer-blogs.nvidia.com/wp-content/uploads/2018/06/pasted-image-0-16-768x341.png\" alt=\"\"></p>\n<p><a href=\"https://developer.nvidia.com/blog/tensorrt-4-accelerates-translation-speech-recommender/\" target=\"_blank\">https://developer.nvidia.com/blog/tensorrt-4-accelerates-translation-speech-recommender/</a><br>\n<a href=\"https://on-demand.gputechconf.com/gtc/2018/presentation/s8822-optimizing-nmt-with-tensorrt.pdf\" target=\"_blank\">https://on-demand.gputechconf.com/gtc/2018/presentation/s8822-optimizing-nmt-with-tensorrt.pdf</a></p>",
      "rawMarkdown": "tensorRT beam search\n\n![](https://developer-blogs.nvidia.com/wp-content/uploads/2018/06/pasted-image-0-16-768x341.png)\n\nhttps://developer.nvidia.com/blog/tensorrt-4-accelerates-translation-speech-recommender/\nhttps://on-demand.gputechconf.com/gtc/2018/presentation/s8822-optimizing-nmt-with-tensorrt.pdf",
      "votes": null
    },
    {
      "id": "1266700",
      "postDate": "04/08/2021 02:19:06",
      "content": "<p>on a side note, i wonder did anyone  train with soft label token as input before?</p>\n<p><img src=\"https://i.ibb.co/rsKMqRr/Selection-051.png\" alt=\"\"></p>",
      "rawMarkdown": "on a side note, i wonder did anyone  train with soft label token as input before?\n\n![](https://i.ibb.co/rsKMqRr/Selection-051.png)",
      "votes": null
    },
    {
      "id": "1267524",
      "postDate": "04/08/2021 15:15:20",
      "content": "<p>Thanks alot for sharing :) .</p>",
      "rawMarkdown": "Thanks alot for sharing :) .",
      "votes": null
    },
    {
      "id": "1268709",
      "postDate": "04/09/2021 17:06:27",
      "content": "<p>how to analyze beam search problem:</p>\n<p><a href=\"https://www.aclweb.org/anthology/2020.findings-emnlp.276.pdf\" target=\"_blank\">https://www.aclweb.org/anthology/2020.findings-emnlp.276.pdf</a><br>\nOn Long-Tailed Phenomena in Neural Machine Translation</p>\n<p>\"Beam Search Analysis To better establish the link<br>\nbetween token level classification and beam search<br>\ninference, we study the distribution of positional<br>\nscores, i.e. the probabilities selected during each<br>\nstep of decoding, for the top hypothesis finally selected during beam search. \"</p>\n<p>\"These observations show that the<br>\napproximate inference procedure of beam-search<br>\nrelies significantly on low confidence predictions.\"</p>",
      "rawMarkdown": "how to analyze beam search problem:\n\nhttps://www.aclweb.org/anthology/2020.findings-emnlp.276.pdf\nOn Long-Tailed Phenomena in Neural Machine Translation\n\n\"Beam Search Analysis To better establish the link\nbetween token level classification and beam search\ninference, we study the distribution of positional\nscores, i.e. the probabilities selected during each\nstep of decoding, for the top hypothesis finally selected during beam search. \"\n\n\"These observations show that the\napproximate inference procedure of beam-search\nrelies significantly on low confidence predictions.\"",
      "votes": null
    },
    {
      "id": "1428177",
      "postDate": "08/03/2021 09:11:44",
      "content": "<p>Here is <a href=\"https://github.com/nofreewill42/bms/blob/master/model_architecture/beam_search.py\" target=\"_blank\">my implementation of beam search</a>.<br>\nInitialization: you give the models and different weights if you like.<br>\nInference: you can give weights for each model on the fly (I used this for giving more weight to models that were trained on more similar \"ratioed\" images to the one we do inference on)</p>\n<p>A model needs to have the functions</p>\n<ul>\n<li>encoder_output(imgs_tensor) -&gt; enc_out: Tensor of shape [bs,#tokens,d_model]</li>\n<li>decoder_output(enc_out, cache, prev_tokens) -&gt; bpe_probs, cache</li>\n</ul>\n<p>The implementation of my whole solution got out of hands as we approached the end of the competition, but you can see that too, if interested.</p>",
      "rawMarkdown": "Here is [my implementation of beam search](https://github.com/nofreewill42/bms/blob/master/model_architecture/beam_search.py).\nInitialization: you give the models and different weights if you like.\nInference: you can give weights for each model on the fly (I used this for giving more weight to models that were trained on more similar \"ratioed\" images to the one we do inference on)\n\nA model needs to have the functions\n - encoder_output(imgs_tensor) -> enc_out: Tensor of shape [bs,#tokens,d_model]\n - decoder_output(enc_out, cache, prev_tokens) -> bpe_probs, cache\n\nThe implementation of my whole solution got out of hands as we approached the end of the competition, but you can see that too, if interested.",
      "votes": null
    },
    {
      "id": "1447900",
      "postDate": "08/04/2021 16:13:11",
      "content": "<p>Thanks!, Nice work</p>",
      "rawMarkdown": "Thanks!, Nice work",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1265341,
      "author_name": "tugstugi",
      "author_url": "",
      "post_date": "04/06/2021 20:10:20",
      "content": "<p><a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> what is the LB score of your provided trained weight? Is LB20.3 from a model from a private data input source?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1265462,
          "author_name": "datafan07",
          "author_url": "",
          "post_date": "04/06/2021 23:00:06",
          "content": "<p>I think it was 2 epoch training with instead of 1</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1265587,
          "author_name": "yasufuminakama",
          "author_url": "",
          "post_date": "04/07/2021 03:42:31",
          "content": "<p>LB20.34 is using 2 epoch trained weight, so I haven't checked LB of 1 epoch trained weight.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1265477,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "04/06/2021 23:36:06",
      "content": "<p>Thanks!</p>\n<p>beam search is one of the must todo task in this competition</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1265539,
      "author_name": "atsunorifujita",
      "author_url": "",
      "post_date": "04/07/2021 02:11:09",
      "content": "<p>Thanks, great work!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1265917,
      "author_name": "serigne",
      "author_url": "",
      "post_date": "04/07/2021 10:27:55",
      "content": "<p>Thanks <br>\nMy beam search implementation gave me +1.xx boost on CV . But it's terribly  slow (not batch wise)</p>\n<p>That's why I can't submit. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1265976,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "04/07/2021 11:39:28",
          "content": "<p>same problem for me too!</p>\n<p>how i wish that a group of us can pay a professional cuda SW engineer to code for us. there are 419 participiants … each of us fork out $20 …</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1266468,
      "author_name": "drhabib",
      "author_url": "",
      "post_date": "04/07/2021 19:23:08",
      "content": "<p>I applied beam search to our best model .. score went from <code>1.34 --&gt; 1.35</code> =) .. I will update if we could improve our score using beam search and some analysis =) Thank you for sharing. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1266487,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "04/07/2021 19:50:52",
          "content": "<p>you can measure the topK accuracy for each token prediction (like imagenet) during training.<br>\nthis gives you an idea of how far the truth label is from argmax.</p>\n<p>with lb score &lt;2, i think your cross-entropy loss is in the range of 0.010 to 0.015, this is about 99% accurate. Hence the effect of top-k may diminish</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1266492,
          "author_name": "tugstugi",
          "author_url": "",
          "post_date": "04/07/2021 19:54:19",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> You are correct, with my latest model the gain is now only around 0.2… So better the model, the effect is far less.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1266497,
          "author_name": "tugstugi",
          "author_url": "",
          "post_date": "04/07/2021 19:59:46",
          "content": "<p><a href=\"https://www.kaggle.com/drhabib\" target=\"_blank\">@drhabib</a> Have you measured the metric for your train fold?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1266502,
          "author_name": "drhabib",
          "author_url": "",
          "post_date": "04/07/2021 20:03:00",
          "content": "<p>Nope .. only for <code>valid</code> and its <code>1.29</code>. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1266503,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "04/07/2021 20:03:27",
          "content": "<p>you can try softening the softmax output during training. e.g. using label smoothing.<br>\ntop-K works better if the argmax don't over dominate the other values.</p>\n<p>since we are training very long epochs, the argmax values grow very big, resulting in only one dominant path at beam search. you can verity this by:</p>\n<ol>\n<li>check the beam path values for strong models (little improvement)</li>\n<li>check the beam path values for weak models (more improvement)</li>\n</ol>\n<p>it is like ensemble. we can use p**0.5 in average. </p>\n<p>for max path, there may exist a good probability reshaping function.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1266508,
          "author_name": "tugstugi",
          "author_url": "",
          "post_date": "04/07/2021 20:07:18",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> yes, my next things to try are label smoothing and teacher forcing scheduling. But every experiment takes too long time :D</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1266700,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "04/08/2021 02:19:06",
          "content": "<p>on a side note, i wonder did anyone  train with soft label token as input before?</p>\n<p><img src=\"https://i.ibb.co/rsKMqRr/Selection-051.png\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1266478,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "04/07/2021 19:44:46",
      "content": "<p>TPU top-k beam search: <a href=\"https://github.com/tensorflow/tensor2tensor/blob/master/tensor2tensor/utils/beam_search.py\" target=\"_blank\">https://github.com/tensorflow/tensor2tensor/blob/master/tensor2tensor/utils/beam_search.py</a><br>\nfor tf user only</p>\n<p>the ability to watch the live search is amazing</p>\n<pre><code>def beam_search(symbols_to_logits_fn,\n                initial_ids,\n                beam_size,\n                decode_length,\n                vocab_size,\n                alpha,\n                states=None,\n                eos_id=EOS_ID,\n                stop_early=True,\n                use_tpu=False,\n                use_top_k_with_unique=True):\n  \"\"\"Beam search with length penalties.\n  Requires a function that can take the currently decoded symbols and return\n  the logits for the next symbol. The implementation is inspired by\n  https://arxiv.org/abs/1609.08144.\n  When running, the beam search steps can be visualized by using tfdbg to watch\n  the operations generating the output ids for each beam step.  These operations\n  have the pattern:\n    (alive|finished)_topk_(seq,scores)\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 1266509,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "04/07/2021 20:08:13",
          "content": "<p>tensorRT beam search</p>\n<p><img src=\"https://developer-blogs.nvidia.com/wp-content/uploads/2018/06/pasted-image-0-16-768x341.png\" alt=\"\"></p>\n<p><a href=\"https://developer.nvidia.com/blog/tensorrt-4-accelerates-translation-speech-recommender/\" target=\"_blank\">https://developer.nvidia.com/blog/tensorrt-4-accelerates-translation-speech-recommender/</a><br>\n<a href=\"https://on-demand.gputechconf.com/gtc/2018/presentation/s8822-optimizing-nmt-with-tensorrt.pdf\" target=\"_blank\">https://on-demand.gputechconf.com/gtc/2018/presentation/s8822-optimizing-nmt-with-tensorrt.pdf</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1267524,
      "author_name": "sayedathar11",
      "author_url": "",
      "post_date": "04/08/2021 15:15:20",
      "content": "<p>Thanks alot for sharing :) .</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1268709,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "04/09/2021 17:06:27",
      "content": "<p>how to analyze beam search problem:</p>\n<p><a href=\"https://www.aclweb.org/anthology/2020.findings-emnlp.276.pdf\" target=\"_blank\">https://www.aclweb.org/anthology/2020.findings-emnlp.276.pdf</a><br>\nOn Long-Tailed Phenomena in Neural Machine Translation</p>\n<p>\"Beam Search Analysis To better establish the link<br>\nbetween token level classification and beam search<br>\ninference, we study the distribution of positional<br>\nscores, i.e. the probabilities selected during each<br>\nstep of decoding, for the top hypothesis finally selected during beam search. \"</p>\n<p>\"These observations show that the<br>\napproximate inference procedure of beam-search<br>\nrelies significantly on low confidence predictions.\"</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1428177,
      "author_name": "nofreewill",
      "author_url": "",
      "post_date": "08/03/2021 09:11:44",
      "content": "<p>Here is <a href=\"https://github.com/nofreewill42/bms/blob/master/model_architecture/beam_search.py\" target=\"_blank\">my implementation of beam search</a>.<br>\nInitialization: you give the models and different weights if you like.<br>\nInference: you can give weights for each model on the fly (I used this for giving more weight to models that were trained on more similar \"ratioed\" images to the one we do inference on)</p>\n<p>A model needs to have the functions</p>\n<ul>\n<li>encoder_output(imgs_tensor) -&gt; enc_out: Tensor of shape [bs,#tokens,d_model]</li>\n<li>decoder_output(enc_out, cache, prev_tokens) -&gt; bpe_probs, cache</li>\n</ul>\n<p>The implementation of my whole solution got out of hands as we approached the end of the competition, but you can see that too, if interested.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1447900,
      "author_name": "rajavarman",
      "author_url": "",
      "post_date": "08/04/2021 16:13:11",
      "content": "<p>Thanks!, Nice work</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1265339": "I have updated the Y.Nakamas public notebook with a batched beam search decoder: https://www.kaggle.com/tugstugi/batched-beam-search-inference\n\nDepending on the model, it can bring up to 1.X improvement on LB. The original code is taken from https://github.com/IBM/pytorch-seq2seq/blob/master/seq2seq/models/TopKDecoder.py (The _inflate function seems to be not working, so I had to replace it with torch.repeat_interleave)",
    "1265341": "yasufuminakama what is the LB score of your provided trained weight? Is LB20.3 from a model from a private data input source?",
    "1265462": "I think it was 2 epoch training with instead of 1",
    "1265477": "Thanks!\n\nbeam search is one of the must todo task in this competition",
    "1265539": "Thanks, great work!",
    "1265587": "LB20.34 is using 2 epoch trained weight, so I haven't checked LB of 1 epoch trained weight.",
    "1265917": "Thanks \nMy beam search implementation gave me +1.xx boost on CV . But it's terribly  slow (not batch wise)\n\nThat's why I can't submit.",
    "1265976": "same problem for me too!\n\nhow i wish that a group of us can pay a professional cuda SW engineer to code for us. there are 419 participiants ... each of us fork out $20 ...",
    "1266468": "I applied beam search to our best model .. score went from `1.34 --> 1.35` =) .. I will update if we could improve our score using beam search and some analysis =) Thank you for sharing.",
    "1266478": "TPU top-k beam search: https://github.com/tensorflow/tensor2tensor/blob/master/tensor2tensor/utils/beam_search.py\nfor tf user only\n\nthe ability to watch the live search is amazing\n\n```\ndef beam_search(symbols_to_logits_fn,\n                initial_ids,\n                beam_size,\n                decode_length,\n                vocab_size,\n                alpha,\n                states=None,\n                eos_id=EOS_ID,\n                stop_early=True,\n                use_tpu=False,\n                use_top_k_with_unique=True):\n  \"\"\"Beam search with length penalties.\n  Requires a function that can take the currently decoded symbols and return\n  the logits for the next symbol. The implementation is inspired by\n  https://arxiv.org/abs/1609.08144.\n  When running, the beam search steps can be visualized by using tfdbg to watch\n  the operations generating the output ids for each beam step.  These operations\n  have the pattern:\n    (alive|finished)_topk_(seq,scores)\n```",
    "1266487": "you can measure the topK accuracy for each token prediction (like imagenet) during training.\nthis gives you an idea of how far the truth label is from argmax.\n\n\nwith lb score <2, i think your cross-entropy loss is in the range of 0.010 to 0.015, this is about 99% accurate. Hence the effect of top-k may diminish",
    "1266492": "hengck23 You are correct, with my latest model the gain is now only around 0.2... So better the model, the effect is far less.",
    "1266497": "drhabib Have you measured the metric for your train fold?",
    "1266502": "Nope .. only for `valid` and its `1.29`.",
    "1266503": "you can try softening the softmax output during training. e.g. using label smoothing.\ntop-K works better if the argmax don't over dominate the other values.\n\nsince we are training very long epochs, the argmax values grow very big, resulting in only one dominant path at beam search. you can verity this by:\n\n1.  check the beam path values for strong models (little improvement)\n2. check the beam path values for weak models (more improvement)\n\nit is like ensemble. we can use p**0.5 in average. \n\nfor max path, there may exist a good probability reshaping function.",
    "1266508": "hengck23 yes, my next things to try are label smoothing and teacher forcing scheduling. But every experiment takes too long time :D",
    "1266509": "tensorRT beam search\n\n![](https://developer-blogs.nvidia.com/wp-content/uploads/2018/06/pasted-image-0-16-768x341.png)\n\nhttps://developer.nvidia.com/blog/tensorrt-4-accelerates-translation-speech-recommender/\nhttps://on-demand.gputechconf.com/gtc/2018/presentation/s8822-optimizing-nmt-with-tensorrt.pdf",
    "1266700": "on a side note, i wonder did anyone  train with soft label token as input before?\n\n![](https://i.ibb.co/rsKMqRr/Selection-051.png)",
    "1267524": "Thanks alot for sharing :) .",
    "1268709": "how to analyze beam search problem:\n\nhttps://www.aclweb.org/anthology/2020.findings-emnlp.276.pdf\nOn Long-Tailed Phenomena in Neural Machine Translation\n\n\"Beam Search Analysis To better establish the link\nbetween token level classification and beam search\ninference, we study the distribution of positional\nscores, i.e. the probabilities selected during each\nstep of decoding, for the top hypothesis finally selected during beam search. \"\n\n\"These observations show that the\napproximate inference procedure of beam-search\nrelies significantly on low confidence predictions.\"",
    "1428177": "Here is [my implementation of beam search](https://github.com/nofreewill42/bms/blob/master/model_architecture/beam_search.py).\nInitialization: you give the models and different weights if you like.\nInference: you can give weights for each model on the fly (I used this for giving more weight to models that were trained on more similar \"ratioed\" images to the one we do inference on)\n\nA model needs to have the functions\n - encoder_output(imgs_tensor) -> enc_out: Tensor of shape [bs,#tokens,d_model]\n - decoder_output(enc_out, cache, prev_tokens) -> bpe_probs, cache\n\nThe implementation of my whole solution got out of hands as we approached the end of the competition, but you can see that too, if interested.",
    "1447900": "Thanks!, Nice work"
  },
  "source": "meta"
}