{
  "id": 438119,
  "title": "LB [0.444] - How to improve your score - Optimizing the Decoding Parameters with Optuna",
  "url": "/competitions/bengaliai-speech/discussion/438119",
  "author_name": "Sinan Calisir",
  "post_date": "2023-09-09T15:41:36.622000",
  "votes": 10,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n<p>I would like to show you how to improve your scores further by optimizing the decoding parameters with Optuna. </p>\n<blockquote>\n  <p>TLDR;<br>\n  We find better parameters for decoding and the WER improves from .445 to .444 in LB and the CV improves to 0.2738 from 0.2773.</p>\n</blockquote>\n<p>It's being suggested by the authors of the <a href=\"https://github.com/kensho-technologies/pyctcdecode\" target=\"_blank\">pyctcdecode</a> developers that we should perform a parameter search because it can improve our results on specific tasks other than English such as ours.</p>\n<blockquote>\n  <p>(Note: pyctcdecode contains several free hyperparameters that can strongly influence error rate and wall time. Default values for these parameters were (merely) chosen in order to yield good performance for one particular use case. For best results, especially when working with languages other than English, users are encouraged to perform a hyperparameter optimization study on their own data.)</p>\n</blockquote>\n<p>So I gave it a try to find the best parameters on a small subset of the validation split (because of the time constraints we will only use 5k). Here is the list of decoding params that we can tune:</p>\n<pre><code>\n\nDEFAULT_ALPHA = \nDEFAULT_BETA = \nDEFAULT_UNK_LOGP_OFFSET = -\nDEFAULT_BEAM_WIDTH = \nDEFAULT_HOTWORD_WEIGHT = \nDEFAULT_PRUNE_LOGP = -\nDEFAULT_PRUNE_BEAMS = \nDEFAULT_MIN_TOKEN_LOGP = -\nDEFAULT_SCORE_LM_BOUNDARY = \n\n\nAVG_TOKEN_LEN =   \nMIN_TOKEN_CLIP_P =   \nLOG_BASE_CHANGE_FACTOR =  / math.log10(math.e)  \n</code></pre>\n<p>We calculate the base WER score of the public model which is 0.2773. Then, we define a small search space like this:</p>\n<pre><code> ():\n    \n    alpha = trial.suggest_float(, , )\n    beta = trial.suggest_float(, , )\n    beam_width = trial.suggest_categorical(, [, , ])\n    gts = valid[].values.tolist()\n    decode_params = {\n        : alpha,\n        : beta,\n        : beam_width\n    }\n    preds = decode(logits, params=decode_params, pp=)\n    wer_score = score(gts, preds)\n     wer_score\n</code></pre>\n<p>After completing the search, the best parameter found by the optuna is as follows:</p>\n<p>{'alpha': 0.3802723523729998, 'beta': 0.053996879617918436, 'beam_width': 768} compared to what we had in the public notebook which was {\"beam_width\": 512} with a score 0.2738.</p>\n<p>If we submit the inference notebook with these parameters it gives a small boost to the LB. You can find the additional details in this <a href=\"https://www.kaggle.com/snnclsr/0-444-optimize-decoding-parameters-with-optuna\" target=\"_blank\">notebook</a>. Upvotes are really appreciated if you find this work useful. </p>",
  "messages": [
    {
      "id": 2430809,
      "postDate": "2023-09-09T15:41:36.623Z",
      "content": "<p>Hi everyone,</p>\n<p>I would like to show you how to improve your scores further by optimizing the decoding parameters with Optuna. </p>\n<blockquote>\n  <p>TLDR;<br>\n  We find better parameters for decoding and the WER improves from .445 to .444 in LB and the CV improves to 0.2738 from 0.2773.</p>\n</blockquote>\n<p>It's being suggested by the authors of the <a href=\"https://github.com/kensho-technologies/pyctcdecode\" target=\"_blank\">pyctcdecode</a> developers that we should perform a parameter search because it can improve our results on specific tasks other than English such as ours.</p>\n<blockquote>\n  <p>(Note: pyctcdecode contains several free hyperparameters that can strongly influence error rate and wall time. Default values for these parameters were (merely) chosen in order to yield good performance for one particular use case. For best results, especially when working with languages other than English, users are encouraged to perform a hyperparameter optimization study on their own data.)</p>\n</blockquote>\n<p>So I gave it a try to find the best parameters on a small subset of the validation split (because of the time constraints we will only use 5k). Here is the list of decoding params that we can tune:</p>\n<pre><code>\n\nDEFAULT_ALPHA = \nDEFAULT_BETA = \nDEFAULT_UNK_LOGP_OFFSET = -\nDEFAULT_BEAM_WIDTH = \nDEFAULT_HOTWORD_WEIGHT = \nDEFAULT_PRUNE_LOGP = -\nDEFAULT_PRUNE_BEAMS = \nDEFAULT_MIN_TOKEN_LOGP = -\nDEFAULT_SCORE_LM_BOUNDARY = \n\n\nAVG_TOKEN_LEN =   \nMIN_TOKEN_CLIP_P =   \nLOG_BASE_CHANGE_FACTOR =  / math.log10(math.e)  \n</code></pre>\n<p>We calculate the base WER score of the public model which is 0.2773. Then, we define a small search space like this:</p>\n<pre><code> ():\n    \n    alpha = trial.suggest_float(, , )\n    beta = trial.suggest_float(, , )\n    beam_width = trial.suggest_categorical(, [, , ])\n    gts = valid[].values.tolist()\n    decode_params = {\n        : alpha,\n        : beta,\n        : beam_width\n    }\n    preds = decode(logits, params=decode_params, pp=)\n    wer_score = score(gts, preds)\n     wer_score\n</code></pre>\n<p>After completing the search, the best parameter found by the optuna is as follows:</p>\n<p>{'alpha': 0.3802723523729998, 'beta': 0.053996879617918436, 'beam_width': 768} compared to what we had in the public notebook which was {\"beam_width\": 512} with a score 0.2738.</p>\n<p>If we submit the inference notebook with these parameters it gives a small boost to the LB. You can find the additional details in this <a href=\"https://www.kaggle.com/snnclsr/0-444-optimize-decoding-parameters-with-optuna\" target=\"_blank\">notebook</a>. Upvotes are really appreciated if you find this work useful. </p>",
      "rawMarkdown": "Hi everyone,\n\nI would like to show you how to improve your scores further by optimizing the decoding parameters with Optuna. \n\n> TLDR;\nWe find better parameters for decoding and the WER improves from .445 to .444 in LB and the CV improves to 0.2738 from 0.2773.\n\nIt's being suggested by the authors of the [pyctcdecode](https://github.com/kensho-technologies/pyctcdecode) developers that we should perform a parameter search because it can improve our results on specific tasks other than English such as ours.\n\n> (Note: pyctcdecode contains several free hyperparameters that can strongly influence error rate and wall time. Default values for these parameters were (merely) chosen in order to yield good performance for one particular use case. For best results, especially when working with languages other than English, users are encouraged to perform a hyperparameter optimization study on their own data.)\n\nSo I gave it a try to find the best parameters on a small subset of the validation split (because of the time constraints we will only use 5k). Here is the list of decoding params that we can tune:\n```python\n# from: https://github.com/kensho-technologies/pyctcdecode/blob/main/pyctcdecode/constants.py\n# default parameters for decoding (can be modified)\nDEFAULT_ALPHA = 0.5\nDEFAULT_BETA = 1.5\nDEFAULT_UNK_LOGP_OFFSET = -10.0\nDEFAULT_BEAM_WIDTH = 100\nDEFAULT_HOTWORD_WEIGHT = 10.0\nDEFAULT_PRUNE_LOGP = -10.0\nDEFAULT_PRUNE_BEAMS = False\nDEFAULT_MIN_TOKEN_LOGP = -5.0\nDEFAULT_SCORE_LM_BOUNDARY = True\n\n# other constants for decoding\nAVG_TOKEN_LEN = 6  # average number of characters expected per token (used for UNK scoring)\nMIN_TOKEN_CLIP_P = 1e-15  # clipping to avoid underflow in case of malformed logit input\nLOG_BASE_CHANGE_FACTOR = 1.0 / math.log10(math.e)  # kenlm returns base10 but we like natural\n```\n\nWe calculate the base WER score of the public model which is 0.2773. Then, we define a small search space like this:\n```python\ndef objective(trial):\n    \"\"\"\n    alpha: weight for language model during shallow fusion\n    beta: weight for length score adjustment of during scoring\n    unk_score_offset: amount of log score offset for unknown tokens\n    lm_score_boundary: whether to have kenlm respect boundaries when scoring\n    \"\"\"\n    alpha = trial.suggest_float(\"alpha\", 0.0, 2.0)\n    beta = trial.suggest_float(\"beta\", 0.0, 2.0)\n    beam_width = trial.suggest_categorical(\"beam_width\", [256, 512, 768])\n    gts = valid[\"sentence\"].values.tolist()\n    decode_params = {\n        \"alpha\": alpha,\n        \"beta\": beta,\n        \"beam_width\": beam_width\n    }\n    preds = decode(logits, params=decode_params, pp=True)\n    wer_score = score(gts, preds)\n    return wer_score\n```\nAfter completing the search, the best parameter found by the optuna is as follows:\n\n{'alpha': 0.3802723523729998, 'beta': 0.053996879617918436, 'beam_width': 768} compared to what we had in the public notebook which was {\"beam_width\": 512} with a score 0.2738.\n\nIf we submit the inference notebook with these parameters it gives a small boost to the LB. You can find the additional details in this [notebook](https://www.kaggle.com/snnclsr/0-444-optimize-decoding-parameters-with-optuna). Upvotes are really appreciated if you find this work useful. ",
      "votes": 10
    },
    {
      "id": 2431385,
      "postDate": "2023-09-10T04:55:27.460Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2431385,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-09-10T04:55:27.460000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2430809": "Hi everyone,\n\nI would like to show you how to improve your scores further by optimizing the decoding parameters with Optuna. \n\n> TLDR;\nWe find better parameters for decoding and the WER improves from .445 to .444 in LB and the CV improves to 0.2738 from 0.2773.\n\nIt's being suggested by the authors of the [pyctcdecode](https://github.com/kensho-technologies/pyctcdecode) developers that we should perform a parameter search because it can improve our results on specific tasks other than English such as ours.\n\n> (Note: pyctcdecode contains several free hyperparameters that can strongly influence error rate and wall time. Default values for these parameters were (merely) chosen in order to yield good performance for one particular use case. For best results, especially when working with languages other than English, users are encouraged to perform a hyperparameter optimization study on their own data.)\n\nSo I gave it a try to find the best parameters on a small subset of the validation split (because of the time constraints we will only use 5k). Here is the list of decoding params that we can tune:\n```python\n# from: https://github.com/kensho-technologies/pyctcdecode/blob/main/pyctcdecode/constants.py\n# default parameters for decoding (can be modified)\nDEFAULT_ALPHA = 0.5\nDEFAULT_BETA = 1.5\nDEFAULT_UNK_LOGP_OFFSET = -10.0\nDEFAULT_BEAM_WIDTH = 100\nDEFAULT_HOTWORD_WEIGHT = 10.0\nDEFAULT_PRUNE_LOGP = -10.0\nDEFAULT_PRUNE_BEAMS = False\nDEFAULT_MIN_TOKEN_LOGP = -5.0\nDEFAULT_SCORE_LM_BOUNDARY = True\n\n# other constants for decoding\nAVG_TOKEN_LEN = 6  # average number of characters expected per token (used for UNK scoring)\nMIN_TOKEN_CLIP_P = 1e-15  # clipping to avoid underflow in case of malformed logit input\nLOG_BASE_CHANGE_FACTOR = 1.0 / math.log10(math.e)  # kenlm returns base10 but we like natural\n```\n\nWe calculate the base WER score of the public model which is 0.2773. Then, we define a small search space like this:\n```python\ndef objective(trial):\n    \"\"\"\n    alpha: weight for language model during shallow fusion\n    beta: weight for length score adjustment of during scoring\n    unk_score_offset: amount of log score offset for unknown tokens\n    lm_score_boundary: whether to have kenlm respect boundaries when scoring\n    \"\"\"\n    alpha = trial.suggest_float(\"alpha\", 0.0, 2.0)\n    beta = trial.suggest_float(\"beta\", 0.0, 2.0)\n    beam_width = trial.suggest_categorical(\"beam_width\", [256, 512, 768])\n    gts = valid[\"sentence\"].values.tolist()\n    decode_params = {\n        \"alpha\": alpha,\n        \"beta\": beta,\n        \"beam_width\": beam_width\n    }\n    preds = decode(logits, params=decode_params, pp=True)\n    wer_score = score(gts, preds)\n    return wer_score\n```\nAfter completing the search, the best parameter found by the optuna is as follows:\n\n{'alpha': 0.3802723523729998, 'beta': 0.053996879617918436, 'beam_width': 768} compared to what we had in the public notebook which was {\"beam_width\": 512} with a score 0.2738.\n\nIf we submit the inference notebook with these parameters it gives a small boost to the LB. You can find the additional details in this [notebook](https://www.kaggle.com/snnclsr/0-444-optimize-decoding-parameters-with-optuna). Upvotes are really appreciated if you find this work useful. ",
    "2431385": ""
  }
}