{
  "id": 56074,
  "title": "An LSTM CTC solution",
  "url": "/competitions/tensorflow-speech-recognition-challenge/writeups/huschen-an-lstm-ctc-solution",
  "author_name": "",
  "post_date": "2018-05-05T18:15:16.783Z",
  "votes": 6,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Most models in this competition are CNN based and learns to categorise words. </p>\n\n<p>This model is a Convolution Residual, backward LSTM network using Connectionist Temporal Classification (CTC) cost, written in TensorFlow. <a href=\"https://github.com/huschen/kaggle_speech_recognition\">Source code</a>. It is more of a characters(phonemes) learning model.</p>\n\n<h2>Performance (Private Score):</h2>\n\n<p>Single model, single training run: 0.89357</p>\n\n<p>Average over 5 folds (runs): 0.89874</p>\n\n<p>No test data is used in training.</p>\n\n<h2>Comparision with Word-Learning Models:</h2>\n\n<p>In this particular competition and data set, character learning model has its limitations:</p>\n\n<ul>\n<li><p>High false negative rate (i.e. the tendency to miss predict key words as unknown words). Simply because it has to get all the characters right to predict the correct word. An ambiguous 'off' sound might be recognised as 'oP' or 'Uf'. A 'seven' with an unheard accent might be predicted as 'sIven'. Averaging over different fold runs sometimes helps.</p></li>\n<li><p>More sensitive to incomplete word recordings (both in training and testing). Learning (maybe memorising) ‘LE’ ‘LEF’, 'EFT', etc. as 'LEFT' is harder for c-models than w-models, because c-models has to 'learn' the character that is not there in the sound.. It is quite a random guess in recognizing 'O' as 'go', 'off' or 'on' in the competition.</p></li>\n<li><p>Need prior knowledge of pronunciations and preprocessing, as there is no one-to-one mapping between characters and phonemes. But it is possible the models can learn the mapping (and memorizing exceptions) with larger training data.</p></li>\n<li><p>In the case of this particular Kaggle competition, the word-character (30-28) ratio is too small for c-models being efficient.</p></li>\n</ul>\n\n<p>But some advantages as well, possibly more suitable in practical uses:</p>\n\n<ul>\n<li><p>Generalisation. The c-model is able to learn unseen words, e.g. recognizing 'night' from learning 'NIne' and 'rIGHT', recognizing 'follow' as 'foow' (missing 'l' sound) from learning 'Four' 'dOWn' and 'nO'.</p></li>\n<li><p>Easier to scale up to larger vocabularies, with the increase of the number of LSTM hidden units, as well as the increase of the training time when learning longer words. Bi-Directional LSTM will be needed to capture more complex patterns.</p></li>\n<li><p>In practise, the property of high false negative (and low false positive) rate is likely a desired feature for key commands detecting applications. It is possible that the models will be customized to end users (with the balance of being general) so that the accents won't be a problem.</p></li>\n</ul>\n\n<p>Let me know what you think:-)</p>",
  "messages": [
    {
      "id": "323588",
      "postDate": "05/05/2018 16:05:11",
      "content": "<p>Most models in this competition are CNN based and learns to categorise words. </p>\n\n<p>This model is a Convolution Residual, backward LSTM network using Connectionist Temporal Classification (CTC) cost, written in TensorFlow. <a href=\"https://github.com/huschen/kaggle_speech_recognition\">Source code</a>. It is more of a characters(phonemes) learning model.</p>\n\n<h2>Performance (Private Score):</h2>\n\n<p>Single model, single training run: 0.89357</p>\n\n<p>Average over 5 folds (runs): 0.89874</p>\n\n<p>No test data is used in training.</p>\n\n<h2>Comparision with Word-Learning Models:</h2>\n\n<p>In this particular competition and data set, character learning model has its limitations:</p>\n\n<ul>\n<li><p>High false negative rate (i.e. the tendency to miss predict key words as unknown words). Simply because it has to get all the characters right to predict the correct word. An ambiguous 'off' sound might be recognised as 'oP' or 'Uf'. A 'seven' with an unheard accent might be predicted as 'sIven'. Averaging over different fold runs sometimes helps.</p></li>\n<li><p>More sensitive to incomplete word recordings (both in training and testing). Learning (maybe memorising) ‘LE’ ‘LEF’, 'EFT', etc. as 'LEFT' is harder for c-models than w-models, because c-models has to 'learn' the character that is not there in the sound.. It is quite a random guess in recognizing 'O' as 'go', 'off' or 'on' in the competition.</p></li>\n<li><p>Need prior knowledge of pronunciations and preprocessing, as there is no one-to-one mapping between characters and phonemes. But it is possible the models can learn the mapping (and memorizing exceptions) with larger training data.</p></li>\n<li><p>In the case of this particular Kaggle competition, the word-character (30-28) ratio is too small for c-models being efficient.</p></li>\n</ul>\n\n<p>But some advantages as well, possibly more suitable in practical uses:</p>\n\n<ul>\n<li><p>Generalisation. The c-model is able to learn unseen words, e.g. recognizing 'night' from learning 'NIne' and 'rIGHT', recognizing 'follow' as 'foow' (missing 'l' sound) from learning 'Four' 'dOWn' and 'nO'.</p></li>\n<li><p>Easier to scale up to larger vocabularies, with the increase of the number of LSTM hidden units, as well as the increase of the training time when learning longer words. Bi-Directional LSTM will be needed to capture more complex patterns.</p></li>\n<li><p>In practise, the property of high false negative (and low false positive) rate is likely a desired feature for key commands detecting applications. It is possible that the models will be customized to end users (with the balance of being general) so that the accents won't be a problem.</p></li>\n</ul>\n\n<p>Let me know what you think:-)</p>",
      "rawMarkdown": "Most models in this competition are CNN based and learns to categorise words. \n\nThis model is a Convolution Residual, backward LSTM network using Connectionist Temporal Classification (CTC) cost, written in TensorFlow. [Source code](https://github.com/huschen/kaggle_speech_recognition). It is more of a characters(phonemes) learning model.\n\n##Performance (Private Score):\nSingle model, single training run: 0.89357\n\nAverage over 5 folds (runs): 0.89874\n\nNo test data is used in training.\n\n##Comparision with Word-Learning Models:\nIn this particular competition and data set, character learning model has its limitations:\n\n* High false negative rate (i.e. the tendency to miss predict key words as unknown words). Simply because it has to get all the characters right to predict the correct word. An ambiguous 'off' sound might be recognised as 'oP' or 'Uf'. A 'seven' with an unheard accent might be predicted as 'sIven'. Averaging over different fold runs sometimes helps.\n\n* More sensitive to incomplete word recordings (both in training and testing). Learning (maybe memorising) ‘LE’ ‘LEF’, 'EFT', etc. as 'LEFT' is harder for c-models than w-models, because c-models has to 'learn' the character that is not there in the sound.. It is quite a random guess in recognizing 'O' as 'go', 'off' or 'on' in the competition.\n\n* Need prior knowledge of pronunciations and preprocessing, as there is no one-to-one mapping between characters and phonemes. But it is possible the models can learn the mapping (and memorizing exceptions) with larger training data.\n\n* In the case of this particular Kaggle competition, the word-character (30-28) ratio is too small for c-models being efficient.\n\nBut some advantages as well, possibly more suitable in practical uses:\n\n* Generalisation. The c-model is able to learn unseen words, e.g. recognizing 'night' from learning 'NIne' and 'rIGHT', recognizing 'follow' as 'foow' (missing 'l' sound) from learning 'Four' 'dOWn' and 'nO'.\n\n* Easier to scale up to larger vocabularies, with the increase of the number of LSTM hidden units, as well as the increase of the training time when learning longer words. Bi-Directional LSTM will be needed to capture more complex patterns.\n\n* In practise, the property of high false negative (and low false positive) rate is likely a desired feature for key commands detecting applications. It is possible that the models will be customized to end users (with the balance of being general) so that the accents won't be a problem.\n\nLet me know what you think:-)",
      "votes": null
    },
    {
      "id": "323645",
      "postDate": "05/05/2018 18:50:58",
      "content": "<p>The repo looks great! Thanks for sharing. Just started the training :P</p>",
      "rawMarkdown": "The repo looks great! Thanks for sharing. Just started the training :P",
      "votes": null
    },
    {
      "id": "323649",
      "postDate": "05/05/2018 19:03:15",
      "content": "<p>Thank you:-) </p>",
      "rawMarkdown": "Thank you:-)",
      "votes": null
    },
    {
      "id": "340871",
      "postDate": "06/10/2018 14:19:03",
      "content": "<p>\"Single model, single training run: 0.89357\" this is very good results!</p>",
      "rawMarkdown": "\"Single model, single training run: 0.89357\" this is very good results!",
      "votes": null
    },
    {
      "id": "340878",
      "postDate": "06/10/2018 14:45:39",
      "content": "<p>Thank you, but the model is unlikely to achieve your impressive score 0.91:-)\nAny insights into the dataset and my comparison of word learning and character learning models?  </p>",
      "rawMarkdown": "Thank you, but the model is unlikely to achieve your impressive score 0.91:-)\nAny insights into the dataset and my comparison of word learning and character learning models?",
      "votes": null
    },
    {
      "id": "349635",
      "postDate": "06/28/2018 12:16:01",
      "content": "<p>hello i am new in the speech field we are working in the task audio to text transcription. kindly guide me.</p>",
      "rawMarkdown": "hello i am new in the speech field we are working in the task audio to text transcription. kindly guide me.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 323645,
      "author_name": "seesee",
      "author_url": "",
      "post_date": "05/05/2018 18:50:58",
      "content": "<p>The repo looks great! Thanks for sharing. Just started the training :P</p>",
      "votes": null,
      "replies": [
        {
          "id": 323649,
          "author_name": "shijing",
          "author_url": "",
          "post_date": "05/05/2018 19:03:15",
          "content": "<p>Thank you:-) </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 340871,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "06/10/2018 14:19:03",
      "content": "<p>\"Single model, single training run: 0.89357\" this is very good results!</p>",
      "votes": null,
      "replies": [
        {
          "id": 340878,
          "author_name": "shijing",
          "author_url": "",
          "post_date": "06/10/2018 14:45:39",
          "content": "<p>Thank you, but the model is unlikely to achieve your impressive score 0.91:-)\nAny insights into the dataset and my comparison of word learning and character learning models?  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 349635,
      "author_name": "nadityacse",
      "author_url": "",
      "post_date": "06/28/2018 12:16:01",
      "content": "<p>hello i am new in the speech field we are working in the task audio to text transcription. kindly guide me.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "323588": "Most models in this competition are CNN based and learns to categorise words. \n\nThis model is a Convolution Residual, backward LSTM network using Connectionist Temporal Classification (CTC) cost, written in TensorFlow. [Source code](https://github.com/huschen/kaggle_speech_recognition). It is more of a characters(phonemes) learning model.\n\n##Performance (Private Score):\nSingle model, single training run: 0.89357\n\nAverage over 5 folds (runs): 0.89874\n\nNo test data is used in training.\n\n##Comparision with Word-Learning Models:\nIn this particular competition and data set, character learning model has its limitations:\n\n* High false negative rate (i.e. the tendency to miss predict key words as unknown words). Simply because it has to get all the characters right to predict the correct word. An ambiguous 'off' sound might be recognised as 'oP' or 'Uf'. A 'seven' with an unheard accent might be predicted as 'sIven'. Averaging over different fold runs sometimes helps.\n\n* More sensitive to incomplete word recordings (both in training and testing). Learning (maybe memorising) ‘LE’ ‘LEF’, 'EFT', etc. as 'LEFT' is harder for c-models than w-models, because c-models has to 'learn' the character that is not there in the sound.. It is quite a random guess in recognizing 'O' as 'go', 'off' or 'on' in the competition.\n\n* Need prior knowledge of pronunciations and preprocessing, as there is no one-to-one mapping between characters and phonemes. But it is possible the models can learn the mapping (and memorizing exceptions) with larger training data.\n\n* In the case of this particular Kaggle competition, the word-character (30-28) ratio is too small for c-models being efficient.\n\nBut some advantages as well, possibly more suitable in practical uses:\n\n* Generalisation. The c-model is able to learn unseen words, e.g. recognizing 'night' from learning 'NIne' and 'rIGHT', recognizing 'follow' as 'foow' (missing 'l' sound) from learning 'Four' 'dOWn' and 'nO'.\n\n* Easier to scale up to larger vocabularies, with the increase of the number of LSTM hidden units, as well as the increase of the training time when learning longer words. Bi-Directional LSTM will be needed to capture more complex patterns.\n\n* In practise, the property of high false negative (and low false positive) rate is likely a desired feature for key commands detecting applications. It is possible that the models will be customized to end users (with the balance of being general) so that the accents won't be a problem.\n\nLet me know what you think:-)",
    "323645": "The repo looks great! Thanks for sharing. Just started the training :P",
    "323649": "Thank you:-)",
    "340871": "\"Single model, single training run: 0.89357\" this is very good results!",
    "340878": "Thank you, but the model is unlikely to achieve your impressive score 0.91:-)\nAny insights into the dataset and my comparison of word learning and character learning models?",
    "349635": "hello i am new in the speech field we are working in the task audio to text transcription. kindly guide me."
  },
  "source": "meta"
}