{
  "id": 437165,
  "title": "Why RNN are not now used for speech recognition?Why they are now taken over by tranformer models",
  "url": "/competitions/bengaliai-speech/discussion/437165",
  "author_name": "",
  "post_date": "2023-09-05T17:35:00.501862800Z",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>RNNs are giving a very low accuracy when we go for speech recognition , why it is now failing to do its task which it was primarily made for</p>",
  "messages": [
    {
      "id": "2425176",
      "postDate": "09/05/2023 17:35:00",
      "content": "<p>RNNs are giving a very low accuracy when we go for speech recognition , why it is now failing to do its task which it was primarily made for</p>",
      "rawMarkdown": "RNNs are giving a very low accuracy when we go for speech recognition , why it is now failing to do its task which it was primarily made for",
      "votes": null
    },
    {
      "id": "2425262",
      "postDate": "09/05/2023 18:25:16",
      "content": "<p>RNNs (Recurrent Neural Networks) were initially considered promising for sequence processing tasks, including speech recognition. However, they have inherent limitations that might impact their efficacy in more complex tasks such as speech recognition:</p>\n<p>Vanishing Gradient Problem: During backpropagation, the gradients in an RNN can become very small (vanish) or very large (explode). A vanishing gradient means that the weights of the network get very weak updates, making it hard for RNNs to learn long-term dependencies in data.</p>\n<p>Long-term Dependencies: RNNs struggle to learn long-term dependencies in a sequence. For speech recognition tasks, where context and information from previous words can be crucial, this becomes a significant issue.</p>\n<p>Modeling Complex Structures: Speech is a complex sequence of sounds that often requires understanding context, intonation, and several other factors. More advanced architectures like LSTM (Long Short-Term Memory) or GRU (Gated Recurrent Unit) were developed specifically to address some of these challenges better than traditional RNNs.</p>\n<p>Computational Complexity: For very long sequences, RNNs can become impractical due to high computational complexity.</p>\n<p>Modern Approaches: Nowadays, convolutional neural networks (CNNs) and attention-based models (e.g., Transformers) have become popular and effective in many speech recognition tasks.</p>\n<p>In conclusion, while RNNs were designed for sequence processing, they have certain limitations that make them less effective in more complex tasks like speech recognition compared to newer architectures. If you're finding that RNNs are yielding low accuracy in your task, you might consider exploring other architectures like LSTM, GRU, or even Transformers.</p>",
      "rawMarkdown": "RNNs (Recurrent Neural Networks) were initially considered promising for sequence processing tasks, including speech recognition. However, they have inherent limitations that might impact their efficacy in more complex tasks such as speech recognition:\n\nVanishing Gradient Problem: During backpropagation, the gradients in an RNN can become very small (vanish) or very large (explode). A vanishing gradient means that the weights of the network get very weak updates, making it hard for RNNs to learn long-term dependencies in data.\n\nLong-term Dependencies: RNNs struggle to learn long-term dependencies in a sequence. For speech recognition tasks, where context and information from previous words can be crucial, this becomes a significant issue.\n\nModeling Complex Structures: Speech is a complex sequence of sounds that often requires understanding context, intonation, and several other factors. More advanced architectures like LSTM (Long Short-Term Memory) or GRU (Gated Recurrent Unit) were developed specifically to address some of these challenges better than traditional RNNs.\n\nComputational Complexity: For very long sequences, RNNs can become impractical due to high computational complexity.\n\nModern Approaches: Nowadays, convolutional neural networks (CNNs) and attention-based models (e.g., Transformers) have become popular and effective in many speech recognition tasks.\n\nIn conclusion, while RNNs were designed for sequence processing, they have certain limitations that make them less effective in more complex tasks like speech recognition compared to newer architectures. If you're finding that RNNs are yielding low accuracy in your task, you might consider exploring other architectures like LSTM, GRU, or even Transformers.",
      "votes": null
    },
    {
      "id": "2425291",
      "postDate": "09/05/2023 19:00:52",
      "content": "<p>To piggyback on HuBERT's great answer, a lot of why transformers are better is that you can parallelize their training. For RNN's, you gotta wait until the entire sequence has been passed through your net, compute your loss, then do <a href=\"https://d2l.ai/chapter_recurrent-neural-networks/bptt.html\" target=\"_blank\">backprop through time</a>. This stinks because you can't do anything until your sequence is done processing. No weight updates until you've done all your matrix multiplications.</p>\n<p>However if you have a transformer, you can parallelize the heck out of it. Each input element attends (pays attentions) to the others, but you can calculate Attention(x1, x2) at the same time as Attention(x1, x3) if you have the money for some good GPU's. Thus, you get a lot more training data fed to your model, hence, better results. </p>\n<p>Who knows though. Maybe RNN's and LSTM's will come back in style. Models have a way of dying out in popularity only to come back again with new math tricks added.</p>",
      "rawMarkdown": "To piggyback on HuBERT's great answer, a lot of why transformers are better is that you can parallelize their training. For RNN's, you gotta wait until the entire sequence has been passed through your net, compute your loss, then do [backprop through time](https://d2l.ai/chapter_recurrent-neural-networks/bptt.html). This stinks because you can't do anything until your sequence is done processing. No weight updates until you've done all your matrix multiplications.\n\nHowever if you have a transformer, you can parallelize the heck out of it. Each input element attends (pays attentions) to the others, but you can calculate Attention(x1, x2) at the same time as Attention(x1, x3) if you have the money for some good GPU's. Thus, you get a lot more training data fed to your model, hence, better results. \n\nWho knows though. Maybe RNN's and LSTM's will come back in style. Models have a way of dying out in popularity only to come back again with new math tricks added.",
      "votes": null
    },
    {
      "id": "2427216",
      "postDate": "09/07/2023 05:52:29",
      "content": "<p>TLDR: Attention is all you need.</p>",
      "rawMarkdown": "TLDR: Attention is all you need.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2425262,
      "author_name": "hubert101",
      "author_url": "",
      "post_date": "09/05/2023 18:25:16",
      "content": "<p>RNNs (Recurrent Neural Networks) were initially considered promising for sequence processing tasks, including speech recognition. However, they have inherent limitations that might impact their efficacy in more complex tasks such as speech recognition:</p>\n<p>Vanishing Gradient Problem: During backpropagation, the gradients in an RNN can become very small (vanish) or very large (explode). A vanishing gradient means that the weights of the network get very weak updates, making it hard for RNNs to learn long-term dependencies in data.</p>\n<p>Long-term Dependencies: RNNs struggle to learn long-term dependencies in a sequence. For speech recognition tasks, where context and information from previous words can be crucial, this becomes a significant issue.</p>\n<p>Modeling Complex Structures: Speech is a complex sequence of sounds that often requires understanding context, intonation, and several other factors. More advanced architectures like LSTM (Long Short-Term Memory) or GRU (Gated Recurrent Unit) were developed specifically to address some of these challenges better than traditional RNNs.</p>\n<p>Computational Complexity: For very long sequences, RNNs can become impractical due to high computational complexity.</p>\n<p>Modern Approaches: Nowadays, convolutional neural networks (CNNs) and attention-based models (e.g., Transformers) have become popular and effective in many speech recognition tasks.</p>\n<p>In conclusion, while RNNs were designed for sequence processing, they have certain limitations that make them less effective in more complex tasks like speech recognition compared to newer architectures. If you're finding that RNNs are yielding low accuracy in your task, you might consider exploring other architectures like LSTM, GRU, or even Transformers.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2425291,
      "author_name": "msthil",
      "author_url": "",
      "post_date": "09/05/2023 19:00:52",
      "content": "<p>To piggyback on HuBERT's great answer, a lot of why transformers are better is that you can parallelize their training. For RNN's, you gotta wait until the entire sequence has been passed through your net, compute your loss, then do <a href=\"https://d2l.ai/chapter_recurrent-neural-networks/bptt.html\" target=\"_blank\">backprop through time</a>. This stinks because you can't do anything until your sequence is done processing. No weight updates until you've done all your matrix multiplications.</p>\n<p>However if you have a transformer, you can parallelize the heck out of it. Each input element attends (pays attentions) to the others, but you can calculate Attention(x1, x2) at the same time as Attention(x1, x3) if you have the money for some good GPU's. Thus, you get a lot more training data fed to your model, hence, better results. </p>\n<p>Who knows though. Maybe RNN's and LSTM's will come back in style. Models have a way of dying out in popularity only to come back again with new math tricks added.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2427216,
      "author_name": "botdeveloper11",
      "author_url": "",
      "post_date": "09/07/2023 05:52:29",
      "content": "<p>TLDR: Attention is all you need.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2425176": "RNNs are giving a very low accuracy when we go for speech recognition , why it is now failing to do its task which it was primarily made for",
    "2425262": "RNNs (Recurrent Neural Networks) were initially considered promising for sequence processing tasks, including speech recognition. However, they have inherent limitations that might impact their efficacy in more complex tasks such as speech recognition:\n\nVanishing Gradient Problem: During backpropagation, the gradients in an RNN can become very small (vanish) or very large (explode). A vanishing gradient means that the weights of the network get very weak updates, making it hard for RNNs to learn long-term dependencies in data.\n\nLong-term Dependencies: RNNs struggle to learn long-term dependencies in a sequence. For speech recognition tasks, where context and information from previous words can be crucial, this becomes a significant issue.\n\nModeling Complex Structures: Speech is a complex sequence of sounds that often requires understanding context, intonation, and several other factors. More advanced architectures like LSTM (Long Short-Term Memory) or GRU (Gated Recurrent Unit) were developed specifically to address some of these challenges better than traditional RNNs.\n\nComputational Complexity: For very long sequences, RNNs can become impractical due to high computational complexity.\n\nModern Approaches: Nowadays, convolutional neural networks (CNNs) and attention-based models (e.g., Transformers) have become popular and effective in many speech recognition tasks.\n\nIn conclusion, while RNNs were designed for sequence processing, they have certain limitations that make them less effective in more complex tasks like speech recognition compared to newer architectures. If you're finding that RNNs are yielding low accuracy in your task, you might consider exploring other architectures like LSTM, GRU, or even Transformers.",
    "2425291": "To piggyback on HuBERT's great answer, a lot of why transformers are better is that you can parallelize their training. For RNN's, you gotta wait until the entire sequence has been passed through your net, compute your loss, then do [backprop through time](https://d2l.ai/chapter_recurrent-neural-networks/bptt.html). This stinks because you can't do anything until your sequence is done processing. No weight updates until you've done all your matrix multiplications.\n\nHowever if you have a transformer, you can parallelize the heck out of it. Each input element attends (pays attentions) to the others, but you can calculate Attention(x1, x2) at the same time as Attention(x1, x3) if you have the money for some good GPU's. Thus, you get a lot more training data fed to your model, hence, better results. \n\nWho knows though. Maybe RNN's and LSTM's will come back in style. Models have a way of dying out in popularity only to come back again with new math tricks added.",
    "2427216": "TLDR: Attention is all you need."
  },
  "source": "meta"
}