{
  "id": 97241,
  "title": "FreeSound Tagging vs. Speech Recognition",
  "url": "/competitions/freesound-audio-tagging-2019/discussion/97241",
  "author_name": "",
  "post_date": "2019-06-26T11:18:26.357705700Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Dear all fellow participants,</p>\n\n<p>This is a kind of a newbie question but I have it in mind and almost forget to ask all your opinion. </p>\n\n<p>1) Do you think the problem we have here, Freesound Audio Tagging, is just a special (simpler) kind of speech recognition task? [i.e. in speech recognition, we have to recognize many words anywhere in the audio and the number of possible words are many more than the labels we have here]</p>\n\n<p>2) If this task is much easier, why don't we can just use SOTA speech recognition algorithm (e.g. Deep Speech) ... </p>\n\n<p>3) Or there is some perspective that this Freesound tagging problem is more difficult than recognize words from audio ?</p>\n\n<p>Hope to hear what you guys think! And also would like to see opinions from organizers who are experts in the field as well <a href=\"/fredericfont\">@fredericfont</a> <a href=\"/plakal\">@plakal</a> <a href=\"/eduardofonseca\">@eduardofonseca</a> </p>",
  "messages": [
    {
      "id": "561369",
      "postDate": "06/26/2019 11:18:26",
      "content": "<p>Dear all fellow participants,</p>\n\n<p>This is a kind of a newbie question but I have it in mind and almost forget to ask all your opinion. </p>\n\n<p>1) Do you think the problem we have here, Freesound Audio Tagging, is just a special (simpler) kind of speech recognition task? [i.e. in speech recognition, we have to recognize many words anywhere in the audio and the number of possible words are many more than the labels we have here]</p>\n\n<p>2) If this task is much easier, why don't we can just use SOTA speech recognition algorithm (e.g. Deep Speech) ... </p>\n\n<p>3) Or there is some perspective that this Freesound tagging problem is more difficult than recognize words from audio ?</p>\n\n<p>Hope to hear what you guys think! And also would like to see opinions from organizers who are experts in the field as well <a href=\"/fredericfont\">@fredericfont</a> <a href=\"/plakal\">@plakal</a> <a href=\"/eduardofonseca\">@eduardofonseca</a> </p>",
      "rawMarkdown": "Dear all fellow participants,\n\nThis is a kind of a newbie question but I have it in mind and almost forget to ask all your opinion. \n\n1) Do you think the problem we have here, Freesound Audio Tagging, is just a special (simpler) kind of speech recognition task? [i.e. in speech recognition, we have to recognize many words anywhere in the audio and the number of possible words are many more than the labels we have here]\n\n2) If this task is much easier, why don't we can just use SOTA speech recognition algorithm (e.g. Deep Speech) ... \n\n3) Or there is some perspective that this Freesound tagging problem is more difficult than recognize words from audio ?\n\nHope to hear what you guys think! And also would like to see opinions from organizers who are experts in the field as well @fredericfont @plakal @eduardofonseca",
      "votes": null
    },
    {
      "id": "565698",
      "postDate": "07/01/2019 09:35:03",
      "content": "<p>1) Speech recognition (SR) is many-to-many problem, where you don't know how may inputs corresponds to how many outputs. Moreover, the targets in SR are dependent on each other, which is not true in this competition. These difference may seem vague, but they are crucial - here you know at a glance how many inputs do you have, you usually have one target, and if more, they are more or less independent of each other. That's why the methods here and in SR are quite different.</p>\n\n<p>2) Deep Speech is not that great I think now. One of the main features of Deep Speech is that the criterion function is minimizing the loss over all possible sequences of frames and targets. It could be prepared to work with classification but it does not default and would not be natural (in my opinion) for classification problem.</p>\n\n<p>I think we have a lot of tricks here to classify the sounds, like data augmentation. They are probably crucial for this competition. Again it can't be used with SR, because f.e mixup would work only if targets in SR would be independent. TTA - ok to use, but speech recognition often has to be used with real-time speed, and the machines are not there yet. </p>\n\n<p>Deep speech doesn't utilize language properties cleverly, it just fuses language with audio at some point. But even that doesn't occur here. In SR, using Language Model in recognition is an important component and field of research.</p>\n\n<p>3) It's hard to say what is harder, it's just different. But the facts are that SR needs a lot more data to work decent, and the variety of sounds, accents, voices is huge, and there are a lot of things to say at the language layer. Both the variance and data are incredibly huge in SR. I doubt there's a single system in the world that can really work similar to human - in some search queries in a browser it's ok, but then it's not working in specialized speech f.e. medical, and opposite.</p>",
      "rawMarkdown": "1) Speech recognition (SR) is many-to-many problem, where you don't know how may inputs corresponds to how many outputs. Moreover, the targets in SR are dependent on each other, which is not true in this competition. These difference may seem vague, but they are crucial - here you know at a glance how many inputs do you have, you usually have one target, and if more, they are more or less independent of each other. That's why the methods here and in SR are quite different.\n\n2) Deep Speech is not that great I think now. One of the main features of Deep Speech is that the criterion function is minimizing the loss over all possible sequences of frames and targets. It could be prepared to work with classification but it does not default and would not be natural (in my opinion) for classification problem.\n\nI think we have a lot of tricks here to classify the sounds, like data augmentation. They are probably crucial for this competition. Again it can't be used with SR, because f.e mixup would work only if targets in SR would be independent. TTA - ok to use, but speech recognition often has to be used with real-time speed, and the machines are not there yet. \n\nDeep speech doesn't utilize language properties cleverly, it just fuses language with audio at some point. But even that doesn't occur here. In SR, using Language Model in recognition is an important component and field of research.\n\n3) It's hard to say what is harder, it's just different. But the facts are that SR needs a lot more data to work decent, and the variety of sounds, accents, voices is huge, and there are a lot of things to say at the language layer. Both the variance and data are incredibly huge in SR. I doubt there's a single system in the world that can really work similar to human - in some search queries in a browser it's ok, but then it's not working in specialized speech f.e. medical, and opposite.",
      "votes": null
    },
    {
      "id": "566345",
      "postDate": "07/02/2019 04:38:56",
      "content": "<p>Hi <a href=\"/davids1992\">@davids1992</a> ,</p>\n\n<p>Truly appreciate your thoughtful comments!! Let me slowly think about all your writing!</p>\n\n<p>Thanks again.</p>",
      "rawMarkdown": "Hi @davids1992 ,\n\nTruly appreciate your thoughtful comments!! Let me slowly think about all your writing!\n\nThanks again.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 565698,
      "author_name": "davids1992",
      "author_url": "",
      "post_date": "07/01/2019 09:35:03",
      "content": "<p>1) Speech recognition (SR) is many-to-many problem, where you don't know how may inputs corresponds to how many outputs. Moreover, the targets in SR are dependent on each other, which is not true in this competition. These difference may seem vague, but they are crucial - here you know at a glance how many inputs do you have, you usually have one target, and if more, they are more or less independent of each other. That's why the methods here and in SR are quite different.</p>\n\n<p>2) Deep Speech is not that great I think now. One of the main features of Deep Speech is that the criterion function is minimizing the loss over all possible sequences of frames and targets. It could be prepared to work with classification but it does not default and would not be natural (in my opinion) for classification problem.</p>\n\n<p>I think we have a lot of tricks here to classify the sounds, like data augmentation. They are probably crucial for this competition. Again it can't be used with SR, because f.e mixup would work only if targets in SR would be independent. TTA - ok to use, but speech recognition often has to be used with real-time speed, and the machines are not there yet. </p>\n\n<p>Deep speech doesn't utilize language properties cleverly, it just fuses language with audio at some point. But even that doesn't occur here. In SR, using Language Model in recognition is an important component and field of research.</p>\n\n<p>3) It's hard to say what is harder, it's just different. But the facts are that SR needs a lot more data to work decent, and the variety of sounds, accents, voices is huge, and there are a lot of things to say at the language layer. Both the variance and data are incredibly huge in SR. I doubt there's a single system in the world that can really work similar to human - in some search queries in a browser it's ok, but then it's not working in specialized speech f.e. medical, and opposite.</p>",
      "votes": null,
      "replies": [
        {
          "id": 566345,
          "author_name": "ratthachat",
          "author_url": "",
          "post_date": "07/02/2019 04:38:56",
          "content": "<p>Hi <a href=\"/davids1992\">@davids1992</a> ,</p>\n\n<p>Truly appreciate your thoughtful comments!! Let me slowly think about all your writing!</p>\n\n<p>Thanks again.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "561369": "Dear all fellow participants,\n\nThis is a kind of a newbie question but I have it in mind and almost forget to ask all your opinion. \n\n1) Do you think the problem we have here, Freesound Audio Tagging, is just a special (simpler) kind of speech recognition task? [i.e. in speech recognition, we have to recognize many words anywhere in the audio and the number of possible words are many more than the labels we have here]\n\n2) If this task is much easier, why don't we can just use SOTA speech recognition algorithm (e.g. Deep Speech) ... \n\n3) Or there is some perspective that this Freesound tagging problem is more difficult than recognize words from audio ?\n\nHope to hear what you guys think! And also would like to see opinions from organizers who are experts in the field as well @fredericfont @plakal @eduardofonseca",
    "565698": "1) Speech recognition (SR) is many-to-many problem, where you don't know how may inputs corresponds to how many outputs. Moreover, the targets in SR are dependent on each other, which is not true in this competition. These difference may seem vague, but they are crucial - here you know at a glance how many inputs do you have, you usually have one target, and if more, they are more or less independent of each other. That's why the methods here and in SR are quite different.\n\n2) Deep Speech is not that great I think now. One of the main features of Deep Speech is that the criterion function is minimizing the loss over all possible sequences of frames and targets. It could be prepared to work with classification but it does not default and would not be natural (in my opinion) for classification problem.\n\nI think we have a lot of tricks here to classify the sounds, like data augmentation. They are probably crucial for this competition. Again it can't be used with SR, because f.e mixup would work only if targets in SR would be independent. TTA - ok to use, but speech recognition often has to be used with real-time speed, and the machines are not there yet. \n\nDeep speech doesn't utilize language properties cleverly, it just fuses language with audio at some point. But even that doesn't occur here. In SR, using Language Model in recognition is an important component and field of research.\n\n3) It's hard to say what is harder, it's just different. But the facts are that SR needs a lot more data to work decent, and the variety of sounds, accents, voices is huge, and there are a lot of things to say at the language layer. Both the variance and data are incredibly huge in SR. I doubt there's a single system in the world that can really work similar to human - in some search queries in a browser it's ok, but then it's not working in specialized speech f.e. medical, and opposite.",
    "566345": "Hi @davids1992 ,\n\nTruly appreciate your thoughtful comments!! Let me slowly think about all your writing!\n\nThanks again."
  },
  "source": "meta"
}