{
  "id": 46046,
  "title": "Up, off",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/46046",
  "author_name": "JihaoLiu",
  "post_date": "2017-12-20T02:22:31.578000",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>As can be seen from the confusion matrix, the neural network is difficult to distinguish between the two categories \"up\" and \"off\".\nI've tried to make the feature map larger, but get a worse result.</p>",
  "messages": [
    {
      "id": 260517,
      "postDate": "2017-12-20T11:26:00.410Z",
      "content": "<p>Both words are rather short. You could try playing them slower and average both predictions (slow &amp; default version of the record). This gave me a small improvement (low 87% to higher 87%). The <code>librosa</code> method for time stretching is slow. You can't use it online. I therefore dumped the files to disk. Like so:</p>\n\n<pre><code>from librosa import effects\nfrom tqdm import tqdm\nfrom glob import glob\nimport numpy as np\nfrom scipy.io import wavfile as wf\nfrom os.path import join as jp, basename as bn\n\n\ndef main():\n  tta_speed = 0.9  # slow down (i.e. &amp;lt; 1.0)\n  samples_per_sec = 16000\n  test_fns = sorted(glob('data/test/audio/*.wav'))\n  tta_dir = 'data/tta_test/audio'\n  for fn in tqdm(test_fns):\n    basename = bn(fn)\n    rate, data = wf.read(fn)\n    assert len(data) == samples_per_sec\n    data = np.float32(data) / 32767\n    data = effects.time_stretch(data, tta_speed)\n    data = data[-samples_per_sec:]\n    out_fn = jp(tta_dir, basename)\n    wf.write(out_fn, rate, np.int16(data * 32767))\n\n\nif __name__ == '__main__':\n  main()\n</code></pre>\n\n<p>Though, I didn't specifically check for \"up\" vs \"off\".</p>",
      "rawMarkdown": "Both words are rather short. You could try playing them slower and average both predictions (slow &amp; default version of the record). This gave me a small improvement (low 87% to higher 87%). The `librosa` method for time stretching is slow. You can't use it online. I therefore dumped the files to disk. Like so:\n\n    from librosa import effects\n    from tqdm import tqdm\n    from glob import glob\n    import numpy as np\n    from scipy.io import wavfile as wf\n    from os.path import join as jp, basename as bn\n    \n    \n    def main():\n      tta_speed = 0.9  # slow down (i.e. &lt; 1.0)\n      samples_per_sec = 16000\n      test_fns = sorted(glob('data/test/audio/*.wav'))\n      tta_dir = 'data/tta_test/audio'\n      for fn in tqdm(test_fns):\n        basename = bn(fn)\n        rate, data = wf.read(fn)\n        assert len(data) == samples_per_sec\n        data = np.float32(data) / 32767\n        data = effects.time_stretch(data, tta_speed)\n        data = data[-samples_per_sec:]\n        out_fn = jp(tta_dir, basename)\n        wf.write(out_fn, rate, np.int16(data * 32767))\n    \n    \n    if __name__ == '__main__':\n      main()\n    \n\nThough, I didn't specifically check for \"up\" vs \"off\".",
      "votes": 7,
      "replies": [
        {
          "id": 260553,
          "postDate": "2017-12-20T13:00:32.490Z",
          "content": "<p>Thank you so much.</p>",
          "rawMarkdown": "Thank you so much."
        }
      ]
    },
    {
      "id": 260300,
      "postDate": "2017-12-20T02:22:31.580Z",
      "content": "<p>As can be seen from the confusion matrix, the neural network is difficult to distinguish between the two categories \"up\" and \"off\".\nI've tried to make the feature map larger, but get a worse result.</p>",
      "rawMarkdown": "As can be seen from the confusion matrix, the neural network is difficult to distinguish between the two categories \"up\" and \"off\".\nI've tried to make the feature map larger, but get a worse result.",
      "votes": 1
    },
    {
      "id": 264267,
      "postDate": "2018-01-02T18:04:12.353Z",
      "content": "<p>You could train a One vs One model when your global model predicts either \"up\" or \"off\".</p>",
      "rawMarkdown": "You could train a One vs One model when your global model predicts either \"up\" or \"off\"."
    },
    {
      "id": 260699,
      "postDate": "2017-12-20T18:41:37.720Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 260893,
          "postDate": "2017-12-21T05:50:25.207Z",
          "content": "<p>OK. I am not familiar with attention model. This is a good chance to learn something about it.\nThank you so much.</p>",
          "rawMarkdown": "OK. I am not familiar with attention model. This is a good chance to learn something about it.\nThank you so much."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 260517,
      "author_name": "See--",
      "author_url": "",
      "post_date": "2017-12-20T11:26:00.410000",
      "content": "<p>Both words are rather short. You could try playing them slower and average both predictions (slow &amp; default version of the record). This gave me a small improvement (low 87% to higher 87%). The <code>librosa</code> method for time stretching is slow. You can't use it online. I therefore dumped the files to disk. Like so:</p>\n\n<pre><code>from librosa import effects\nfrom tqdm import tqdm\nfrom glob import glob\nimport numpy as np\nfrom scipy.io import wavfile as wf\nfrom os.path import join as jp, basename as bn\n\n\ndef main():\n  tta_speed = 0.9  # slow down (i.e. &amp;lt; 1.0)\n  samples_per_sec = 16000\n  test_fns = sorted(glob('data/test/audio/*.wav'))\n  tta_dir = 'data/tta_test/audio'\n  for fn in tqdm(test_fns):\n    basename = bn(fn)\n    rate, data = wf.read(fn)\n    assert len(data) == samples_per_sec\n    data = np.float32(data) / 32767\n    data = effects.time_stretch(data, tta_speed)\n    data = data[-samples_per_sec:]\n    out_fn = jp(tta_dir, basename)\n    wf.write(out_fn, rate, np.int16(data * 32767))\n\n\nif __name__ == '__main__':\n  main()\n</code></pre>\n\n<p>Though, I didn't specifically check for \"up\" vs \"off\".</p>",
      "votes": 7,
      "replies": [
        {
          "id": 260553,
          "author_name": "JihaoLiu",
          "author_url": "",
          "post_date": "2017-12-20T13:00:32.490000",
          "content": "<p>Thank you so much.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 264267,
      "author_name": "Marvin Lerousseau",
      "author_url": "",
      "post_date": "2018-01-02T18:04:12.353000",
      "content": "<p>You could train a One vs One model when your global model predicts either \"up\" or \"off\".</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 260699,
      "author_name": "",
      "author_url": "",
      "post_date": "2017-12-20T18:41:37.720000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 260893,
          "author_name": "JihaoLiu",
          "author_url": "",
          "post_date": "2017-12-21T05:50:25.207000",
          "content": "<p>OK. I am not familiar with attention model. This is a good chance to learn something about it.\nThank you so much.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "260517": "Both words are rather short. You could try playing them slower and average both predictions (slow &amp; default version of the record). This gave me a small improvement (low 87% to higher 87%). The `librosa` method for time stretching is slow. You can't use it online. I therefore dumped the files to disk. Like so:\n\n    from librosa import effects\n    from tqdm import tqdm\n    from glob import glob\n    import numpy as np\n    from scipy.io import wavfile as wf\n    from os.path import join as jp, basename as bn\n    \n    \n    def main():\n      tta_speed = 0.9  # slow down (i.e. &lt; 1.0)\n      samples_per_sec = 16000\n      test_fns = sorted(glob('data/test/audio/*.wav'))\n      tta_dir = 'data/tta_test/audio'\n      for fn in tqdm(test_fns):\n        basename = bn(fn)\n        rate, data = wf.read(fn)\n        assert len(data) == samples_per_sec\n        data = np.float32(data) / 32767\n        data = effects.time_stretch(data, tta_speed)\n        data = data[-samples_per_sec:]\n        out_fn = jp(tta_dir, basename)\n        wf.write(out_fn, rate, np.int16(data * 32767))\n    \n    \n    if __name__ == '__main__':\n      main()\n    \n\nThough, I didn't specifically check for \"up\" vs \"off\".",
    "260300": "As can be seen from the confusion matrix, the neural network is difficult to distinguish between the two categories \"up\" and \"off\".\nI've tried to make the feature map larger, but get a worse result.",
    "264267": "You could train a One vs One model when your global model predicts either \"up\" or \"off\".",
    "260699": ""
  }
}