{
  "id": 45313,
  "title": "Load test data into notebook",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/45313",
  "author_name": "",
  "post_date": "2017-12-09T04:37:53.368245200Z",
  "votes": null,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hi,\nI am new to Kaggle. Kindly help me How to load test data into note book?\n<code>check_output([\"ls\",\"../input/\"])</code> <br>\nOutput :: <br>\nsample_submission.csv <br>\ntrain  </p>\n\n<p><code>files = glob(os.path.join('../input/', 'test/audio/*wav'))</code> <br>\nOutput:: <br>\n[]</p>",
  "messages": [
    {
      "id": "255459",
      "postDate": "12/09/2017 04:37:53",
      "content": "<p>Hi,\nI am new to Kaggle. Kindly help me How to load test data into note book?\n<code>check_output([\"ls\",\"../input/\"])</code> <br>\nOutput :: <br>\nsample_submission.csv <br>\ntrain  </p>\n\n<p><code>files = glob(os.path.join('../input/', 'test/audio/*wav'))</code> <br>\nOutput:: <br>\n[]</p>",
      "rawMarkdown": "Hi,\nI am new to Kaggle. Kindly help me How to load test data into note book?\n```check_output([\"ls\",\"../input/\"])```  \nOutput ::  \nsample_submission.csv  \ntrain  \n\n```files = glob(os.path.join('../input/', 'test/audio/*wav'))```   \nOutput::  \n[]",
      "votes": null
    },
    {
      "id": "255547",
      "postDate": "12/09/2017 10:55:50",
      "content": "<p>That's a tricky question. The obvious answer is look at the kernels, but I still think the \"mainstream\" ideas are using this sound to image + deep learning as described here:</p>\n\n<p><a href=\"https://www.tensorflow.org/tutorials/audio_recognition\">https://www.tensorflow.org/tutorials/audio_recognition</a></p>\n\n<p>It surprisingly works well, but it may have limitations.</p>\n\n<p>My approach: First thing: Since I am not an expert, I will not do great in such a short time, but I am happy to learn a lot.</p>\n\n<p>Treat sound as sound: Read about FFT, DCT, STFT until you understand an practice.</p>\n\n<ol>\n<li>Read from the experts. (see below)</li>\n<li>Play with the files. (see below)</li>\n<li>Extract features that work with a classifier of you choice (xgboost, lightgbm, catboost, fast_rgf, ...)</li>\n</ol>\n\n<p>Experts, please extend this!</p>\n\n<hr>\n\n<p>(1) On readings:</p>\n\n<p>Deep Voice:\n    <a href=\"https://github.com/israelg99/deepvoice\">https://github.com/israelg99/deepvoice</a>\n    <a href=\"https://arxiv.org/pdf/1702.07825v2.pdf\">https://arxiv.org/pdf/1702.07825v2.pdf</a></p>\n\n<p>Tacotron:\n    <a href=\"https://arxiv.org/pdf/1703.10135.pdf\">https://arxiv.org/pdf/1703.10135.pdf</a>\n    <a href=\"https://github.com/keithito/tacotron\">https://github.com/keithito/tacotron</a>\n    <a href=\"https://github.com/Kyubyong/tacotron\">https://github.com/Kyubyong/tacotron</a>\n    <a href=\"https://github.com/barronalex/Tacotron\">https://github.com/barronalex/Tacotron</a>\n    <a href=\"https://google.github.io/tacotron/\">https://google.github.io/tacotron/</a></p>\n\n<p>WaveNet: Google Deep Mind: When applied to text-to-speech, it yields state-ofthe-art performance, with human listeners rating it as significantly more natural sounding than the best parametric and concatenative systems for both English and Mandarin. A single WaveNet can capture the characteristics of many different speakers with equal fidelity, and can switch between them by conditioning on the speaker identity.\n    <a href=\"https://arxiv.org/pdf/1609.03499v2.pdf\">https://arxiv.org/pdf/1609.03499v2.pdf</a></p>\n\n<p>(2) On playing with files:</p>\n\n<h1><a href=\"https://storage.googleapis.com/kaggle-forum-message-attachments/251419/7997/00-WAV-File-Inventory.html\">https://storage.googleapis.com/kaggle-forum-message-attachments/251419/7997/00-WAV-File-Inventory.html</a></h1>\n\n<p>suppressWarnings( library(dplyr) )\nlibrary(tibble)     # rownames_to_column\nlibrary(stringr)    # word\nlibrary(tools)      # md5sum\nlibrary(tuneR)      # readWave\nsuppressMessages( library(doParallel) )  # foreach</p>\n\n<p>trainFilenames &lt;- list.files(path = 'input/train', pattern='^.*\\.wav$', recursive = TRUE, full.names = TRUE)</p>\n\n<p>length(trainFilenames)\nhead(trainFilenames)</p>\n\n<p>cluster &lt;- makePSOCKcluster(detectCores())\nregisterDoParallel(cluster)</p>\n\n<p>trainSoundInfo &lt;- foreach(i = 1:length(trainFilenames), .combine = rbind) %dopar%\n{\n    soundWave &lt;- tuneR::readWave(trainFilenames[i])</p>\n\n<pre><code>data.frame(\n    filename          = trainFilenames[i],\n    samples           = length(soundWave@left),\n    durationSeconds   = length(soundWave@left)/soundWave@samp.rate,\n    samplingRateHertz = soundWave@samp.rate,\n    channels          = ifelse(soundWave@stereo, 'Stereo', 'Mono'),\n    pcmIntFormat      = soundWave@pcm,\n    bitsPerSample     = soundWave@bit,\n    md5sum            = tools::md5sum(trainFilenames[i]),\n    row.names         = NULL,\n    stringsAsFactors  = FALSE)\n</code></pre>\n\n<p>}</p>\n\n<p>stopCluster(cluster)</p>",
      "rawMarkdown": "That's a tricky question. The obvious answer is look at the kernels, but I still think the \"mainstream\" ideas are using this sound to image + deep learning as described here:\n\nhttps://www.tensorflow.org/tutorials/audio_recognition\n\nIt surprisingly works well, but it may have limitations.\n\nMy approach: First thing: Since I am not an expert, I will not do great in such a short time, but I am happy to learn a lot.\n\nTreat sound as sound: Read about FFT, DCT, STFT until you understand an practice.\n\n1. Read from the experts. (see below)\n2. Play with the files. (see below)\n3. Extract features that work with a classifier of you choice (xgboost, lightgbm, catboost, fast_rgf, ...)\n\nExperts, please extend this!\n\n---\n(1) On readings:\n\nDeep Voice:\n\thttps://github.com/israelg99/deepvoice\n\thttps://arxiv.org/pdf/1702.07825v2.pdf\n\nTacotron:\n\thttps://arxiv.org/pdf/1703.10135.pdf\n\thttps://github.com/keithito/tacotron\n\thttps://github.com/Kyubyong/tacotron\n\thttps://github.com/barronalex/Tacotron\n\thttps://google.github.io/tacotron/\n\nWaveNet: Google Deep Mind: When applied to text-to-speech, it yields state-ofthe-art performance, with human listeners rating it as significantly more natural sounding than the best parametric and concatenative systems for both English and Mandarin. A single WaveNet can capture the characteristics of many different speakers with equal fidelity, and can switch between them by conditioning on the speaker identity.\n\thttps://arxiv.org/pdf/1609.03499v2.pdf\n\n\n(2) On playing with files:\n\n# https://kaggle2.blob.core.windows.net/forum-message-attachments/251419/7997/00-WAV-File-Inventory.html\n\nsuppressWarnings( library(dplyr) )\nlibrary(tibble)     # rownames_to_column\nlibrary(stringr)    # word\nlibrary(tools)      # md5sum\nlibrary(tuneR)      # readWave\nsuppressMessages( library(doParallel) )  # foreach\n\ntrainFilenames &lt;- list.files(path = 'input/train', pattern='^.*\\\\.wav$', recursive = TRUE, full.names = TRUE)\n\nlength(trainFilenames)\nhead(trainFilenames)\n\ncluster &lt;- makePSOCKcluster(detectCores())\nregisterDoParallel(cluster)\n\ntrainSoundInfo &lt;- foreach(i = 1:length(trainFilenames), .combine = rbind) %dopar%\n{\n\tsoundWave &lt;- tuneR::readWave(trainFilenames[i])\n\n\tdata.frame(\n\t\tfilename          = trainFilenames[i],\n\t\tsamples           = length(soundWave@left),\n\t\tdurationSeconds   = length(soundWave@left)/soundWave@samp.rate,\n\t\tsamplingRateHertz = soundWave@samp.rate,\n\t\tchannels          = ifelse(soundWave@stereo, 'Stereo', 'Mono'),\n\t\tpcmIntFormat      = soundWave@pcm,\n\t\tbitsPerSample     = soundWave@bit,\n\t\tmd5sum            = tools::md5sum(trainFilenames[i]),\n\t\trow.names         = NULL,\n\t\tstringsAsFactors  = FALSE)\n}\n\nstopCluster(cluster)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 255547,
      "author_name": "batangas",
      "author_url": "",
      "post_date": "12/09/2017 10:55:50",
      "content": "<p>That's a tricky question. The obvious answer is look at the kernels, but I still think the \"mainstream\" ideas are using this sound to image + deep learning as described here:</p>\n\n<p><a href=\"https://www.tensorflow.org/tutorials/audio_recognition\">https://www.tensorflow.org/tutorials/audio_recognition</a></p>\n\n<p>It surprisingly works well, but it may have limitations.</p>\n\n<p>My approach: First thing: Since I am not an expert, I will not do great in such a short time, but I am happy to learn a lot.</p>\n\n<p>Treat sound as sound: Read about FFT, DCT, STFT until you understand an practice.</p>\n\n<ol>\n<li>Read from the experts. (see below)</li>\n<li>Play with the files. (see below)</li>\n<li>Extract features that work with a classifier of you choice (xgboost, lightgbm, catboost, fast_rgf, ...)</li>\n</ol>\n\n<p>Experts, please extend this!</p>\n\n<hr>\n\n<p>(1) On readings:</p>\n\n<p>Deep Voice:\n    <a href=\"https://github.com/israelg99/deepvoice\">https://github.com/israelg99/deepvoice</a>\n    <a href=\"https://arxiv.org/pdf/1702.07825v2.pdf\">https://arxiv.org/pdf/1702.07825v2.pdf</a></p>\n\n<p>Tacotron:\n    <a href=\"https://arxiv.org/pdf/1703.10135.pdf\">https://arxiv.org/pdf/1703.10135.pdf</a>\n    <a href=\"https://github.com/keithito/tacotron\">https://github.com/keithito/tacotron</a>\n    <a href=\"https://github.com/Kyubyong/tacotron\">https://github.com/Kyubyong/tacotron</a>\n    <a href=\"https://github.com/barronalex/Tacotron\">https://github.com/barronalex/Tacotron</a>\n    <a href=\"https://google.github.io/tacotron/\">https://google.github.io/tacotron/</a></p>\n\n<p>WaveNet: Google Deep Mind: When applied to text-to-speech, it yields state-ofthe-art performance, with human listeners rating it as significantly more natural sounding than the best parametric and concatenative systems for both English and Mandarin. A single WaveNet can capture the characteristics of many different speakers with equal fidelity, and can switch between them by conditioning on the speaker identity.\n    <a href=\"https://arxiv.org/pdf/1609.03499v2.pdf\">https://arxiv.org/pdf/1609.03499v2.pdf</a></p>\n\n<p>(2) On playing with files:</p>\n\n<h1><a href=\"https://storage.googleapis.com/kaggle-forum-message-attachments/251419/7997/00-WAV-File-Inventory.html\">https://storage.googleapis.com/kaggle-forum-message-attachments/251419/7997/00-WAV-File-Inventory.html</a></h1>\n\n<p>suppressWarnings( library(dplyr) )\nlibrary(tibble)     # rownames_to_column\nlibrary(stringr)    # word\nlibrary(tools)      # md5sum\nlibrary(tuneR)      # readWave\nsuppressMessages( library(doParallel) )  # foreach</p>\n\n<p>trainFilenames &lt;- list.files(path = 'input/train', pattern='^.*\\.wav$', recursive = TRUE, full.names = TRUE)</p>\n\n<p>length(trainFilenames)\nhead(trainFilenames)</p>\n\n<p>cluster &lt;- makePSOCKcluster(detectCores())\nregisterDoParallel(cluster)</p>\n\n<p>trainSoundInfo &lt;- foreach(i = 1:length(trainFilenames), .combine = rbind) %dopar%\n{\n    soundWave &lt;- tuneR::readWave(trainFilenames[i])</p>\n\n<pre><code>data.frame(\n    filename          = trainFilenames[i],\n    samples           = length(soundWave@left),\n    durationSeconds   = length(soundWave@left)/soundWave@samp.rate,\n    samplingRateHertz = soundWave@samp.rate,\n    channels          = ifelse(soundWave@stereo, 'Stereo', 'Mono'),\n    pcmIntFormat      = soundWave@pcm,\n    bitsPerSample     = soundWave@bit,\n    md5sum            = tools::md5sum(trainFilenames[i]),\n    row.names         = NULL,\n    stringsAsFactors  = FALSE)\n</code></pre>\n\n<p>}</p>\n\n<p>stopCluster(cluster)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "255459": "Hi,\nI am new to Kaggle. Kindly help me How to load test data into note book?\n```check_output([\"ls\",\"../input/\"])```  \nOutput ::  \nsample_submission.csv  \ntrain  \n\n```files = glob(os.path.join('../input/', 'test/audio/*wav'))```   \nOutput::  \n[]",
    "255547": "That's a tricky question. The obvious answer is look at the kernels, but I still think the \"mainstream\" ideas are using this sound to image + deep learning as described here:\n\nhttps://www.tensorflow.org/tutorials/audio_recognition\n\nIt surprisingly works well, but it may have limitations.\n\nMy approach: First thing: Since I am not an expert, I will not do great in such a short time, but I am happy to learn a lot.\n\nTreat sound as sound: Read about FFT, DCT, STFT until you understand an practice.\n\n1. Read from the experts. (see below)\n2. Play with the files. (see below)\n3. Extract features that work with a classifier of you choice (xgboost, lightgbm, catboost, fast_rgf, ...)\n\nExperts, please extend this!\n\n---\n(1) On readings:\n\nDeep Voice:\n\thttps://github.com/israelg99/deepvoice\n\thttps://arxiv.org/pdf/1702.07825v2.pdf\n\nTacotron:\n\thttps://arxiv.org/pdf/1703.10135.pdf\n\thttps://github.com/keithito/tacotron\n\thttps://github.com/Kyubyong/tacotron\n\thttps://github.com/barronalex/Tacotron\n\thttps://google.github.io/tacotron/\n\nWaveNet: Google Deep Mind: When applied to text-to-speech, it yields state-ofthe-art performance, with human listeners rating it as significantly more natural sounding than the best parametric and concatenative systems for both English and Mandarin. A single WaveNet can capture the characteristics of many different speakers with equal fidelity, and can switch between them by conditioning on the speaker identity.\n\thttps://arxiv.org/pdf/1609.03499v2.pdf\n\n\n(2) On playing with files:\n\n# https://kaggle2.blob.core.windows.net/forum-message-attachments/251419/7997/00-WAV-File-Inventory.html\n\nsuppressWarnings( library(dplyr) )\nlibrary(tibble)     # rownames_to_column\nlibrary(stringr)    # word\nlibrary(tools)      # md5sum\nlibrary(tuneR)      # readWave\nsuppressMessages( library(doParallel) )  # foreach\n\ntrainFilenames &lt;- list.files(path = 'input/train', pattern='^.*\\\\.wav$', recursive = TRUE, full.names = TRUE)\n\nlength(trainFilenames)\nhead(trainFilenames)\n\ncluster &lt;- makePSOCKcluster(detectCores())\nregisterDoParallel(cluster)\n\ntrainSoundInfo &lt;- foreach(i = 1:length(trainFilenames), .combine = rbind) %dopar%\n{\n\tsoundWave &lt;- tuneR::readWave(trainFilenames[i])\n\n\tdata.frame(\n\t\tfilename          = trainFilenames[i],\n\t\tsamples           = length(soundWave@left),\n\t\tdurationSeconds   = length(soundWave@left)/soundWave@samp.rate,\n\t\tsamplingRateHertz = soundWave@samp.rate,\n\t\tchannels          = ifelse(soundWave@stereo, 'Stereo', 'Mono'),\n\t\tpcmIntFormat      = soundWave@pcm,\n\t\tbitsPerSample     = soundWave@bit,\n\t\tmd5sum            = tools::md5sum(trainFilenames[i]),\n\t\trow.names         = NULL,\n\t\tstringsAsFactors  = FALSE)\n}\n\nstopCluster(cluster)"
  },
  "source": "meta"
}