{
  "id": 44970,
  "title": "Getting a list of files, labels",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/44970",
  "author_name": "",
  "post_date": "2017-12-04T22:29:36.597810500Z",
  "votes": 2,
  "comment_count": 1,
  "views": 0,
  "content": "<p>For anyone looking to have a csv of files (and labels of the train set), I wrote a couple of scripts I thought I'd share. For performance reasons, it was quickest to get them from the archive:</p>\n\n<h2>Create a csv of the train files with labels:</h2>\n\n<pre><code>OUTFILE=./input/train_files.csv\necho \"fname,label\" &gt; $OUTFILE\n7za l ./input/train.7z | grep '.wav$' | awk '{print $(NF)}' | awk -F'/' '{print $4,$3}' OFS=\",\" &gt;&gt; $OUTFILE\n</code></pre>\n\n<h2>Create a csv of the test file names:</h2>\n\n<pre><code>OUTFILE=./input/test_files.csv\necho \"fname\" &gt; $OUTFILE\n7za l ./input/test.7z | grep '.wav$' | awk '{print $(NF)}' | cut -f3 -d'/' &gt;&gt; $OUTFILE\n</code></pre>",
  "messages": [
    {
      "id": "253409",
      "postDate": "12/04/2017 22:29:36",
      "content": "<p>For anyone looking to have a csv of files (and labels of the train set), I wrote a couple of scripts I thought I'd share. For performance reasons, it was quickest to get them from the archive:</p>\n\n<h2>Create a csv of the train files with labels:</h2>\n\n<pre><code>OUTFILE=./input/train_files.csv\necho \"fname,label\" &gt; $OUTFILE\n7za l ./input/train.7z | grep '.wav$' | awk '{print $(NF)}' | awk -F'/' '{print $4,$3}' OFS=\",\" &gt;&gt; $OUTFILE\n</code></pre>\n\n<h2>Create a csv of the test file names:</h2>\n\n<pre><code>OUTFILE=./input/test_files.csv\necho \"fname\" &gt; $OUTFILE\n7za l ./input/test.7z | grep '.wav$' | awk '{print $(NF)}' | cut -f3 -d'/' &gt;&gt; $OUTFILE\n</code></pre>",
      "rawMarkdown": "For anyone looking to have a csv of files (and labels of the train set), I wrote a couple of scripts I thought I'd share. For performance reasons, it was quickest to get them from the archive:\n\n## Create a csv of the train files with labels:\n\n    OUTFILE=./input/train_files.csv\n    echo \"fname,label\" &gt; $OUTFILE\n    7za l ./input/train.7z | grep '.wav$' | awk '{print $(NF)}' | awk -F'/' '{print $4,$3}' OFS=\",\" &gt;&gt; $OUTFILE\n\n## Create a csv of the test file names:\n    OUTFILE=./input/test_files.csv\n    echo \"fname\" &gt; $OUTFILE\n    7za l ./input/test.7z | grep '.wav$' | awk '{print $(NF)}' | cut -f3 -d'/' &gt;&gt; $OUTFILE",
      "votes": null
    },
    {
      "id": "253539",
      "postDate": "12/05/2017 06:33:50",
      "content": "<p>and for PyTorch:</p>\n\n<p>```\ndef find_classes(fullDir):\n    classes = [d for d in os.listdir(fullDir) if os.path.isdir(os.path.join(fullDir, d))]\n    classes.sort()\n    class_to_idx = {classes[i]: i for i in range(len(classes))}\n    num_to_class = dict(zip(range(len(classes)), classes))</p>\n\n<pre><code>train = []\nfor index, label in enumerate(classes):\n    path = fullDir + label + '/'\n    for file in listdir(path):\n        train.append(['{}/{}'.format(label, file), label, index])\n\ndf = pd.DataFrame(train, columns=['file', 'category', 'category_id', ])\n\nreturn classes, class_to_idx, num_to_class, df\n</code></pre>\n\n<p>```</p>",
      "rawMarkdown": "and for PyTorch:\n\n```\ndef find_classes(fullDir):\n    classes = [d for d in os.listdir(fullDir) if os.path.isdir(os.path.join(fullDir, d))]\n    classes.sort()\n    class_to_idx = {classes[i]: i for i in range(len(classes))}\n    num_to_class = dict(zip(range(len(classes)), classes))\n\n    train = []\n    for index, label in enumerate(classes):\n        path = fullDir + label + '/'\n        for file in listdir(path):\n            train.append(['{}/{}'.format(label, file), label, index])\n\n    df = pd.DataFrame(train, columns=['file', 'category', 'category_id', ])\n\n    return classes, class_to_idx, num_to_class, df\n```",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 253539,
      "author_name": "solomonk",
      "author_url": "",
      "post_date": "12/05/2017 06:33:50",
      "content": "<p>and for PyTorch:</p>\n\n<p>```\ndef find_classes(fullDir):\n    classes = [d for d in os.listdir(fullDir) if os.path.isdir(os.path.join(fullDir, d))]\n    classes.sort()\n    class_to_idx = {classes[i]: i for i in range(len(classes))}\n    num_to_class = dict(zip(range(len(classes)), classes))</p>\n\n<pre><code>train = []\nfor index, label in enumerate(classes):\n    path = fullDir + label + '/'\n    for file in listdir(path):\n        train.append(['{}/{}'.format(label, file), label, index])\n\ndf = pd.DataFrame(train, columns=['file', 'category', 'category_id', ])\n\nreturn classes, class_to_idx, num_to_class, df\n</code></pre>\n\n<p>```</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "253409": "For anyone looking to have a csv of files (and labels of the train set), I wrote a couple of scripts I thought I'd share. For performance reasons, it was quickest to get them from the archive:\n\n## Create a csv of the train files with labels:\n\n    OUTFILE=./input/train_files.csv\n    echo \"fname,label\" &gt; $OUTFILE\n    7za l ./input/train.7z | grep '.wav$' | awk '{print $(NF)}' | awk -F'/' '{print $4,$3}' OFS=\",\" &gt;&gt; $OUTFILE\n\n## Create a csv of the test file names:\n    OUTFILE=./input/test_files.csv\n    echo \"fname\" &gt; $OUTFILE\n    7za l ./input/test.7z | grep '.wav$' | awk '{print $(NF)}' | cut -f3 -d'/' &gt;&gt; $OUTFILE",
    "253539": "and for PyTorch:\n\n```\ndef find_classes(fullDir):\n    classes = [d for d in os.listdir(fullDir) if os.path.isdir(os.path.join(fullDir, d))]\n    classes.sort()\n    class_to_idx = {classes[i]: i for i in range(len(classes))}\n    num_to_class = dict(zip(range(len(classes)), classes))\n\n    train = []\n    for index, label in enumerate(classes):\n        path = fullDir + label + '/'\n        for file in listdir(path):\n            train.append(['{}/{}'.format(label, file), label, index])\n\n    df = pd.DataFrame(train, columns=['file', 'category', 'category_id', ])\n\n    return classes, class_to_idx, num_to_class, df\n```"
  },
  "source": "meta"
}