{
  "id": 209844,
  "title": "What are the flac files and how do they relate to tfrecord files?",
  "url": "/competitions/rfcx-species-audio-detection/discussion/209844",
  "author_name": "",
  "post_date": "2021-01-08T18:37:07.281206700Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>This is not completely clear from the description of the data. What are the tfrecord files and how do they relate to the flac files under <code>\\train</code>?</p>",
  "messages": [
    {
      "id": "1144911",
      "postDate": "01/08/2021 18:37:07",
      "content": "<p>This is not completely clear from the description of the data. What are the tfrecord files and how do they relate to the flac files under <code>\\train</code>?</p>",
      "rawMarkdown": "This is not completely clear from the description of the data. What are the tfrecord files and how do they relate to the flac files under `\\train`?",
      "votes": null
    },
    {
      "id": "1144970",
      "postDate": "01/08/2021 19:27:39",
      "content": "<p>In this competition you have an option of using either audio files in FLAC format, or either tfrecords files. Data section has detailed explanation what is encoded in tfrecords files. You might want to read TF documentation to get a info how to process them.</p>\n<blockquote>\n  <p>tfrecords/{train,test} - competition data in the TFRecord format, which includes recording_id, audio_wav (encoded in 16-bit PCM format), and label_info (for train only)</p>\n</blockquote>",
      "rawMarkdown": "In this competition you have an option of using either audio files in FLAC format, or either tfrecords files. Data section has detailed explanation what is encoded in tfrecords files. You might want to read TF documentation to get a info how to process them.\n\n> tfrecords/{train,test} - competition data in the TFRecord format, which includes recording_id, audio_wav (encoded in 16-bit PCM format), and label_info (for train only)",
      "votes": null
    },
    {
      "id": "1145513",
      "postDate": "01/09/2021 07:42:26",
      "content": "<p>I have same questio.<br>\nI think one tfrecords file contains 3min audio. But one train file audio has 1 minute audio. Am I doing something wrong? </p>\n<p>Below is tfrecord reading script.</p>\n<pre><code>def read_labeled_tfrecord(example):\n    LABELED_TFREC_FORMAT = {\n        \"recording_id\"    : tf.io.FixedLenFeature([], tf.string, default_value=''),\n        \"audio_wav\"    : tf.io.FixedLenFeature([], tf.string, default_value=''), \n        \"label_info\"   : tf.io.FixedLenFeature([], tf.string, default_value='') \n    }\n    example = tf.io.parse_single_example(example, LABELED_TFREC_FORMAT)\n    idnum = example['recording_id']\n    audio = decode_audio(example['audio_wav'])    \n    first_split = string_split_semicolon(example['label_info'])\n    remove_quotes = tf.strings.regex_replace(first_split, '\"', \"\") \n    second_split = string_split_comma(remove_quotes)  \n    species_id = tf.gather_nd(second_split, [0, 0])  \n    return audio, species_id, idnum\n</code></pre>",
      "rawMarkdown": "I have same questio.\nI think one tfrecords file contains 3min audio. But one train file audio has 1 minute audio. Am I doing something wrong? \n\nBelow is tfrecord reading script.\n```\ndef read_labeled_tfrecord(example):\n    LABELED_TFREC_FORMAT = {\n        \"recording_id\"    : tf.io.FixedLenFeature([], tf.string, default_value=''),\n        \"audio_wav\"    : tf.io.FixedLenFeature([], tf.string, default_value=''), \n        \"label_info\"   : tf.io.FixedLenFeature([], tf.string, default_value='') \n    }\n    example = tf.io.parse_single_example(example, LABELED_TFREC_FORMAT)\n    idnum = example['recording_id']\n    audio = decode_audio(example['audio_wav'])    \n    first_split = string_split_semicolon(example['label_info'])\n    remove_quotes = tf.strings.regex_replace(first_split, '\"', \"\") \n    second_split = string_split_comma(remove_quotes)  \n    species_id = tf.gather_nd(second_split, [0, 0])  \n    return audio, species_id, idnum\n```",
      "votes": null
    },
    {
      "id": "1145518",
      "postDate": "01/09/2021 07:50:18",
      "content": "<p>I understand. sample rate = 48000 gets same minute.</p>",
      "rawMarkdown": "I understand. sample rate = 48000 gets same minute.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1144970,
      "author_name": "atamazian",
      "author_url": "",
      "post_date": "01/08/2021 19:27:39",
      "content": "<p>In this competition you have an option of using either audio files in FLAC format, or either tfrecords files. Data section has detailed explanation what is encoded in tfrecords files. You might want to read TF documentation to get a info how to process them.</p>\n<blockquote>\n  <p>tfrecords/{train,test} - competition data in the TFRecord format, which includes recording_id, audio_wav (encoded in 16-bit PCM format), and label_info (for train only)</p>\n</blockquote>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1145513,
      "author_name": "rautaki0127",
      "author_url": "",
      "post_date": "01/09/2021 07:42:26",
      "content": "<p>I have same questio.<br>\nI think one tfrecords file contains 3min audio. But one train file audio has 1 minute audio. Am I doing something wrong? </p>\n<p>Below is tfrecord reading script.</p>\n<pre><code>def read_labeled_tfrecord(example):\n    LABELED_TFREC_FORMAT = {\n        \"recording_id\"    : tf.io.FixedLenFeature([], tf.string, default_value=''),\n        \"audio_wav\"    : tf.io.FixedLenFeature([], tf.string, default_value=''), \n        \"label_info\"   : tf.io.FixedLenFeature([], tf.string, default_value='') \n    }\n    example = tf.io.parse_single_example(example, LABELED_TFREC_FORMAT)\n    idnum = example['recording_id']\n    audio = decode_audio(example['audio_wav'])    \n    first_split = string_split_semicolon(example['label_info'])\n    remove_quotes = tf.strings.regex_replace(first_split, '\"', \"\") \n    second_split = string_split_comma(remove_quotes)  \n    species_id = tf.gather_nd(second_split, [0, 0])  \n    return audio, species_id, idnum\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 1145518,
          "author_name": "rautaki0127",
          "author_url": "",
          "post_date": "01/09/2021 07:50:18",
          "content": "<p>I understand. sample rate = 48000 gets same minute.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1144911": "This is not completely clear from the description of the data. What are the tfrecord files and how do they relate to the flac files under `\\train`?",
    "1144970": "In this competition you have an option of using either audio files in FLAC format, or either tfrecords files. Data section has detailed explanation what is encoded in tfrecords files. You might want to read TF documentation to get a info how to process them.\n\n> tfrecords/{train,test} - competition data in the TFRecord format, which includes recording_id, audio_wav (encoded in 16-bit PCM format), and label_info (for train only)",
    "1145513": "I have same questio.\nI think one tfrecords file contains 3min audio. But one train file audio has 1 minute audio. Am I doing something wrong? \n\nBelow is tfrecord reading script.\n```\ndef read_labeled_tfrecord(example):\n    LABELED_TFREC_FORMAT = {\n        \"recording_id\"    : tf.io.FixedLenFeature([], tf.string, default_value=''),\n        \"audio_wav\"    : tf.io.FixedLenFeature([], tf.string, default_value=''), \n        \"label_info\"   : tf.io.FixedLenFeature([], tf.string, default_value='') \n    }\n    example = tf.io.parse_single_example(example, LABELED_TFREC_FORMAT)\n    idnum = example['recording_id']\n    audio = decode_audio(example['audio_wav'])    \n    first_split = string_split_semicolon(example['label_info'])\n    remove_quotes = tf.strings.regex_replace(first_split, '\"', \"\") \n    second_split = string_split_comma(remove_quotes)  \n    species_id = tf.gather_nd(second_split, [0, 0])  \n    return audio, species_id, idnum\n```",
    "1145518": "I understand. sample rate = 48000 gets same minute."
  },
  "source": "meta"
}