{
  "id": 452341,
  "title": "Dataset Explanation ",
  "url": "/competitions/automatic-speech-recognition-asr/discussion/452341",
  "author_name": "",
  "post_date": "2023-11-01T20:03:54.451117600Z",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hi everyone, </p>\n<p>Can somebody explain the dataset? What is the purpose competition and how should the submission file be?</p>\n<p>Thank you in advance,</p>",
  "messages": [
    {
      "id": "2508586",
      "postDate": "11/01/2023 20:03:54",
      "content": "<p>Hi everyone, </p>\n<p>Can somebody explain the dataset? What is the purpose competition and how should the submission file be?</p>\n<p>Thank you in advance,</p>",
      "rawMarkdown": "Hi everyone, \n\nCan somebody explain the dataset? What is the purpose competition and how should the submission file be?\n\nThank you in advance,",
      "votes": null
    },
    {
      "id": "2510334",
      "postDate": "11/02/2023 22:15:34",
      "content": "<p>Hi Gamze! This competition is actually a homework assignment for the Introduction to Deep Learning course being run at CMU and students are given additional writeups and a starter notebook, hence a lot of details are not mentioned here.</p>\n<p>The dataset has speech MFCCs and their corresponding phoneme transcripts as NumPy arrays. The task is to train a sequence model using CTC loss to predict the phonemes for test MFCCs. The submission file has the phonemes converted into ARPAbet representation and concatenated in the form of a string. A sample submission file is present in the 'test' directory.</p>\n<p>Code to create the submission file and other helper code is provided in the starter notebook that is given to students. We have not yet started making these competitions complete as a standalone for external participants. But we will look to do this in future iterations of the course.</p>",
      "rawMarkdown": "Hi Gamze! This competition is actually a homework assignment for the Introduction to Deep Learning course being run at CMU and students are given additional writeups and a starter notebook, hence a lot of details are not mentioned here.\n\nThe dataset has speech MFCCs and their corresponding phoneme transcripts as NumPy arrays. The task is to train a sequence model using CTC loss to predict the phonemes for test MFCCs. The submission file has the phonemes converted into ARPAbet representation and concatenated in the form of a string. A sample submission file is present in the 'test' directory.\n\nCode to create the submission file and other helper code is provided in the starter notebook that is given to students. We have not yet started making these competitions complete as a standalone for external participants. But we will look to do this in future iterations of the course.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2510334,
      "author_name": "harshit2997cmu",
      "author_url": "",
      "post_date": "11/02/2023 22:15:34",
      "content": "<p>Hi Gamze! This competition is actually a homework assignment for the Introduction to Deep Learning course being run at CMU and students are given additional writeups and a starter notebook, hence a lot of details are not mentioned here.</p>\n<p>The dataset has speech MFCCs and their corresponding phoneme transcripts as NumPy arrays. The task is to train a sequence model using CTC loss to predict the phonemes for test MFCCs. The submission file has the phonemes converted into ARPAbet representation and concatenated in the form of a string. A sample submission file is present in the 'test' directory.</p>\n<p>Code to create the submission file and other helper code is provided in the starter notebook that is given to students. We have not yet started making these competitions complete as a standalone for external participants. But we will look to do this in future iterations of the course.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2508586": "Hi everyone, \n\nCan somebody explain the dataset? What is the purpose competition and how should the submission file be?\n\nThank you in advance,",
    "2510334": "Hi Gamze! This competition is actually a homework assignment for the Introduction to Deep Learning course being run at CMU and students are given additional writeups and a starter notebook, hence a lot of details are not mentioned here.\n\nThe dataset has speech MFCCs and their corresponding phoneme transcripts as NumPy arrays. The task is to train a sequence model using CTC loss to predict the phonemes for test MFCCs. The submission file has the phonemes converted into ARPAbet representation and concatenated in the form of a string. A sample submission file is present in the 'test' directory.\n\nCode to create the submission file and other helper code is provided in the starter notebook that is given to students. We have not yet started making these competitions complete as a standalone for external participants. But we will look to do this in future iterations of the course."
  },
  "source": "meta"
}