{
  "id": 672487,
  "title": "Corrupt audio files ",
  "url": "/competitions/dl-sprint-4-0-bengali-long-form-speech-recognition/discussion/672487",
  "author_name": "",
  "post_date": "2026-02-08T15:44:26.515965200Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>what to do if there is a corrupt audio file ? will the hidden test sets contain corrupt files ? what should the model transcribe when corrupt files are found ? </p>",
  "messages": [
    {
      "id": "3403481",
      "postDate": "02/08/2026 15:44:26",
      "content": "<p>what to do if there is a corrupt audio file ? will the hidden test sets contain corrupt files ? what should the model transcribe when corrupt files are found ? </p>",
      "rawMarkdown": "what to do if there is a corrupt audio file ? will the hidden test sets contain corrupt files ? what should the model transcribe when corrupt files are found ?",
      "votes": null
    },
    {
      "id": "3403487",
      "postDate": "02/08/2026 15:50:51",
      "content": "<p>There will be no corrupted files in the hidden set. If you encounter any corrupted files in the training set, possibly train_89, simply discard them and train your model on the remaining data. There is no need to worry about inference on the test or hidden sets. We will ensure their integrity.</p>",
      "rawMarkdown": "There will be no corrupted files in the hidden set. If you encounter any corrupted files in the training set, possibly train_89, simply discard them and train your model on the remaining data. There is no need to worry about inference on the test or hidden sets. We will ensure their integrity.",
      "votes": null
    },
    {
      "id": "3403518",
      "postDate": "02/08/2026 17:05:23",
      "content": "<p>Develop an automatic recognition system that accurately transcribes long-form Bengali audio into readable text.</p>",
      "rawMarkdown": "Develop an automatic recognition system that accurately transcribes long-form Bengali audio into readable text.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3403487,
      "author_name": "shadmantabib",
      "author_url": "",
      "post_date": "02/08/2026 15:50:51",
      "content": "<p>There will be no corrupted files in the hidden set. If you encounter any corrupted files in the training set, possibly train_89, simply discard them and train your model on the remaining data. There is no need to worry about inference on the test or hidden sets. We will ensure their integrity.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3403518,
      "author_name": "thegoanpanda",
      "author_url": "",
      "post_date": "02/08/2026 17:05:23",
      "content": "<p>Develop an automatic recognition system that accurately transcribes long-form Bengali audio into readable text.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3403481": "what to do if there is a corrupt audio file ? will the hidden test sets contain corrupt files ? what should the model transcribe when corrupt files are found ?",
    "3403487": "There will be no corrupted files in the hidden set. If you encounter any corrupted files in the training set, possibly train_89, simply discard them and train your model on the remaining data. There is no need to worry about inference on the test or hidden sets. We will ensure their integrity.",
    "3403518": "Develop an automatic recognition system that accurately transcribes long-form Bengali audio into readable text."
  },
  "source": "meta"
}