{
  "id": 435610,
  "title": "Do I finetune wav2vec2 with text that has punctuations or do I remove punctuations from training dataset?",
  "url": "/competitions/bengaliai-speech/discussion/435610",
  "author_name": "",
  "post_date": "2023-08-30T06:20:43.432034500Z",
  "votes": 3,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Here was another <a href=\"url\" target=\"_blank\">https://www.kaggle.com/competitions/bengaliai-speech/discussion/429246</a> with this but I am still confused.<br>\nQ1: <strong>Do we need to submit our predictions with or without  punctuations?</strong> Do organizers using test set with or without punctuations to evaluate our model?<br>\nQ2: <strong>My model producing text without punctuations, do I need to figure out how to add punctuations?</strong></p>\n<p>The sample given in the competition  has no quotations, commas etc.<br>\nThanks<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3600075%2F7dcbf45d914aa34cb0062b2e14210038%2FScreen%20Shot%202023-08-29%20at%2011.12.13%20PM.png?generation=1693376046110972&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "2415131",
      "postDate": "08/30/2023 06:20:43",
      "content": "<p>Here was another <a href=\"url\" target=\"_blank\">https://www.kaggle.com/competitions/bengaliai-speech/discussion/429246</a> with this but I am still confused.<br>\nQ1: <strong>Do we need to submit our predictions with or without  punctuations?</strong> Do organizers using test set with or without punctuations to evaluate our model?<br>\nQ2: <strong>My model producing text without punctuations, do I need to figure out how to add punctuations?</strong></p>\n<p>The sample given in the competition  has no quotations, commas etc.<br>\nThanks<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3600075%2F7dcbf45d914aa34cb0062b2e14210038%2FScreen%20Shot%202023-08-29%20at%2011.12.13%20PM.png?generation=1693376046110972&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Here was another [https://www.kaggle.com/competitions/bengaliai-speech/discussion/429246](url) with this but I am still confused.\nQ1: **Do we need to submit our predictions with or without  punctuations?** Do organizers using test set with or without punctuations to evaluate our model?\nQ2: **My model producing text without punctuations, do I need to figure out how to add punctuations?**\n\nThe sample given in the competition  has no quotations, commas etc.\nThanks\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3600075%2F7dcbf45d914aa34cb0062b2e14210038%2FScreen%20Shot%202023-08-29%20at%2011.12.13%20PM.png?generation=1693376046110972&alt=media)",
      "votes": null
    },
    {
      "id": "2415271",
      "postDate": "08/30/2023 08:04:33",
      "content": "<p>The punctuations are not removed from the test set. But they are normalized, so make sure to normalize your train and validation set. Regarding training Wav2Vec2, it's better to remove punctuations as no accoustic feature corresponds with the punctuations. You have to figure out how to punctuate the output of Wav2Vec2. A naive way to do this is adding '।' (dari) at the end. Almost all public notebooks already utilized this trick. 'Dari' is equivalent to full stop in English. </p>",
      "rawMarkdown": "The punctuations are not removed from the test set. But they are normalized, so make sure to normalize your train and validation set. Regarding training Wav2Vec2, it's better to remove punctuations as no accoustic feature corresponds with the punctuations. You have to figure out how to punctuate the output of Wav2Vec2. A naive way to do this is adding '।' (dari) at the end. Almost all public notebooks already utilized this trick. 'Dari' is equivalent to full stop in English.",
      "votes": null
    },
    {
      "id": "2422235",
      "postDate": "09/03/2023 20:08:18",
      "content": "<p><a href=\"https://www.kaggle.com/umongsain\" target=\"_blank\">@umongsain</a> what do you mean by \"punctuations are normalized\"?</p>",
      "rawMarkdown": "umongsain what do you mean by \"punctuations are normalized\"?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2415271,
      "author_name": "umongsain",
      "author_url": "",
      "post_date": "08/30/2023 08:04:33",
      "content": "<p>The punctuations are not removed from the test set. But they are normalized, so make sure to normalize your train and validation set. Regarding training Wav2Vec2, it's better to remove punctuations as no accoustic feature corresponds with the punctuations. You have to figure out how to punctuate the output of Wav2Vec2. A naive way to do this is adding '।' (dari) at the end. Almost all public notebooks already utilized this trick. 'Dari' is equivalent to full stop in English. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2422235,
          "author_name": "manwithaflower",
          "author_url": "",
          "post_date": "09/03/2023 20:08:18",
          "content": "<p><a href=\"https://www.kaggle.com/umongsain\" target=\"_blank\">@umongsain</a> what do you mean by \"punctuations are normalized\"?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2415131": "Here was another [https://www.kaggle.com/competitions/bengaliai-speech/discussion/429246](url) with this but I am still confused.\nQ1: **Do we need to submit our predictions with or without  punctuations?** Do organizers using test set with or without punctuations to evaluate our model?\nQ2: **My model producing text without punctuations, do I need to figure out how to add punctuations?**\n\nThe sample given in the competition  has no quotations, commas etc.\nThanks\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3600075%2F7dcbf45d914aa34cb0062b2e14210038%2FScreen%20Shot%202023-08-29%20at%2011.12.13%20PM.png?generation=1693376046110972&alt=media)",
    "2415271": "The punctuations are not removed from the test set. But they are normalized, so make sure to normalize your train and validation set. Regarding training Wav2Vec2, it's better to remove punctuations as no accoustic feature corresponds with the punctuations. You have to figure out how to punctuate the output of Wav2Vec2. A naive way to do this is adding '।' (dari) at the end. Almost all public notebooks already utilized this trick. 'Dari' is equivalent to full stop in English.",
    "2422235": "umongsain what do you mean by \"punctuations are normalized\"?"
  },
  "source": "meta"
}