{
  "id": 447969,
  "title": "90th Place Solution for the Bengali.AI Speech Recognition",
  "url": "/competitions/bengaliai-speech/discussion/447969",
  "author_name": "Yurnero",
  "post_date": "2023-10-18T01:30:47.208000",
  "votes": 3,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I would like to thank the hosts for the wonderful task! I feel so bad that I had only 3 days for this competition and didn't join earlier.</p>\n<p><strong>Context</strong></p>\n<ul>\n<li>Business context: <a href=\"https://www.kaggle.com/competitions/bengaliai-speech/overview\" target=\"_blank\">https://www.kaggle.com/competitions/bengaliai-speech/overview</a></li>\n<li>Data context: <a href=\"https://www.kaggle.com/competitions/bengaliai-speech/data\" target=\"_blank\">https://www.kaggle.com/competitions/bengaliai-speech/data</a></li>\n</ul>\n<p><strong>Overview of the Approach</strong></p>\n<p>As almost everyone else I used pretrained on bengali speech model. Main parameters for fine-tuning are listed below</p>\n<pre><code>=,\n=,\n=,\n=,\n=,\n=,\n=,\n=,\n=,\n=e-,\n=\n</code></pre>\n<p>and the following configuration was used</p>\n<pre><code>.freeze_feature_extractor()\n.freeze_feature_encoder()\n</code></pre>\n<p><strong>Details of the submission</strong></p>\n<p>It took me 16h to fine-tune a model on 10% of train data for 3 epochs on my GPU — thats where I felt the need for time.</p>\n<p>I found that increasing <code>beam_width</code> in decoder from default to 2048 lead to 0.003 boost both on public and private (and inference time is ~8h, so we cant really increase further).</p>\n<p><strong>Sources</strong></p>\n<p><a href=\"https://www.kaggle.com/code/takanashihumbert/bengali-sr-wav2vec-v1-bengali-training\" target=\"_blank\">https://www.kaggle.com/code/takanashihumbert/bengali-sr-wav2vec-v1-bengali-training</a> — notebook that I used for fine-tuning<br>\n<a href=\"https://www.kaggle.com/competitions/bengaliai-speech/discussion/435300\" target=\"_blank\">https://www.kaggle.com/competitions/bengaliai-speech/discussion/435300</a> — datasets and model checkpoints for resource efficient training</p>",
  "messages": [
    {
      "id": 2486508,
      "postDate": "2023-10-18T01:30:47.210Z",
      "content": "<p>I would like to thank the hosts for the wonderful task! I feel so bad that I had only 3 days for this competition and didn't join earlier.</p>\n<p><strong>Context</strong></p>\n<ul>\n<li>Business context: <a href=\"https://www.kaggle.com/competitions/bengaliai-speech/overview\" target=\"_blank\">https://www.kaggle.com/competitions/bengaliai-speech/overview</a></li>\n<li>Data context: <a href=\"https://www.kaggle.com/competitions/bengaliai-speech/data\" target=\"_blank\">https://www.kaggle.com/competitions/bengaliai-speech/data</a></li>\n</ul>\n<p><strong>Overview of the Approach</strong></p>\n<p>As almost everyone else I used pretrained on bengali speech model. Main parameters for fine-tuning are listed below</p>\n<pre><code>=,\n=,\n=,\n=,\n=,\n=,\n=,\n=,\n=,\n=e-,\n=\n</code></pre>\n<p>and the following configuration was used</p>\n<pre><code>.freeze_feature_extractor()\n.freeze_feature_encoder()\n</code></pre>\n<p><strong>Details of the submission</strong></p>\n<p>It took me 16h to fine-tune a model on 10% of train data for 3 epochs on my GPU — thats where I felt the need for time.</p>\n<p>I found that increasing <code>beam_width</code> in decoder from default to 2048 lead to 0.003 boost both on public and private (and inference time is ~8h, so we cant really increase further).</p>\n<p><strong>Sources</strong></p>\n<p><a href=\"https://www.kaggle.com/code/takanashihumbert/bengali-sr-wav2vec-v1-bengali-training\" target=\"_blank\">https://www.kaggle.com/code/takanashihumbert/bengali-sr-wav2vec-v1-bengali-training</a> — notebook that I used for fine-tuning<br>\n<a href=\"https://www.kaggle.com/competitions/bengaliai-speech/discussion/435300\" target=\"_blank\">https://www.kaggle.com/competitions/bengaliai-speech/discussion/435300</a> — datasets and model checkpoints for resource efficient training</p>",
      "rawMarkdown": "I would like to thank the hosts for the wonderful task! I feel so bad that I had only 3 days for this competition and didn't join earlier.\n\n**Context**\n\n- Business context: https://www.kaggle.com/competitions/bengaliai-speech/overview\n- Data context: https://www.kaggle.com/competitions/bengaliai-speech/data\n\n**Overview of the Approach**\n\nAs almost everyone else I used pretrained on bengali speech model. Main parameters for fine-tuning are listed below\n\n    lr_scheduler_type='cosine',\n    weight_decay=0.01,\n    num_train_epochs=3,\n    per_device_train_batch_size=2,\n    gradient_accumulation_steps=8,\n    fp16=True,\n    save_steps=500,\n    eval_steps=500,\n    logging_steps=500,\n    learning_rate=1e-5,\n    warmup_ratio=0.1\n\nand the following configuration was used\n\n    model.freeze_feature_extractor()\n    model.freeze_feature_encoder()\n\n**Details of the submission**\n\nIt took me 16h to fine-tune a model on 10% of train data for 3 epochs on my GPU — thats where I felt the need for time.\n\nI found that increasing `beam_width` in decoder from default to 2048 lead to 0.003 boost both on public and private (and inference time is ~8h, so we cant really increase further).\n\n**Sources**\n\nhttps://www.kaggle.com/code/takanashihumbert/bengali-sr-wav2vec-v1-bengali-training — notebook that I used for fine-tuning\nhttps://www.kaggle.com/competitions/bengaliai-speech/discussion/435300 — datasets and model checkpoints for resource efficient training",
      "votes": 2
    },
    {
      "id": 2486628,
      "postDate": "2023-10-18T04:35:32.887Z",
      "content": "<p>Changing <code>beam_width</code> to 2048 in this <a href=\"https://www.kaggle.com/code/takanashihumbert/bengali-sr-wav2vec-v1-bengali-inference\" target=\"_blank\">public notebook</a> provided a private score = 0.518. This score is about 85th place in the private LB.</p>",
      "rawMarkdown": "Changing `beam_width` to 2048 in this [public notebook](https://www.kaggle.com/code/takanashihumbert/bengali-sr-wav2vec-v1-bengali-inference) provided a private score = 0.518. This score is about 85th place in the private LB.",
      "replies": [
        {
          "id": 2486631,
          "postDate": "2023-10-18T04:40:14.973Z",
          "content": "<p>Hm, thats very interesting. Because this was my second submission and I got 0.521</p>",
          "rawMarkdown": "Hm, thats very interesting. Because this was my second submission and I got 0.521",
          "replies": [
            {
              "id": 2486638,
              "postDate": "2023-10-18T04:50:44.587Z",
              "content": "<p>Did you use the 4th version of the notebook?</p>",
              "rawMarkdown": "Did you use the 4th version of the notebook?"
            },
            {
              "id": 2486646,
              "postDate": "2023-10-18T04:58:13.937Z",
              "content": "<p>Yeah. The one with 0.439 public score</p>",
              "rawMarkdown": "Yeah. The one with 0.439 public score"
            },
            {
              "id": 2486662,
              "postDate": "2023-10-18T05:13:20.767Z",
              "content": "<p>I checked it. The modification also included reducing the <code>batch_size</code> from 16 to 1</p>",
              "rawMarkdown": "I checked it. The modification also included reducing the `batch_size` from 16 to 1",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2487869,
      "postDate": "2023-10-18T20:01:18.037Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2486628,
      "author_name": "Kristina Kalbasiuk",
      "author_url": "",
      "post_date": "2023-10-18T04:35:32.887000",
      "content": "<p>Changing <code>beam_width</code> to 2048 in this <a href=\"https://www.kaggle.com/code/takanashihumbert/bengali-sr-wav2vec-v1-bengali-inference\" target=\"_blank\">public notebook</a> provided a private score = 0.518. This score is about 85th place in the private LB.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2486631,
          "author_name": "Yurnero",
          "author_url": "",
          "post_date": "2023-10-18T04:40:14.973000",
          "content": "<p>Hm, thats very interesting. Because this was my second submission and I got 0.521</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2486638,
              "author_name": "Kristina Kalbasiuk",
              "author_url": "",
              "post_date": "2023-10-18T04:50:44.587000",
              "content": "<p>Did you use the 4th version of the notebook?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2486646,
              "author_name": "Yurnero",
              "author_url": "",
              "post_date": "2023-10-18T04:58:13.937000",
              "content": "<p>Yeah. The one with 0.439 public score</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2486662,
              "author_name": "Kristina Kalbasiuk",
              "author_url": "",
              "post_date": "2023-10-18T05:13:20.767000",
              "content": "<p>I checked it. The modification also included reducing the <code>batch_size</code> from 16 to 1</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2487869,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-10-18T20:01:18.037000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2486508": "I would like to thank the hosts for the wonderful task! I feel so bad that I had only 3 days for this competition and didn't join earlier.\n\n**Context**\n\n- Business context: https://www.kaggle.com/competitions/bengaliai-speech/overview\n- Data context: https://www.kaggle.com/competitions/bengaliai-speech/data\n\n**Overview of the Approach**\n\nAs almost everyone else I used pretrained on bengali speech model. Main parameters for fine-tuning are listed below\n\n    lr_scheduler_type='cosine',\n    weight_decay=0.01,\n    num_train_epochs=3,\n    per_device_train_batch_size=2,\n    gradient_accumulation_steps=8,\n    fp16=True,\n    save_steps=500,\n    eval_steps=500,\n    logging_steps=500,\n    learning_rate=1e-5,\n    warmup_ratio=0.1\n\nand the following configuration was used\n\n    model.freeze_feature_extractor()\n    model.freeze_feature_encoder()\n\n**Details of the submission**\n\nIt took me 16h to fine-tune a model on 10% of train data for 3 epochs on my GPU — thats where I felt the need for time.\n\nI found that increasing `beam_width` in decoder from default to 2048 lead to 0.003 boost both on public and private (and inference time is ~8h, so we cant really increase further).\n\n**Sources**\n\nhttps://www.kaggle.com/code/takanashihumbert/bengali-sr-wav2vec-v1-bengali-training — notebook that I used for fine-tuning\nhttps://www.kaggle.com/competitions/bengaliai-speech/discussion/435300 — datasets and model checkpoints for resource efficient training",
    "2486628": "Changing `beam_width` to 2048 in this [public notebook](https://www.kaggle.com/code/takanashihumbert/bengali-sr-wav2vec-v1-bengali-inference) provided a private score = 0.518. This score is about 85th place in the private LB.",
    "2487869": ""
  }
}