{
  "id": 442691,
  "title": "Looking for Teammates",
  "url": "/competitions/bengaliai-speech/discussion/442691",
  "author_name": "",
  "post_date": "2023-09-23T19:57:08.980231800Z",
  "votes": 1,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I'm looking for 1-2 two teammates from the Gold/Silver Zone. <br>\nFollowing is my current setup.</p>\n<pre><code>     : ai4bharat/indicwav2vec_v1_bengali\n    : \n        : .\n      : % Train + Valid + Fleurs + openslr  + small portion of common voice\n</code></pre>\n<p>I still didn't compute the&nbsp;effects of including other datasets (Fleurs + openslr 37 + tiny amount of common voice). I'll share it in a few days.</p>\n<p>I hope this is useful.</p>",
  "messages": [
    {
      "id": "2453101",
      "postDate": "09/23/2023 19:57:08",
      "content": "<p>I'm looking for 1-2 two teammates from the Gold/Silver Zone. <br>\nFollowing is my current setup.</p>\n<pre><code>     : ai4bharat/indicwav2vec_v1_bengali\n    : \n        : .\n      : % Train + Valid + Fleurs + openslr  + small portion of common voice\n</code></pre>\n<p>I still didn't compute the&nbsp;effects of including other datasets (Fleurs + openslr 37 + tiny amount of common voice). I'll share it in a few days.</p>\n<p>I hope this is useful.</p>",
      "rawMarkdown": "I'm looking for 1-2 two teammates from the Gold/Silver Zone. \nFollowing is my current setup.\n\n```\nModel     : ai4bharat/indicwav2vec_v1_bengali\nEpochs    : 5\nLB        : 0.420\nData      : 20% Train + Valid + Fleurs + openslr 37 + small portion of common voice\n```\n\nI still didn't compute the effects of including other datasets (Fleurs + openslr 37 + tiny amount of common voice). I'll share it in a few days.\n\nI hope this is useful.",
      "votes": null
    },
    {
      "id": "2453269",
      "postDate": "09/23/2023 23:06:50",
      "content": "<p>Do you include all these data into training step by step or all at once. Personally, I have used 40% of competition data once, and the result is disappointing. </p>",
      "rawMarkdown": "Do you include all these data into training step by step or all at once. Personally, I have used 40% of competition data once, and the result is disappointing.",
      "votes": null
    },
    {
      "id": "2453522",
      "postDate": "09/24/2023 05:39:06",
      "content": "<p>All at once and then make a random split. I filtered the training data and used 20% of the filtered data. I will update it</p>",
      "rawMarkdown": "All at once and then make a random split. I filtered the training data and used 20% of the filtered data. I will update it",
      "votes": null
    },
    {
      "id": "2453625",
      "postDate": "09/24/2023 07:43:21",
      "content": "<p>Hi Balaji. I want to try unsupervised training for wav2vec model on unlabeled data. If you are willing, we can cooperate ;) </p>",
      "rawMarkdown": "Hi Balaji. I want to try unsupervised training for wav2vec model on unlabeled data. If you are willing, we can cooperate ;)",
      "votes": null
    },
    {
      "id": "2453631",
      "postDate": "09/24/2023 07:58:05",
      "content": "<p>I have trained for 10 epochs on some of the data you have used, and it gets a similar score as yours. I will probably give the competition dataset another shot. </p>\n<p>I think perhaps you can train on more common_voice, it gets me from 0.445 to 0.428. Have you think of some other way to better fit with the OOD</p>",
      "rawMarkdown": "I have trained for 10 epochs on some of the data you have used, and it gets a similar score as yours. I will probably give the competition dataset another shot. \n\nI think perhaps you can train on more common_voice, it gets me from 0.445 to 0.428. Have you think of some other way to better fit with the OOD",
      "votes": null
    },
    {
      "id": "2454136",
      "postDate": "09/24/2023 15:44:16",
      "content": "<p>I am using the common voice mentioned in the below link<br>\nI am planning to add one or two datasets to handle OOD</p>\n<p><a href=\"https://www.kaggle.com/competitions/bengaliai-speech/discussion/435300\" target=\"_blank\">https://www.kaggle.com/competitions/bengaliai-speech/discussion/435300</a></p>",
      "rawMarkdown": "I am using the common voice mentioned in the below link\nI am planning to add one or two datasets to handle OOD\n\nhttps://www.kaggle.com/competitions/bengaliai-speech/discussion/435300",
      "votes": null
    },
    {
      "id": "2454142",
      "postDate": "09/24/2023 15:46:02",
      "content": "<p>Hi HuBERT, </p>\n<p>My current GPU setup wont be good enough to try unsupervised training. I am planning to focus on adding few more datasets.</p>\n<p>To improve the score, I believe we need to tackle </p>\n<ol>\n<li>Punctuations</li>\n<li>OOD words</li>\n<li>Spelling error</li>\n</ol>\n<p>I believe Language Modeling may be the key</p>",
      "rawMarkdown": "Hi HuBERT, \n\nMy current GPU setup wont be good enough to try unsupervised training. I am planning to focus on adding few more datasets.\n\nTo improve the score, I believe we need to tackle \n1. Punctuations\n2. OOD words\n3. Spelling error\n\nI believe Language Modeling may be the key",
      "votes": null
    },
    {
      "id": "2457523",
      "postDate": "09/27/2023 02:37:40",
      "content": "<p>Hello Balaji,<br>\nI am doing adapter traning on  ai4bharat/indicwav2vec_v1_bengali with 100 % data.<br>\nI have many A100s , and my 10 epoch training is able to finish in one day.  But I haven't made a submission yet.  Are you interested in let me join in your team?</p>\n<blockquote>\n  <p>Hi HuBERT, </p>\n  <p>My current GPU setup wont be good enough to try unsupervised training. I am planning to focus on adding few more datasets.</p>\n  <p>To improve the score, I believe we need to tackle </p>\n  <ol>\n  <li>Punctuations</li>\n  <li>OOD words</li>\n  <li>Spelling error</li>\n  </ol>\n  <p>I believe Language Modeling may be the key</p>\n</blockquote>",
      "rawMarkdown": "Hello Balaji,\nI am doing adapter traning on  ai4bharat/indicwav2vec_v1_bengali with 100 % data.\nI have many A100s , and my 10 epoch training is able to finish in one day.  But I haven't made a submission yet.  Are you interested in let me join in your team?\n\n\n\n> Hi HuBERT, \n> \n> My current GPU setup wont be good enough to try unsupervised training. I am planning to focus on adding few more datasets.\n> \n> To improve the score, I believe we need to tackle \n> 1. Punctuations\n> 2. OOD words\n> 3. Spelling error\n> \n> I believe Language Modeling may be the key",
      "votes": null
    },
    {
      "id": "2457861",
      "postDate": "09/27/2023 07:45:33",
      "content": "<p>I would recommend you to submit. Not just myself, but many others will be interested in working with you if you can achieve Silver or Gold zone.</p>\n<p>All the best </p>",
      "rawMarkdown": "I would recommend you to submit. Not just myself, but many others will be interested in working with you if you can achieve Silver or Gold zone.\n\nAll the best",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2453269,
      "author_name": "renyiwei",
      "author_url": "",
      "post_date": "09/23/2023 23:06:50",
      "content": "<p>Do you include all these data into training step by step or all at once. Personally, I have used 40% of competition data once, and the result is disappointing. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2453522,
          "author_name": "dhakshiin1601",
          "author_url": "",
          "post_date": "09/24/2023 05:39:06",
          "content": "<p>All at once and then make a random split. I filtered the training data and used 20% of the filtered data. I will update it</p>",
          "votes": null,
          "replies": [
            {
              "id": 2453631,
              "author_name": "renyiwei",
              "author_url": "",
              "post_date": "09/24/2023 07:58:05",
              "content": "<p>I have trained for 10 epochs on some of the data you have used, and it gets a similar score as yours. I will probably give the competition dataset another shot. </p>\n<p>I think perhaps you can train on more common_voice, it gets me from 0.445 to 0.428. Have you think of some other way to better fit with the OOD</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2454136,
                  "author_name": "dhakshiin1601",
                  "author_url": "",
                  "post_date": "09/24/2023 15:44:16",
                  "content": "<p>I am using the common voice mentioned in the below link<br>\nI am planning to add one or two datasets to handle OOD</p>\n<p><a href=\"https://www.kaggle.com/competitions/bengaliai-speech/discussion/435300\" target=\"_blank\">https://www.kaggle.com/competitions/bengaliai-speech/discussion/435300</a></p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2453625,
      "author_name": "hubert101",
      "author_url": "",
      "post_date": "09/24/2023 07:43:21",
      "content": "<p>Hi Balaji. I want to try unsupervised training for wav2vec model on unlabeled data. If you are willing, we can cooperate ;) </p>",
      "votes": null,
      "replies": [
        {
          "id": 2454142,
          "author_name": "dhakshiin1601",
          "author_url": "",
          "post_date": "09/24/2023 15:46:02",
          "content": "<p>Hi HuBERT, </p>\n<p>My current GPU setup wont be good enough to try unsupervised training. I am planning to focus on adding few more datasets.</p>\n<p>To improve the score, I believe we need to tackle </p>\n<ol>\n<li>Punctuations</li>\n<li>OOD words</li>\n<li>Spelling error</li>\n</ol>\n<p>I believe Language Modeling may be the key</p>",
          "votes": null,
          "replies": [
            {
              "id": 2457523,
              "author_name": "baibizhe1",
              "author_url": "",
              "post_date": "09/27/2023 02:37:40",
              "content": "<p>Hello Balaji,<br>\nI am doing adapter traning on  ai4bharat/indicwav2vec_v1_bengali with 100 % data.<br>\nI have many A100s , and my 10 epoch training is able to finish in one day.  But I haven't made a submission yet.  Are you interested in let me join in your team?</p>\n<blockquote>\n  <p>Hi HuBERT, </p>\n  <p>My current GPU setup wont be good enough to try unsupervised training. I am planning to focus on adding few more datasets.</p>\n  <p>To improve the score, I believe we need to tackle </p>\n  <ol>\n  <li>Punctuations</li>\n  <li>OOD words</li>\n  <li>Spelling error</li>\n  </ol>\n  <p>I believe Language Modeling may be the key</p>\n</blockquote>",
              "votes": null,
              "replies": [
                {
                  "id": 2457861,
                  "author_name": "dhakshiin1601",
                  "author_url": "",
                  "post_date": "09/27/2023 07:45:33",
                  "content": "<p>I would recommend you to submit. Not just myself, but many others will be interested in working with you if you can achieve Silver or Gold zone.</p>\n<p>All the best </p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2453101": "I'm looking for 1-2 two teammates from the Gold/Silver Zone. \nFollowing is my current setup.\n\n```\nModel     : ai4bharat/indicwav2vec_v1_bengali\nEpochs    : 5\nLB        : 0.420\nData      : 20% Train + Valid + Fleurs + openslr 37 + small portion of common voice\n```\n\nI still didn't compute the effects of including other datasets (Fleurs + openslr 37 + tiny amount of common voice). I'll share it in a few days.\n\nI hope this is useful.",
    "2453269": "Do you include all these data into training step by step or all at once. Personally, I have used 40% of competition data once, and the result is disappointing.",
    "2453522": "All at once and then make a random split. I filtered the training data and used 20% of the filtered data. I will update it",
    "2453625": "Hi Balaji. I want to try unsupervised training for wav2vec model on unlabeled data. If you are willing, we can cooperate ;)",
    "2453631": "I have trained for 10 epochs on some of the data you have used, and it gets a similar score as yours. I will probably give the competition dataset another shot. \n\nI think perhaps you can train on more common_voice, it gets me from 0.445 to 0.428. Have you think of some other way to better fit with the OOD",
    "2454136": "I am using the common voice mentioned in the below link\nI am planning to add one or two datasets to handle OOD\n\nhttps://www.kaggle.com/competitions/bengaliai-speech/discussion/435300",
    "2454142": "Hi HuBERT, \n\nMy current GPU setup wont be good enough to try unsupervised training. I am planning to focus on adding few more datasets.\n\nTo improve the score, I believe we need to tackle \n1. Punctuations\n2. OOD words\n3. Spelling error\n\nI believe Language Modeling may be the key",
    "2457523": "Hello Balaji,\nI am doing adapter traning on  ai4bharat/indicwav2vec_v1_bengali with 100 % data.\nI have many A100s , and my 10 epoch training is able to finish in one day.  But I haven't made a submission yet.  Are you interested in let me join in your team?\n\n\n\n> Hi HuBERT, \n> \n> My current GPU setup wont be good enough to try unsupervised training. I am planning to focus on adding few more datasets.\n> \n> To improve the score, I believe we need to tackle \n> 1. Punctuations\n> 2. OOD words\n> 3. Spelling error\n> \n> I believe Language Modeling may be the key",
    "2457861": "I would recommend you to submit. Not just myself, but many others will be interested in working with you if you can achieve Silver or Gold zone.\n\nAll the best"
  },
  "source": "meta"
}