{
  "id": 434651,
  "title": "Can I create a dataset and use it, only audio is in public domain?",
  "url": "/competitions/bengaliai-speech/discussion/434651",
  "author_name": "Blue",
  "post_date": "2023-08-25T23:54:53.503000",
  "votes": 0,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hi, I want to train model from scratch. <br>\nDon't remember where but I saw a github project with Bengali audio files. It didn't have transcripts. </p>\n<h5>Q1: Can I use those audio files and transcribe them using YellowKing's model and use that as dataset to my model?</h5>",
  "messages": [
    {
      "id": 2425655,
      "postDate": "2023-09-06T05:10:46.083Z",
      "content": "<p>Yes you can label data, but you need to publicly disclose all the data you are using.</p>",
      "rawMarkdown": "Yes you can label data, but you need to publicly disclose all the data you are using.",
      "votes": 2,
      "replies": [
        {
          "id": 2429909,
          "postDate": "2023-09-08T22:41:50.077Z",
          "content": "<p>Just to clarify, do we need to share the public data that we use before the end of competition or after it? </p>",
          "rawMarkdown": "Just to clarify, do we need to share the public data that we use before the end of competition or after it? ",
          "votes": 1,
          "replies": [
            {
              "id": 2435060,
              "postDate": "2023-09-12T17:07:31.147Z",
              "content": "<p>Normally, pseudo-labelled or synthetic generated data etc. only needs to be shared after the competition has ended. Otherwise there is little incentive for participants to make the effort. <a href=\"https://www.kaggle.com/reasat\" target=\"_blank\">@reasat</a> correct me please, if I am wrong.</p>",
              "rawMarkdown": "Normally, pseudo-labelled or synthetic generated data etc. only needs to be shared after the competition has ended. Otherwise there is little incentive for participants to make the effort. @reasat correct me please, if I am wrong.",
              "votes": 3
            },
            {
              "id": 2435140,
              "postDate": "2023-09-12T18:05:53.823Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2435143,
              "postDate": "2023-09-12T18:10:26.207Z",
              "content": "<p>Great question. Let me consult with the team and get back to you.</p>",
              "rawMarkdown": "Great question. Let me consult with the team and get back to you.",
              "votes": 2
            },
            {
              "id": 2435158,
              "postDate": "2023-09-12T18:23:24.077Z",
              "content": "<p>Yes, what <a href=\"https://www.kaggle.com/benbla\" target=\"_blank\">@benbla</a> has stated is correct. We would require the data to be disclosed after the competition. </p>",
              "rawMarkdown": "Yes, what @benbla has stated is correct. We would require the data to be disclosed after the competition. ",
              "votes": 3
            },
            {
              "id": 2436931,
              "postDate": "2023-09-13T21:30:07.870Z",
              "content": "<p>Tahsin, <a href=\"https://www.kaggle.com/reasat\" target=\"_blank\">@reasat</a> I found audios that have creative commons license and they DON'T have transcripts. Not sure yet how but I am planning to create transcripts using some API.  So, I only need to make my dataset public if I win the competition?</p>",
              "rawMarkdown": "Tahsin, @reasat I found audios that have creative commons license and they DON'T have transcripts. Not sure yet how but I am planning to create transcripts using some API.  So, I only need to make my dataset public if I win the competition?"
            },
            {
              "id": 2436940,
              "postDate": "2023-09-13T21:55:30.337Z",
              "content": "<p>Yes that is correct.</p>",
              "rawMarkdown": "Yes that is correct.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2408944,
      "postDate": "2023-08-25T23:54:53.503Z",
      "content": "<p>Hi, I want to train model from scratch. <br>\nDon't remember where but I saw a github project with Bengali audio files. It didn't have transcripts. </p>\n<h5>Q1: Can I use those audio files and transcribe them using YellowKing's model and use that as dataset to my model?</h5>",
      "rawMarkdown": "Hi, I want to train model from scratch. \nDon't remember where but I saw a github project with Bengali audio files. It didn't have transcripts. \n##### Q1: Can I use those audio files and transcribe them using YellowKing's model and use that as dataset to my model? \n\n\n"
    }
  ],
  "comments": [
    {
      "id": 2425655,
      "author_name": "Tahsin",
      "author_url": "",
      "post_date": "2023-09-06T05:10:46.083000",
      "content": "<p>Yes you can label data, but you need to publicly disclose all the data you are using.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2429909,
          "author_name": "Man of the year",
          "author_url": "",
          "post_date": "2023-09-08T22:41:50.077000",
          "content": "<p>Just to clarify, do we need to share the public data that we use before the end of competition or after it? </p>",
          "votes": 1,
          "replies": [
            {
              "id": 2435060,
              "author_name": "Benedikt Droste",
              "author_url": "",
              "post_date": "2023-09-12T17:07:31.147000",
              "content": "<p>Normally, pseudo-labelled or synthetic generated data etc. only needs to be shared after the competition has ended. Otherwise there is little incentive for participants to make the effort. <a href=\"https://www.kaggle.com/reasat\" target=\"_blank\">@reasat</a> correct me please, if I am wrong.</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2435140,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-09-12T18:05:53.823000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2435143,
              "author_name": "Tahsin",
              "author_url": "",
              "post_date": "2023-09-12T18:10:26.207000",
              "content": "<p>Great question. Let me consult with the team and get back to you.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2435158,
              "author_name": "Tahsin",
              "author_url": "",
              "post_date": "2023-09-12T18:23:24.077000",
              "content": "<p>Yes, what <a href=\"https://www.kaggle.com/benbla\" target=\"_blank\">@benbla</a> has stated is correct. We would require the data to be disclosed after the competition. </p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2436931,
              "author_name": "Blue",
              "author_url": "",
              "post_date": "2023-09-13T21:30:07.870000",
              "content": "<p>Tahsin, <a href=\"https://www.kaggle.com/reasat\" target=\"_blank\">@reasat</a> I found audios that have creative commons license and they DON'T have transcripts. Not sure yet how but I am planning to create transcripts using some API.  So, I only need to make my dataset public if I win the competition?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2436940,
              "author_name": "Tahsin",
              "author_url": "",
              "post_date": "2023-09-13T21:55:30.337000",
              "content": "<p>Yes that is correct.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2425655": "Yes you can label data, but you need to publicly disclose all the data you are using.",
    "2408944": "Hi, I want to train model from scratch. \nDon't remember where but I saw a github project with Bengali audio files. It didn't have transcripts. \n##### Q1: Can I use those audio files and transcribe them using YellowKing's model and use that as dataset to my model? \n\n\n"
  }
}