{
  "id": 146933,
  "title": "Indonesian Text Data  Embeddings?",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/146933",
  "author_name": "",
  "post_date": "2020-04-28T23:13:12.532084600Z",
  "votes": 6,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Guys This task is not related to this competition but i thought this discussion panel would help me.\nI have a task to create text classification system for Indonesian Tweets. 2 things i want to know:-\n1. Can I use Bert Base Uncased multilingual for creating embeddings and should i  Use it or anything else should i use?\n2. How can I extract features from this data(my professor wants) me to do this additionally.\ni have 2 columns ( comment , sentiment) for 8000 tweets.</p>",
  "messages": [
    {
      "id": "825300",
      "postDate": "04/28/2020 23:13:12",
      "content": "<p>Guys This task is not related to this competition but i thought this discussion panel would help me.\nI have a task to create text classification system for Indonesian Tweets. 2 things i want to know:-\n1. Can I use Bert Base Uncased multilingual for creating embeddings and should i  Use it or anything else should i use?\n2. How can I extract features from this data(my professor wants) me to do this additionally.\ni have 2 columns ( comment , sentiment) for 8000 tweets.</p>",
      "rawMarkdown": "Guys This task is not related to this competition but i thought this discussion panel would help me.\nI have a task to create text classification system for Indonesian Tweets. 2 things i want to know:-\n1. Can I use Bert Base Uncased multilingual for creating embeddings and should i  Use it or anything else should i use?\n2. How can I extract features from this data(my professor wants) me to do this additionally.\ni have 2 columns ( comment , sentiment) for 8000 tweets.",
      "votes": null
    },
    {
      "id": "825324",
      "postDate": "04/28/2020 23:53:22",
      "content": "<p>Try to use at first TF-IDF it works with most languages.</p>",
      "rawMarkdown": "Try to use at first TF-IDF it works with most languages.",
      "votes": null
    },
    {
      "id": "825334",
      "postDate": "04/29/2020 00:10:59",
      "content": "<p>OK\nAnd how to tokenise Indonesian <a href=\"/hamditarek\">@hamditarek</a> </p>",
      "rawMarkdown": "OK\nAnd how to tokenise Indonesian @hamditarek",
      "votes": null
    },
    {
      "id": "825902",
      "postDate": "04/29/2020 10:11:35",
      "content": "<p>Hi Sahib,\nYes, you can use Bert Base Uncased Multilingual as it was trained with Indonesian text data too. If you want a more powerful model, you could also use the latest multilingual model XLM-R.</p>",
      "rawMarkdown": "Hi Sahib,\nYes, you can use Bert Base Uncased Multilingual as it was trained with Indonesian text data too. If you want a more powerful model, you could also use the latest multilingual model XLM-R.",
      "votes": null
    },
    {
      "id": "825921",
      "postDate": "04/29/2020 10:28:17",
      "content": "<p>Buddy thank you so much <a href=\"/ilhamfp31\">@ilhamfp31</a>\nCan you suggest any tutorial for the same \nThat will be very helpful </p>",
      "rawMarkdown": "Buddy thank you so much @ilhamfp31\nCan you suggest any tutorial for the same \nThat will be very helpful",
      "votes": null
    },
    {
      "id": "825928",
      "postDate": "04/29/2020 10:34:17",
      "content": "<p>My current undergraduate final project is on Indonesian text classification using multilingual language model like XLM-R and mBERT! I'll publish it on GitHub tomorrow. Right now I'm still finishing it :D</p>\n\n<p>In the meantime, you can look at kernel such as <a href=\"https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta\">this one</a>.</p>",
      "rawMarkdown": "My current undergraduate final project is on Indonesian text classification using multilingual language model like XLM-R and mBERT! I'll publish it on GitHub tomorrow. Right now I'm still finishing it :D\n\nIn the meantime, you can look at kernel such as [this one](https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta).",
      "votes": null
    },
    {
      "id": "826022",
      "postDate": "04/29/2020 11:49:51",
      "content": "<p>Thank you so much <a href=\"/ilhamfp31\">@ilhamfp31</a> ill be looking forward to your reply</p>",
      "rawMarkdown": "Thank you so much @ilhamfp31 ill be looking forward to your reply",
      "votes": null
    },
    {
      "id": "914639",
      "postDate": "07/04/2020 05:41:05",
      "content": "<p>Hi! Sorry for the late reply, here is my GitHub source code: <a href=\"https://github.com/ilhamfp/indonesian-text-classification-multilingual\">https://github.com/ilhamfp/indonesian-text-classification-multilingual</a></p>",
      "rawMarkdown": "Hi! Sorry for the late reply, here is my GitHub source code: https://github.com/ilhamfp/indonesian-text-classification-multilingual",
      "votes": null
    },
    {
      "id": "915851",
      "postDate": "07/05/2020 06:31:52",
      "content": "<p>thanks bro <a href=\"/ilhamfp31\">@ilhamfp31</a> </p>",
      "rawMarkdown": "thanks bro @ilhamfp31",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 825324,
      "author_name": "hamditarek",
      "author_url": "",
      "post_date": "04/28/2020 23:53:22",
      "content": "<p>Try to use at first TF-IDF it works with most languages.</p>",
      "votes": null,
      "replies": [
        {
          "id": 825334,
          "author_name": "sahib12",
          "author_url": "",
          "post_date": "04/29/2020 00:10:59",
          "content": "<p>OK\nAnd how to tokenise Indonesian <a href=\"/hamditarek\">@hamditarek</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 825902,
      "author_name": "ilhamfp31",
      "author_url": "",
      "post_date": "04/29/2020 10:11:35",
      "content": "<p>Hi Sahib,\nYes, you can use Bert Base Uncased Multilingual as it was trained with Indonesian text data too. If you want a more powerful model, you could also use the latest multilingual model XLM-R.</p>",
      "votes": null,
      "replies": [
        {
          "id": 825921,
          "author_name": "sahib12",
          "author_url": "",
          "post_date": "04/29/2020 10:28:17",
          "content": "<p>Buddy thank you so much <a href=\"/ilhamfp31\">@ilhamfp31</a>\nCan you suggest any tutorial for the same \nThat will be very helpful </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 825928,
          "author_name": "ilhamfp31",
          "author_url": "",
          "post_date": "04/29/2020 10:34:17",
          "content": "<p>My current undergraduate final project is on Indonesian text classification using multilingual language model like XLM-R and mBERT! I'll publish it on GitHub tomorrow. Right now I'm still finishing it :D</p>\n\n<p>In the meantime, you can look at kernel such as <a href=\"https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta\">this one</a>.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 826022,
          "author_name": "sahib12",
          "author_url": "",
          "post_date": "04/29/2020 11:49:51",
          "content": "<p>Thank you so much <a href=\"/ilhamfp31\">@ilhamfp31</a> ill be looking forward to your reply</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 914639,
          "author_name": "ilhamfp31",
          "author_url": "",
          "post_date": "07/04/2020 05:41:05",
          "content": "<p>Hi! Sorry for the late reply, here is my GitHub source code: <a href=\"https://github.com/ilhamfp/indonesian-text-classification-multilingual\">https://github.com/ilhamfp/indonesian-text-classification-multilingual</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 915851,
          "author_name": "sahib12",
          "author_url": "",
          "post_date": "07/05/2020 06:31:52",
          "content": "<p>thanks bro <a href=\"/ilhamfp31\">@ilhamfp31</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "825300": "Guys This task is not related to this competition but i thought this discussion panel would help me.\nI have a task to create text classification system for Indonesian Tweets. 2 things i want to know:-\n1. Can I use Bert Base Uncased multilingual for creating embeddings and should i  Use it or anything else should i use?\n2. How can I extract features from this data(my professor wants) me to do this additionally.\ni have 2 columns ( comment , sentiment) for 8000 tweets.",
    "825324": "Try to use at first TF-IDF it works with most languages.",
    "825334": "OK\nAnd how to tokenise Indonesian @hamditarek",
    "825902": "Hi Sahib,\nYes, you can use Bert Base Uncased Multilingual as it was trained with Indonesian text data too. If you want a more powerful model, you could also use the latest multilingual model XLM-R.",
    "825921": "Buddy thank you so much @ilhamfp31\nCan you suggest any tutorial for the same \nThat will be very helpful",
    "825928": "My current undergraduate final project is on Indonesian text classification using multilingual language model like XLM-R and mBERT! I'll publish it on GitHub tomorrow. Right now I'm still finishing it :D\n\nIn the meantime, you can look at kernel such as [this one](https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta).",
    "826022": "Thank you so much @ilhamfp31 ill be looking forward to your reply",
    "914639": "Hi! Sorry for the late reply, here is my GitHub source code: https://github.com/ilhamfp/indonesian-text-classification-multilingual",
    "915851": "thanks bro @ilhamfp31"
  },
  "source": "meta"
}