{
  "id": 390761,
  "title": "General recommendations to improve your results",
  "url": "/competitions/ml-olympiad-dialectrecognition/discussion/390761",
  "author_name": "Ali",
  "post_date": "2023-02-27T05:57:15.188000",
  "votes": 3,
  "comment_count": 0,
  "views": 0,
  "content": "<p><strong>General recommendations to improve your results:</strong></p>\n<p>1- Try using deep learning language models (One example in the notebook I shared was Arabert: <a href=\"https://huggingface.co/aubmindlab/bert-base-arabert\" target=\"_blank\">https://huggingface.co/aubmindlab/bert-base-arabert</a> ), however, there are other models that surly might perform better ;-)</p>\n<p>2- In my opinion, Deep Models outperform Linear Models (most of the time, unless the training data is small) </p>\n<p>3- Splitting the training data to train/valid sets using the traditional “train_test_split” method is good but NOT the best! try “KFold” the data, this will give better validation and a better understanding of the results. </p>\n<p>4- Don't let High leaderboard scores scares you (mostly above 0.6+ ) including mine :-), most of this scores \"if not all\" scored when the public leaderboard weight was 1% and 99% private board. For example, My 0.68686 score on the leaderboard is really 0.55804 when the splitting is 50% and 50%.</p>\n<p>5- Ensembling different models might lead to better results ( not always, you have to test) </p>\n<p>6- Try different learning rates and try tune “other” parameters.</p>\n<p>7- Try cleaning (normalizing, stemming …) the data, and try the data as is! </p>\n<p>A good start to “get your feet wet” is this notebook which I shared before: <br>\n<a href=\"https://www.kaggle.com/code/asalhi/starter-training-and-infer-using-arabert\" target=\"_blank\">https://www.kaggle.com/code/asalhi/starter-training-and-infer-using-arabert</a></p>\n<p>I will share a higher score notebook later once the competition ends. (which is based on the current notebook) </p>\n<p>Good Luck All :-)</p>",
  "messages": [
    {
      "id": 2160960,
      "postDate": "2023-02-27T05:57:15.190Z",
      "content": "<p><strong>General recommendations to improve your results:</strong></p>\n<p>1- Try using deep learning language models (One example in the notebook I shared was Arabert: <a href=\"https://huggingface.co/aubmindlab/bert-base-arabert\" target=\"_blank\">https://huggingface.co/aubmindlab/bert-base-arabert</a> ), however, there are other models that surly might perform better ;-)</p>\n<p>2- In my opinion, Deep Models outperform Linear Models (most of the time, unless the training data is small) </p>\n<p>3- Splitting the training data to train/valid sets using the traditional “train_test_split” method is good but NOT the best! try “KFold” the data, this will give better validation and a better understanding of the results. </p>\n<p>4- Don't let High leaderboard scores scares you (mostly above 0.6+ ) including mine :-), most of this scores \"if not all\" scored when the public leaderboard weight was 1% and 99% private board. For example, My 0.68686 score on the leaderboard is really 0.55804 when the splitting is 50% and 50%.</p>\n<p>5- Ensembling different models might lead to better results ( not always, you have to test) </p>\n<p>6- Try different learning rates and try tune “other” parameters.</p>\n<p>7- Try cleaning (normalizing, stemming …) the data, and try the data as is! </p>\n<p>A good start to “get your feet wet” is this notebook which I shared before: <br>\n<a href=\"https://www.kaggle.com/code/asalhi/starter-training-and-infer-using-arabert\" target=\"_blank\">https://www.kaggle.com/code/asalhi/starter-training-and-infer-using-arabert</a></p>\n<p>I will share a higher score notebook later once the competition ends. (which is based on the current notebook) </p>\n<p>Good Luck All :-)</p>",
      "rawMarkdown": "**General recommendations to improve your results:**\n\n1- Try using deep learning language models (One example in the notebook I shared was Arabert: https://huggingface.co/aubmindlab/bert-base-arabert ), however, there are other models that surly might perform better ;-)\n\n2- In my opinion, Deep Models outperform Linear Models (most of the time, unless the training data is small) \n\n3- Splitting the training data to train/valid sets using the traditional “train_test_split” method is good but NOT the best! try “KFold” the data, this will give better validation and a better understanding of the results. \n\n4- Don't let High leaderboard scores scares you (mostly above 0.6+ ) including mine :-), most of this scores \"if not all\" scored when the public leaderboard weight was 1% and 99% private board. For example, My 0.68686 score on the leaderboard is really 0.55804 when the splitting is 50% and 50%.\n\n5- Ensembling different models might lead to better results ( not always, you have to test) \n\n6- Try different learning rates and try tune “other” parameters.\n\n7- Try cleaning (normalizing, stemming ...) the data, and try the data as is! \n\nA good start to “get your feet wet” is this notebook which I shared before: \nhttps://www.kaggle.com/code/asalhi/starter-training-and-infer-using-arabert\n\nI will share a higher score notebook later once the competition ends. (which is based on the current notebook) \n\nGood Luck All :-)",
      "votes": 3
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2160960": "**General recommendations to improve your results:**\n\n1- Try using deep learning language models (One example in the notebook I shared was Arabert: https://huggingface.co/aubmindlab/bert-base-arabert ), however, there are other models that surly might perform better ;-)\n\n2- In my opinion, Deep Models outperform Linear Models (most of the time, unless the training data is small) \n\n3- Splitting the training data to train/valid sets using the traditional “train_test_split” method is good but NOT the best! try “KFold” the data, this will give better validation and a better understanding of the results. \n\n4- Don't let High leaderboard scores scares you (mostly above 0.6+ ) including mine :-), most of this scores \"if not all\" scored when the public leaderboard weight was 1% and 99% private board. For example, My 0.68686 score on the leaderboard is really 0.55804 when the splitting is 50% and 50%.\n\n5- Ensembling different models might lead to better results ( not always, you have to test) \n\n6- Try different learning rates and try tune “other” parameters.\n\n7- Try cleaning (normalizing, stemming ...) the data, and try the data as is! \n\nA good start to “get your feet wet” is this notebook which I shared before: \nhttps://www.kaggle.com/code/asalhi/starter-training-and-infer-using-arabert\n\nI will share a higher score notebook later once the competition ends. (which is based on the current notebook) \n\nGood Luck All :-)"
  }
}