{
  "id": 393609,
  "title": "My Solution - Saudi Dialect Recognition (Ensembling Different Models) ",
  "url": "/competitions/ml-olympiad-dialectrecognition/discussion/393609",
  "author_name": "Ali",
  "post_date": "2023-03-10T05:26:45.249000",
  "votes": 2,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hi!<br>\nMany Thanks for this competition</p>\n<p>First of all, I only focused on the text feature in this competition, I thought this or (the voice) were the only essential features since only those features in the dataset will mean that your score is as general as possible.</p>\n<p>For real-world applications, we will not have \"most of the time\" other information such as gender or speaker id, or voice source</p>\n<p><strong>My Solution:</strong></p>\n<p>1- I started with the notebook I shared on Kaggle here : <br>\n<a href=\"https://www.kaggle.com/code/asalhi/starter-training-and-infer-using-arabert\" target=\"_blank\">https://www.kaggle.com/code/asalhi/starter-training-and-infer-using-arabert</a></p>\n<p>2- I tried different Arabic models and the best two were : </p>\n<ul>\n<li>MARBERT: <a href=\"https://huggingface.co/UBC-NLP/MARBERT\" target=\"_blank\">https://huggingface.co/UBC-NLP/MARBERT</a> (which comes in two versions)</li>\n<li>bert-large-arabic by <a href=\"https://www.kaggle.com/alisafaya\" target=\"_blank\">@alisafaya</a> : &nbsp;<a href=\"https://huggingface.co/asafaya/bert-large-arabic\" target=\"_blank\">https://huggingface.co/asafaya/bert-large-arabic</a><br>\nAli Safaya model was the best for this task! it got me a bit higher results than MARABERT. </li>\n</ul>\n<p>3- I extended my Kaggle shared notebook to use \"StratifiedKFold\" with training data, here is the training notebook:<br>\n<a href=\"https://colab.research.google.com/drive/1AXakV7hk6uzoNl2MZOiuKwnsrC95DnB6?usp=sharing\" target=\"_blank\">https://colab.research.google.com/drive/1AXakV7hk6uzoNl2MZOiuKwnsrC95DnB6?usp=sharing</a></p>\n<p>4- I did ensembling between different models mainly MARBERT and bert-large-arabic, however, this didn't improve the score ( I might need to tune more but didn't have the time to do so !)<br>\nHere is the notebook for inference (with ensembling): <br>\n<a href=\"https://colab.research.google.com/drive/1eo8tWjpiM_Z1MpWsbQwTkYWvyerEuw78?usp=sharing\" target=\"_blank\">https://colab.research.google.com/drive/1eo8tWjpiM_Z1MpWsbQwTkYWvyerEuw78?usp=sharing</a></p>\n<p>5- I tried majority voting in ensembling but the results of (summing the outputs for each class from each weight file) were better for Safaya Model, however, worked better for MARBERT in private. </p>\n<p>6- My best (Public and Private) scores (0.56902, 0.62846) for bert-large-arabic by <a href=\"https://www.kaggle.com/alisafaya\" target=\"_blank\">@alisafaya</a> (no ensembling, but training on 99% of the data for 2 epochs)</p>\n<p>7- One of the best results also was Majority Voting with StratifiedKFold training for MARBERTv1 with scores (Public and Private): 0.54851, 0.59328<br>\nThats it! Again I only trained with text! no other features were involved, I wanted to try training with voice but I don't have the experience or the time currently to go deeper with voice :-( </p>\n<p>I am really interested to see other solutions since I see better results than mine :-D So please share! </p>\n<p>My work is done based on my earlier 2nd place winning solution on <strong>Arabic Sentiment Analysis 2021 @ KAUST</strong> <br>\nwhich you can find a copy here:<br>\n<a href=\"https://www.kaggle.com/code/asalhi/arabic-sentiment-analysis-2nd-place-winning-code\" target=\"_blank\">https://www.kaggle.com/code/asalhi/arabic-sentiment-analysis-2nd-place-winning-code</a></p>\n<p><strong>Possible data leaks that will let you climb the leaderboard (public and private):</strong></p>\n<p>1- Depending on features such as <code>FileName</code>, <code>ShowName</code>, <code>SpeakerAge</code> … etc will give you most likely much better results, however I believe this will make the solution baised to certain inputs only, and since the data sources for train and test is the same this most likely will cause a data leak, see <a href=\"https://www.kaggle.com/amjadkhatabi\" target=\"_blank\">@amjadkhatabi</a> thread: <br>\n<a href=\"https://www.kaggle.com/competitions/ml-olympiad-dialectrecognition/discussion/392304\" target=\"_blank\">https://www.kaggle.com/competitions/ml-olympiad-dialectrecognition/discussion/392304</a></p>\n<p>2- Another data leak that will (if you used it correctly) score up to 1.0 ! is the data itself! <br>\nThe training data and test data come from this dataset: <a href=\"https://www.kaggle.com/datasets/sdaiancai/sada2022\" target=\"_blank\">https://www.kaggle.com/datasets/sdaiancai/sada2022</a><br>\nSo we already have the test results :-) </p>\n<p><em>PS:</em> Many Many thanks to SDAIA and SBA for this huge effort, this dataset is great input of the research in this area! Huge Respect :-) </p>\n<p>Also Many Thanks to the organizers of this competition, specially <a href=\"https://www.kaggle.com/ruqiyas\" target=\"_blank\">@ruqiyas</a> ! Great Efforts</p>\n<p>Best Wishes!</p>\n<p>Ali Salhi</p>",
  "messages": [
    {
      "id": 2175709,
      "postDate": "2023-03-10T05:26:45.250Z",
      "content": "<p>Hi!<br>\nMany Thanks for this competition</p>\n<p>First of all, I only focused on the text feature in this competition, I thought this or (the voice) were the only essential features since only those features in the dataset will mean that your score is as general as possible.</p>\n<p>For real-world applications, we will not have \"most of the time\" other information such as gender or speaker id, or voice source</p>\n<p><strong>My Solution:</strong></p>\n<p>1- I started with the notebook I shared on Kaggle here : <br>\n<a href=\"https://www.kaggle.com/code/asalhi/starter-training-and-infer-using-arabert\" target=\"_blank\">https://www.kaggle.com/code/asalhi/starter-training-and-infer-using-arabert</a></p>\n<p>2- I tried different Arabic models and the best two were : </p>\n<ul>\n<li>MARBERT: <a href=\"https://huggingface.co/UBC-NLP/MARBERT\" target=\"_blank\">https://huggingface.co/UBC-NLP/MARBERT</a> (which comes in two versions)</li>\n<li>bert-large-arabic by <a href=\"https://www.kaggle.com/alisafaya\" target=\"_blank\">@alisafaya</a> : &nbsp;<a href=\"https://huggingface.co/asafaya/bert-large-arabic\" target=\"_blank\">https://huggingface.co/asafaya/bert-large-arabic</a><br>\nAli Safaya model was the best for this task! it got me a bit higher results than MARABERT. </li>\n</ul>\n<p>3- I extended my Kaggle shared notebook to use \"StratifiedKFold\" with training data, here is the training notebook:<br>\n<a href=\"https://colab.research.google.com/drive/1AXakV7hk6uzoNl2MZOiuKwnsrC95DnB6?usp=sharing\" target=\"_blank\">https://colab.research.google.com/drive/1AXakV7hk6uzoNl2MZOiuKwnsrC95DnB6?usp=sharing</a></p>\n<p>4- I did ensembling between different models mainly MARBERT and bert-large-arabic, however, this didn't improve the score ( I might need to tune more but didn't have the time to do so !)<br>\nHere is the notebook for inference (with ensembling): <br>\n<a href=\"https://colab.research.google.com/drive/1eo8tWjpiM_Z1MpWsbQwTkYWvyerEuw78?usp=sharing\" target=\"_blank\">https://colab.research.google.com/drive/1eo8tWjpiM_Z1MpWsbQwTkYWvyerEuw78?usp=sharing</a></p>\n<p>5- I tried majority voting in ensembling but the results of (summing the outputs for each class from each weight file) were better for Safaya Model, however, worked better for MARBERT in private. </p>\n<p>6- My best (Public and Private) scores (0.56902, 0.62846) for bert-large-arabic by <a href=\"https://www.kaggle.com/alisafaya\" target=\"_blank\">@alisafaya</a> (no ensembling, but training on 99% of the data for 2 epochs)</p>\n<p>7- One of the best results also was Majority Voting with StratifiedKFold training for MARBERTv1 with scores (Public and Private): 0.54851, 0.59328<br>\nThats it! Again I only trained with text! no other features were involved, I wanted to try training with voice but I don't have the experience or the time currently to go deeper with voice :-( </p>\n<p>I am really interested to see other solutions since I see better results than mine :-D So please share! </p>\n<p>My work is done based on my earlier 2nd place winning solution on <strong>Arabic Sentiment Analysis 2021 @ KAUST</strong> <br>\nwhich you can find a copy here:<br>\n<a href=\"https://www.kaggle.com/code/asalhi/arabic-sentiment-analysis-2nd-place-winning-code\" target=\"_blank\">https://www.kaggle.com/code/asalhi/arabic-sentiment-analysis-2nd-place-winning-code</a></p>\n<p><strong>Possible data leaks that will let you climb the leaderboard (public and private):</strong></p>\n<p>1- Depending on features such as <code>FileName</code>, <code>ShowName</code>, <code>SpeakerAge</code> … etc will give you most likely much better results, however I believe this will make the solution baised to certain inputs only, and since the data sources for train and test is the same this most likely will cause a data leak, see <a href=\"https://www.kaggle.com/amjadkhatabi\" target=\"_blank\">@amjadkhatabi</a> thread: <br>\n<a href=\"https://www.kaggle.com/competitions/ml-olympiad-dialectrecognition/discussion/392304\" target=\"_blank\">https://www.kaggle.com/competitions/ml-olympiad-dialectrecognition/discussion/392304</a></p>\n<p>2- Another data leak that will (if you used it correctly) score up to 1.0 ! is the data itself! <br>\nThe training data and test data come from this dataset: <a href=\"https://www.kaggle.com/datasets/sdaiancai/sada2022\" target=\"_blank\">https://www.kaggle.com/datasets/sdaiancai/sada2022</a><br>\nSo we already have the test results :-) </p>\n<p><em>PS:</em> Many Many thanks to SDAIA and SBA for this huge effort, this dataset is great input of the research in this area! Huge Respect :-) </p>\n<p>Also Many Thanks to the organizers of this competition, specially <a href=\"https://www.kaggle.com/ruqiyas\" target=\"_blank\">@ruqiyas</a> ! Great Efforts</p>\n<p>Best Wishes!</p>\n<p>Ali Salhi</p>",
      "rawMarkdown": "Hi!\nMany Thanks for this competition\n\nFirst of all, I only focused on the text feature in this competition, I thought this or (the voice) were the only essential features since only those features in the dataset will mean that your score is as general as possible.\n\nFor real-world applications, we will not have \"most of the time\" other information such as gender or speaker id, or voice source\n\n**My Solution:**\n\n1- I started with the notebook I shared on Kaggle here : \nhttps://www.kaggle.com/code/asalhi/starter-training-and-infer-using-arabert\n\n2- I tried different Arabic models and the best two were : \n- MARBERT: https://huggingface.co/UBC-NLP/MARBERT (which comes in two versions)\n- bert-large-arabic by @alisafaya :  https://huggingface.co/asafaya/bert-large-arabic\nAli Safaya model was the best for this task! it got me a bit higher results than MARABERT. \n\n3- I extended my Kaggle shared notebook to use \"StratifiedKFold\" with training data, here is the training notebook:\nhttps://colab.research.google.com/drive/1AXakV7hk6uzoNl2MZOiuKwnsrC95DnB6?usp=sharing\n\n4- I did ensembling between different models mainly MARBERT and bert-large-arabic, however, this didn't improve the score ( I might need to tune more but didn't have the time to do so !)\nHere is the notebook for inference (with ensembling): \nhttps://colab.research.google.com/drive/1eo8tWjpiM_Z1MpWsbQwTkYWvyerEuw78?usp=sharing\n\n5- I tried majority voting in ensembling but the results of (summing the outputs for each class from each weight file) were better for Safaya Model, however, worked better for MARBERT in private. \n\n6- My best (Public and Private) scores (0.56902, 0.62846) for bert-large-arabic by @alisafaya (no ensembling, but training on 99% of the data for 2 epochs)\n\n7- One of the best results also was Majority Voting with StratifiedKFold training for MARBERTv1 with scores (Public and Private): 0.54851, 0.59328\nThats it! Again I only trained with text! no other features were involved, I wanted to try training with voice but I don't have the experience or the time currently to go deeper with voice :-( \n\nI am really interested to see other solutions since I see better results than mine :-D So please share! \n\nMy work is done based on my earlier 2nd place winning solution on **Arabic Sentiment Analysis 2021 @ KAUST** \nwhich you can find a copy here:\nhttps://www.kaggle.com/code/asalhi/arabic-sentiment-analysis-2nd-place-winning-code\n\n**Possible data leaks that will let you climb the leaderboard (public and private):**\n\n1- Depending on features such as `FileName `, `ShowName`, `SpeakerAge` ... etc will give you most likely much better results, however I believe this will make the solution baised to certain inputs only, and since the data sources for train and test is the same this most likely will cause a data leak, see @amjadkhatabi thread: \nhttps://www.kaggle.com/competitions/ml-olympiad-dialectrecognition/discussion/392304\n\n2- Another data leak that will (if you used it correctly) score up to 1.0 ! is the data itself! \nThe training data and test data come from this dataset: https://www.kaggle.com/datasets/sdaiancai/sada2022\nSo we already have the test results :-) \n\n*PS:* Many Many thanks to SDAIA and SBA for this huge effort, this dataset is great input of the research in this area! Huge Respect :-) \n\nAlso Many Thanks to the organizers of this competition, specially @ruqiyas ! Great Efforts\n\n\nBest Wishes!\n\nAli Salhi\n \n\n ",
      "votes": 2
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2175709": "Hi!\nMany Thanks for this competition\n\nFirst of all, I only focused on the text feature in this competition, I thought this or (the voice) were the only essential features since only those features in the dataset will mean that your score is as general as possible.\n\nFor real-world applications, we will not have \"most of the time\" other information such as gender or speaker id, or voice source\n\n**My Solution:**\n\n1- I started with the notebook I shared on Kaggle here : \nhttps://www.kaggle.com/code/asalhi/starter-training-and-infer-using-arabert\n\n2- I tried different Arabic models and the best two were : \n- MARBERT: https://huggingface.co/UBC-NLP/MARBERT (which comes in two versions)\n- bert-large-arabic by @alisafaya :  https://huggingface.co/asafaya/bert-large-arabic\nAli Safaya model was the best for this task! it got me a bit higher results than MARABERT. \n\n3- I extended my Kaggle shared notebook to use \"StratifiedKFold\" with training data, here is the training notebook:\nhttps://colab.research.google.com/drive/1AXakV7hk6uzoNl2MZOiuKwnsrC95DnB6?usp=sharing\n\n4- I did ensembling between different models mainly MARBERT and bert-large-arabic, however, this didn't improve the score ( I might need to tune more but didn't have the time to do so !)\nHere is the notebook for inference (with ensembling): \nhttps://colab.research.google.com/drive/1eo8tWjpiM_Z1MpWsbQwTkYWvyerEuw78?usp=sharing\n\n5- I tried majority voting in ensembling but the results of (summing the outputs for each class from each weight file) were better for Safaya Model, however, worked better for MARBERT in private. \n\n6- My best (Public and Private) scores (0.56902, 0.62846) for bert-large-arabic by @alisafaya (no ensembling, but training on 99% of the data for 2 epochs)\n\n7- One of the best results also was Majority Voting with StratifiedKFold training for MARBERTv1 with scores (Public and Private): 0.54851, 0.59328\nThats it! Again I only trained with text! no other features were involved, I wanted to try training with voice but I don't have the experience or the time currently to go deeper with voice :-( \n\nI am really interested to see other solutions since I see better results than mine :-D So please share! \n\nMy work is done based on my earlier 2nd place winning solution on **Arabic Sentiment Analysis 2021 @ KAUST** \nwhich you can find a copy here:\nhttps://www.kaggle.com/code/asalhi/arabic-sentiment-analysis-2nd-place-winning-code\n\n**Possible data leaks that will let you climb the leaderboard (public and private):**\n\n1- Depending on features such as `FileName `, `ShowName`, `SpeakerAge` ... etc will give you most likely much better results, however I believe this will make the solution baised to certain inputs only, and since the data sources for train and test is the same this most likely will cause a data leak, see @amjadkhatabi thread: \nhttps://www.kaggle.com/competitions/ml-olympiad-dialectrecognition/discussion/392304\n\n2- Another data leak that will (if you used it correctly) score up to 1.0 ! is the data itself! \nThe training data and test data come from this dataset: https://www.kaggle.com/datasets/sdaiancai/sada2022\nSo we already have the test results :-) \n\n*PS:* Many Many thanks to SDAIA and SBA for this huge effort, this dataset is great input of the research in this area! Huge Respect :-) \n\nAlso Many Thanks to the organizers of this competition, specially @ruqiyas ! Great Efforts\n\n\nBest Wishes!\n\nAli Salhi\n \n\n "
  }
}