{
  "id": 114861,
  "title": "Question Answering Datasets/Leaderboards",
  "url": "/competitions/tensorflow2-question-answering/discussion/114861",
  "author_name": "",
  "post_date": "2019-10-29T18:14:29.930993Z",
  "votes": 25,
  "comment_count": 1,
  "views": 0,
  "content": "<p>There are many Question Answering Leaderboards to check for what are the best architectures currently out there.</p>\n\n<p>A good place to always check first is paperswithcode.com\nOver <a href=\"https://paperswithcode.com/task/question-answering\">here</a> we can see what are the different datasets and the best score on those leaderboards. There are the different SQuAD leaderboards, WikiQA, and NaturalQuestions, the same dataset we are going to use for this competition.</p>\n\n<p>Note that usually question answering refers to the task of finding the correct string from given text/documents that answers a given question. There can be other tasks, like determining an answer from search engine results (SearchQA) or from knowledge bases. </p>\n\n<p>As a side note, if you look and see, there are many other datasets that have some historical significance in the field of QA. Notably, bAbi was an extremely simple question answering dataset with simple synthetic examples along the lines of \"Bobby had the milk. Bobby gave the milk to Mary. Who has the milk?\" It was meant to gauge the understanding capabilities of different algorithms. It was a relatively hard task for the simple LSTM-based neural networks of 2015. </p>\n\n<p>Another commonly used QA dataset several years ago was TrecQA. This was developed as part of the Text Retrevial Conference (TREC) challenges. Similar to some of the common datasets now, this dataset had question documents, and the task was to find the string in the documents that best answered the given question. </p>\n\n<p>In 2016, Rajpurkar published \"Stanford Question Answering Dataset (SQuAD), a new reading comprehension dataset consisting of 100,000+ questions posed by crowdworkers on a set of Wikipedia articles.\" Each data point contained a paragraph of text from Wikipedia, a question, and an answer coming from this paragraph. This dataset is now one of the biggest question answering and language comprehension dataset currently available. In 2018, SQuAD 2.0 was also released, which now contained unanswerable questions that could not be answered with the text available in the paragraph. The leaderboard is available <a href=\"https://rajpurkar.github.io/SQuAD-explorer/\">here</a>. This again was a hard task, but the advent of Transformer architectures has lead to superhuman performance on this dataset. Some of the top performers include Google's ALBERT, XLNet, and Facebook's RoBERTa.</p>\n\n<p>Now let's look more carefully into Google's Natural Questions dataset, the same dataset for this competition. Natural Questions is somewhat harder and is similar to the previously mentioned SearchQA task. In this task, entire Wikipedia articles is given, and real questions that users have searched on Google are given. The task is then to predict long-form or short-form answers. It is possible that the answer is not available on the given Wikipedia page as well. The long-form answer is like a bounding box over the Wikipedia page where the answer is present. The short-form answer is a much shorter answer of a couple words. The dataset and leaderboard is available <a href=\"https://ai.google.com/research/NaturalQuestions\">here</a>. Currently, BERT-based models are again doing the best.</p>\n\n<p>I hope this helps!</p>",
  "messages": [
    {
      "id": "660874",
      "postDate": "10/29/2019 18:14:29",
      "content": "<p>There are many Question Answering Leaderboards to check for what are the best architectures currently out there.</p>\n\n<p>A good place to always check first is paperswithcode.com\nOver <a href=\"https://paperswithcode.com/task/question-answering\">here</a> we can see what are the different datasets and the best score on those leaderboards. There are the different SQuAD leaderboards, WikiQA, and NaturalQuestions, the same dataset we are going to use for this competition.</p>\n\n<p>Note that usually question answering refers to the task of finding the correct string from given text/documents that answers a given question. There can be other tasks, like determining an answer from search engine results (SearchQA) or from knowledge bases. </p>\n\n<p>As a side note, if you look and see, there are many other datasets that have some historical significance in the field of QA. Notably, bAbi was an extremely simple question answering dataset with simple synthetic examples along the lines of \"Bobby had the milk. Bobby gave the milk to Mary. Who has the milk?\" It was meant to gauge the understanding capabilities of different algorithms. It was a relatively hard task for the simple LSTM-based neural networks of 2015. </p>\n\n<p>Another commonly used QA dataset several years ago was TrecQA. This was developed as part of the Text Retrevial Conference (TREC) challenges. Similar to some of the common datasets now, this dataset had question documents, and the task was to find the string in the documents that best answered the given question. </p>\n\n<p>In 2016, Rajpurkar published \"Stanford Question Answering Dataset (SQuAD), a new reading comprehension dataset consisting of 100,000+ questions posed by crowdworkers on a set of Wikipedia articles.\" Each data point contained a paragraph of text from Wikipedia, a question, and an answer coming from this paragraph. This dataset is now one of the biggest question answering and language comprehension dataset currently available. In 2018, SQuAD 2.0 was also released, which now contained unanswerable questions that could not be answered with the text available in the paragraph. The leaderboard is available <a href=\"https://rajpurkar.github.io/SQuAD-explorer/\">here</a>. This again was a hard task, but the advent of Transformer architectures has lead to superhuman performance on this dataset. Some of the top performers include Google's ALBERT, XLNet, and Facebook's RoBERTa.</p>\n\n<p>Now let's look more carefully into Google's Natural Questions dataset, the same dataset for this competition. Natural Questions is somewhat harder and is similar to the previously mentioned SearchQA task. In this task, entire Wikipedia articles is given, and real questions that users have searched on Google are given. The task is then to predict long-form or short-form answers. It is possible that the answer is not available on the given Wikipedia page as well. The long-form answer is like a bounding box over the Wikipedia page where the answer is present. The short-form answer is a much shorter answer of a couple words. The dataset and leaderboard is available <a href=\"https://ai.google.com/research/NaturalQuestions\">here</a>. Currently, BERT-based models are again doing the best.</p>\n\n<p>I hope this helps!</p>",
      "rawMarkdown": "There are many Question Answering Leaderboards to check for what are the best architectures currently out there.\n\nA good place to always check first is paperswithcode.com\nOver [here](https://paperswithcode.com/task/question-answering) we can see what are the different datasets and the best score on those leaderboards. There are the different SQuAD leaderboards, WikiQA, and NaturalQuestions, the same dataset we are going to use for this competition.\n\nNote that usually question answering refers to the task of finding the correct string from given text/documents that answers a given question. There can be other tasks, like determining an answer from search engine results (SearchQA) or from knowledge bases. \n\nAs a side note, if you look and see, there are many other datasets that have some historical significance in the field of QA. Notably, bAbi was an extremely simple question answering dataset with simple synthetic examples along the lines of \"Bobby had the milk. Bobby gave the milk to Mary. Who has the milk?\" It was meant to gauge the understanding capabilities of different algorithms. It was a relatively hard task for the simple LSTM-based neural networks of 2015. \n\nAnother commonly used QA dataset several years ago was TrecQA. This was developed as part of the Text Retrevial Conference (TREC) challenges. Similar to some of the common datasets now, this dataset had question documents, and the task was to find the string in the documents that best answered the given question. \n\nIn 2016, Rajpurkar published \"Stanford Question Answering Dataset (SQuAD), a new reading comprehension dataset consisting of 100,000+ questions posed by crowdworkers on a set of Wikipedia articles.\" Each data point contained a paragraph of text from Wikipedia, a question, and an answer coming from this paragraph. This dataset is now one of the biggest question answering and language comprehension dataset currently available. In 2018, SQuAD 2.0 was also released, which now contained unanswerable questions that could not be answered with the text available in the paragraph. The leaderboard is available [here](https://rajpurkar.github.io/SQuAD-explorer/). This again was a hard task, but the advent of Transformer architectures has lead to superhuman performance on this dataset. Some of the top performers include Google's ALBERT, XLNet, and Facebook's RoBERTa.\n\nNow let's look more carefully into Google's Natural Questions dataset, the same dataset for this competition. Natural Questions is somewhat harder and is similar to the previously mentioned SearchQA task. In this task, entire Wikipedia articles is given, and real questions that users have searched on Google are given. The task is then to predict long-form or short-form answers. It is possible that the answer is not available on the given Wikipedia page as well. The long-form answer is like a bounding box over the Wikipedia page where the answer is present. The short-form answer is a much shorter answer of a couple words. The dataset and leaderboard is available [here](https://ai.google.com/research/NaturalQuestions). Currently, BERT-based models are again doing the best.\n\nI hope this helps!",
      "votes": null
    },
    {
      "id": "914139",
      "postDate": "07/03/2020 16:37:52",
      "content": "<p>Thanks man, it helps a lot! 🙌 </p>",
      "rawMarkdown": "Thanks man, it helps a lot! 🙌",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 914139,
      "author_name": "robertgutierrez1",
      "author_url": "",
      "post_date": "07/03/2020 16:37:52",
      "content": "<p>Thanks man, it helps a lot! 🙌 </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "660874": "There are many Question Answering Leaderboards to check for what are the best architectures currently out there.\n\nA good place to always check first is paperswithcode.com\nOver [here](https://paperswithcode.com/task/question-answering) we can see what are the different datasets and the best score on those leaderboards. There are the different SQuAD leaderboards, WikiQA, and NaturalQuestions, the same dataset we are going to use for this competition.\n\nNote that usually question answering refers to the task of finding the correct string from given text/documents that answers a given question. There can be other tasks, like determining an answer from search engine results (SearchQA) or from knowledge bases. \n\nAs a side note, if you look and see, there are many other datasets that have some historical significance in the field of QA. Notably, bAbi was an extremely simple question answering dataset with simple synthetic examples along the lines of \"Bobby had the milk. Bobby gave the milk to Mary. Who has the milk?\" It was meant to gauge the understanding capabilities of different algorithms. It was a relatively hard task for the simple LSTM-based neural networks of 2015. \n\nAnother commonly used QA dataset several years ago was TrecQA. This was developed as part of the Text Retrevial Conference (TREC) challenges. Similar to some of the common datasets now, this dataset had question documents, and the task was to find the string in the documents that best answered the given question. \n\nIn 2016, Rajpurkar published \"Stanford Question Answering Dataset (SQuAD), a new reading comprehension dataset consisting of 100,000+ questions posed by crowdworkers on a set of Wikipedia articles.\" Each data point contained a paragraph of text from Wikipedia, a question, and an answer coming from this paragraph. This dataset is now one of the biggest question answering and language comprehension dataset currently available. In 2018, SQuAD 2.0 was also released, which now contained unanswerable questions that could not be answered with the text available in the paragraph. The leaderboard is available [here](https://rajpurkar.github.io/SQuAD-explorer/). This again was a hard task, but the advent of Transformer architectures has lead to superhuman performance on this dataset. Some of the top performers include Google's ALBERT, XLNet, and Facebook's RoBERTa.\n\nNow let's look more carefully into Google's Natural Questions dataset, the same dataset for this competition. Natural Questions is somewhat harder and is similar to the previously mentioned SearchQA task. In this task, entire Wikipedia articles is given, and real questions that users have searched on Google are given. The task is then to predict long-form or short-form answers. It is possible that the answer is not available on the given Wikipedia page as well. The long-form answer is like a bounding box over the Wikipedia page where the answer is present. The short-form answer is a much shorter answer of a couple words. The dataset and leaderboard is available [here](https://ai.google.com/research/NaturalQuestions). Currently, BERT-based models are again doing the best.\n\nI hope this helps!",
    "914139": "Thanks man, it helps a lot! 🙌"
  },
  "source": "meta"
}