{"cells":[{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import numpy as np\nimport pandas as pd\n\n# Any results you write to the current directory are saved as output.\nimport os\nimport gc\nimport matplotlib.pyplot as plt\nimport json\n\nfrom collections import Counter","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Input data files are available in the \"../input/\" directory.\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","collapsed":true,"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":false},"cell_type":"markdown","source":"# File Description\n\n* simplified-nq-train.jsonl - the training data, in newline-delimited JSON format.\n* simplified-nq-kaggle-test.jsonl - the test data, in newline-delimited JSON format.\n* sample_submission.csv - a sample submission file in the correct format"},{"metadata":{},"cell_type":"markdown","source":"# Data fields\n* document_text - the text of the article in question (with some HTML tags to provide document structure). The text can be tokenized by splitting on whitespace.\n* question_text - the question to be answered\n* long_answer_candidates - a JSON array containing all of the plausible long answers.\n* annotations - a JSON array containing all of the correct long + short answers. Only provided for train.\n* document_url - the URL for the full article. Provided for informational purposes only. This is NOT the simplified version of the article so indices from this cannot be used directly. The content may also no longer match the html used to generate document_text. Only provided for train.\n* example_id - unique ID for the sample."},{"metadata":{},"cell_type":"markdown","source":"# Load Data"},{"metadata":{"trusted":true},"cell_type":"code","source":"path = '/kaggle/input/tensorflow2-question-answering/'\ntrain_path = 'simplified-nq-train.jsonl'\ntest_path = 'simplified-nq-test.jsonl'\nsample_submission_path = 'sample_submission.csv'","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_line_count = 0\nwith open(path+train_path) as f:\n    for line in f:\n        line = json.loads(line)\n        train_line_count += 1\ntrain_line_count","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The dataset is huge, for exploration purpose we are going perform the exploratory analysis over a sample of the dataset. Let's read the training data and extract a sample (hopefully the dataset is shuffled so that the first records are random)"},{"metadata":{"trusted":true},"cell_type":"code","source":"def read_data(path, sample = True, chunksize = 30000):\n    if sample == True:\n        df = []\n        with open(path, 'rt') as reader:\n            for i in range(chunksize):\n                df.append(json.loads(reader.readline()))\n        df = pd.DataFrame(df)\n    else:\n        df = pd.read_json(path, orient = 'records', lines = True)\n        gc.collect()\n    return df","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train = read_data(path+train_path, sample = True, chunksize=100000)\nprint(\"train shape\", train.shape)\ntrain[:10]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test = read_data(path+test_path, sample = False)\nprint(\"test shape\", test.shape)\ntest[:10]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sample_submission = pd.read_csv(path + sample_submission_path)\nprint(\"Sample submission shape\", sample_submission.shape)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The submission file has 692 rows, this means that for each row in the test set (there are 346) we have to predict both the short and long answer.\nThe short answer may also be a Yes or No answer."},{"metadata":{"trusted":true},"cell_type":"code","source":"sample_submission[:10]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"markdown","source":"# Missing Values\nWhat columns we have and whether or not we have missing values"},{"metadata":{"trusted":true},"cell_type":"code","source":"def missing_values(df):\n    df = pd.DataFrame(df.isnull().sum()).reset_index()\n    df.columns = ['features', 'n_missing_values']\n    return df","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"missing_values(train)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"missing_values(test)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Great we don't have missing values."},{"metadata":{},"cell_type":"markdown","source":"# What the train data looks like"},{"metadata":{"trusted":true},"cell_type":"code","source":"train.columns","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Question text\nThe question as a text string. Note that it is not tokenized."},{"metadata":{"trusted":true},"cell_type":"code","source":"# Question text\ntrain.loc[0, 'question_text']","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Document text\n'The text of the article in question (with some HTML tags to provide document structure). The text can be tokenized by splitting on whitespace.'\nThis is a huge wikipedia text, where we need to find the answer for the previous question\nNote that they give the tokenization method they expect us to use."},{"metadata":{"trusted":true},"cell_type":"code","source":"train.loc[0, 'document_text']","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train.loc[0, 'document_text'].split()[:100]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Long answer candidates\n'A JSON array containing all of the plausible long answers'.\nThere seem to be a lot of these. Each one has a start_token and an end_token in the text, to delimit where the long answer is.\nThere is also a 'top_level' field."},{"metadata":{"trusted":true},"cell_type":"code","source":"train.loc[0, 'long_answer_candidates'][:10]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Annotations\n'A JSON array containing all of the correct long + short answers. Only provided for train.'\nThis defines the target we are training on."},{"metadata":{"trusted":true},"cell_type":"code","source":"train['annotations'][:10]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Make a dataframe to accumulate answer types\nanswer_num_annotations = train.annotations.apply(lambda x: len(x))\nCounter(answer_num_annotations)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"So there is only ever one annotation in the train data.\nLet us explore how well the annotations are structured. First we will find out what things are in the annotations."},{"metadata":{"trusted":true},"cell_type":"code","source":"train['annotations'][0]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* This is our target variable. In this case this is telling us that our long answer starts in indices 1952 and end at indices 2019.\n* Also, we have a short answer that starts a indices 1960 and end at indices 1969.\n* In this example we dont have a yes or no answer"},{"metadata":{"trusted":true},"cell_type":"code","source":"set().union(*(d[0].keys() for d in train['annotations']))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"So we know about all our keys."},{"metadata":{},"cell_type":"markdown","source":"## Annotations Analysis\nAnnotations is a JSON array containing all of the correct long + short answers. Only provided for train.\nWe will build a reformulated data frame around the annotations."},{"metadata":{"trusted":true},"cell_type":"code","source":"answer_summary = train['example_id'].to_frame()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"answer_summary['annotation_id'] = train.annotations.apply(lambda x: x[0]['annotation_id'])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"answer_summary[:10]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Yes-No"},{"metadata":{"trusted":true},"cell_type":"code","source":"answer_yes_no = train.annotations.apply(lambda x: x[0]['yes_no_answer'])\nyes_no_answer_counts = Counter(answer_yes_no)\nyes_no_answer_counts","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"ks = [k for k in yes_no_answer_counts.keys()]\nvs = [yes_no_answer_counts[k] for k in ks]\n\nplt.bar(ks, vs)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"percent_yes_no = 1 - yes_no_answer_counts['NONE'] / sum(yes_no_answer_counts.values())\nprint(percent_yes_no, \"of the questions have yes/no answers given\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Clean column and save in summary"},{"metadata":{"trusted":true},"cell_type":"code","source":"answer_yes_no_cleaned = answer_yes_no.apply(lambda x: None if x == 'NONE' else x)\nanswer_summary['has_yes_no'] = answer_yes_no_cleaned.apply(lambda x: x is not None)\nanswer_summary['yes_no'] = answer_yes_no_cleaned","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Short Answers"},{"metadata":{"trusted":true},"cell_type":"code","source":"train.annotations.apply(lambda x: x[0]['short_answers'])[:10]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"answer_short = train.annotations.apply(lambda x: [(y['start_token'], y['end_token']) for y in x[0]['short_answers']])\nnum_short_answers = answer_short.apply(lambda x: len(x))\nshort_answer_counts = Counter(num_short_answers)\nshort_answer_counts","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"ks = [k for k in short_answer_counts.keys()]\nks.sort()\nvs = [short_answer_counts[k] for k in ks]\n\nplt.bar(ks, vs)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"So a large majority of the questions do not have any short answers. Of the remaining, most have only one short answer. The longest list is one question that has 21 answers."},{"metadata":{"trusted":true},"cell_type":"code","source":"percent_short = 1 - short_answer_counts[0] / sum(short_answer_counts.values())\nprint(percent_short, \"of the questions have at least one short answer\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Clean columns and save in summary"},{"metadata":{"trusted":true},"cell_type":"code","source":"answer_summary['has_short_answers'] = num_short_answers.apply(lambda x: x>0)\nanswer_summary['num_short_answers'] = num_short_answers\nanswer_summary['answer_short'] = answer_short.apply(lambda x: x if len(x) > 0 else None)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Long Answers"},{"metadata":{"trusted":true},"cell_type":"code","source":"train.annotations.apply(lambda x: x[0]['long_answer'])[:20]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train.loc[0, 'annotations'][0]['long_answer']","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"So we have a start token and an end token, similar to the short answers.\nThe candidate index is the index into the list of candidate answers."},{"metadata":{"trusted":true},"cell_type":"code","source":"answer_long = train.annotations.apply(lambda x: (x[0]['long_answer']['start_token'], x[0]['long_answer']['end_token']))\nanswer_long_cleaned = answer_long.apply(lambda x: x if x != (-1, -1) else None)\nnum_long_answers = answer_long_cleaned.apply(lambda x: 1 if x else 0)\nlong_answer_counts = Counter(num_long_answers)\nlong_answer_counts","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"ks = [k for k in long_answer_counts.keys()]\nks.sort()\nvs = [long_answer_counts[k] for k in ks]\n\nplt.bar(ks, vs)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"percent_long = 1 - long_answer_counts[0] / sum(long_answer_counts.values())\nprint (percent_long, \"of the questions have at least one long answer\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Update the summary."},{"metadata":{"trusted":true},"cell_type":"code","source":"answer_summary['has_long_answer'] = num_long_answers.apply(lambda x: x>0)\nanswer_summary['num_long_answers'] = num_long_answers\nanswer_summary['answer_long'] = answer_long_cleaned","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Also look at the candidate indices."},{"metadata":{"trusted":true},"cell_type":"code","source":"candidate_indices = train.annotations.apply(lambda x: (x[0]['long_answer']['candidate_index']))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"answer_summary['long_candidate_index'] = candidate_indices","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"answer_summary[:10]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Summary"},{"metadata":{},"cell_type":"markdown","source":"How many of the questions have some answer?"},{"metadata":{"trusted":true},"cell_type":"code","source":"summary = answer_summary.apply(lambda row: \n                               True if (row['has_yes_no'] or row['has_short_answers'] or row['has_long_answer'])\n                               else False, axis=1)\nsummary[:10]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"Counter(summary)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Looks like about half our training set consists of questions for which the document does not contain an answer."},{"metadata":{"trusted":true},"cell_type":"code","source":"answer_summary[\"summary\"] = summary","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Count how many have both a yes-no and at least one short answer; also how many have both short and long answers."},{"metadata":{"trusted":true},"cell_type":"code","source":"answer_summary.groupby(['has_yes_no', 'has_short_answers', 'has_long_answer']).size().reset_index()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"So in the training data there is either a yes-no answer or short answers, but not both."},{"metadata":{},"cell_type":"raw","source":"Here is the final summary dataframe"},{"metadata":{"trusted":true},"cell_type":"code","source":"answer_summary[:10]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":1}