{"cells":[{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load in \n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the \"../input/\" directory.\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# Any results you write to the current directory are saved as output.\nimport os\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nimport matplotlib.patches as patches\nfrom plotly import tools, subplots\nimport plotly.offline as py\npy.init_notebook_mode(connected = True)\nimport plotly.graph_objs as go\nimport plotly.express as px\npd.set_option('max_columns', 1000)\nfrom bokeh.models import Panel, Tabs\nfrom bokeh.io import output_notebook, show\nfrom bokeh.plotting import figure\nimport lightgbm as lgb\nimport plotly.figure_factory as ff\nimport gc\nfrom sklearn.model_selection import KFold\nfrom sklearn.preprocessing import LabelEncoder\nimport json\nfrom keras.preprocessing import text, sequence\nfrom sklearn.feature_extraction.text import CountVectorizer\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","collapsed":true,"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":false},"cell_type":"markdown","source":"# File Description\n\n* simplified-nq-train.jsonl - the training data, in newline-delimited JSON format.\n* simplified-nq-kaggle-test.jsonl - the test data, in newline-delimited JSON format.\n* sample_submission.csv - a sample submission file in the correct format"},{"metadata":{},"cell_type":"markdown","source":"# Data fields\n* document_text - the text of the article in question (with some HTML tags to provide document structure). The text can be tokenized by splitting on whitespace.\n* question_text - the question to be answered\n* long_answer_candidates - a JSON array containing all of the plausible long answers.\n* annotations - a JSON array containing all of the correct long + short answers. Only provided for train.\n* document_url - the URL for the full article. Provided for informational purposes only. This is NOT the simplified version of the article so indices from this cannot be used directly. The content may also no longer match the html used to generate document_text. Only provided for train.\n* example_id - unique ID for the sample."},{"metadata":{},"cell_type":"markdown","source":"Let's check the submission file to understand better what we need to predict\n\n# Submission File\nFor each ID in the test set, you must predict a) a set of start:end token indices, b) a YES/NO answer if applicable (short answers ONLY), or c) a BLANK answer if no prediction can be made. The file should contain a header and have the following format:\n\n* -7853356005143141653_long,6:18\n* -7853356005143141653_short,YES\n* -545833482873225036_long,105:200\n* -545833482873225036_short,\n* -6998273848279890840_long,\n* -6998273848279890840_short,NO\n\nInteresting :)."},{"metadata":{},"cell_type":"markdown","source":"# Evaluation¶\nSubmissions are evaluated using micro F1 between the predicted and expected answers. Predicted long and short answers must match exactly the token indices of one of the ground truth labels ((or match YES/NO if the question has a yes/no short answer). There may be up to five labels for long answers, and more for short. If no answer applies, leave the prediction blank/null.\n\nThe metric in this competition diverges from the original metric in two key respects: 1) short and long answer formats do not receive separate scores, but are instead combined into a micro F1 score across both formats, and 2) this competition's metric does not use confidence scores to find an optimal threshold for predictions.\n\n"},{"metadata":{},"cell_type":"markdown","source":"# Load Data\n\nThe dataset is huge, for exploration purpose we are going perform the exploratory analysis over a sample of the dataset. Let's read the training data and extract a sample (hopefully the dataset is shuffled so that the first records are random)"},{"metadata":{"trusted":true},"cell_type":"code","source":"path = '/kaggle/input/tensorflow2-question-answering/'\ntrain_path = 'simplified-nq-train.jsonl'\ntest_path = 'simplified-nq-test.jsonl'\nsample_submission_path = 'sample_submission.csv'\n\ndef read_data(path, sample = True, chunksize = 30000):\n    if sample == True:\n        df = []\n        with open(path, 'rt') as reader:\n            for i in range(chunksize):\n                df.append(json.loads(reader.readline()))\n        df = pd.DataFrame(df)\n        print('Our sampled dataset have {} rows and {} columns'.format(df.shape[0], df.shape[1]))\n    else:\n        df = pd.read_json(path, orient = 'records', lines = True)\n        print('Our dataset have {} rows and {} columns'.format(df.shape[0], df.shape[1]))\n        gc.collect()\n    return df\n\ntrain = read_data(path+train_path, sample = True)\ntest = read_data(path+test_path, sample = False)\ntrain.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sample_submission = pd.read_csv(path + sample_submission_path)\nprint('Our sample submission have {} rows'.format(sample_submission.shape[0]))\nsample_submission.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"markdown","source":"First let's explore if we have missing values"},{"metadata":{"trusted":true},"cell_type":"markdown","source":"# Missing Values"},{"metadata":{"trusted":true},"cell_type":"code","source":"def missing_values(df):\n    df = pd.DataFrame(df.isnull().sum()).reset_index()\n    df.columns = ['features', 'n_missing_values']\n    return df\nmissing_values(train)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"missing_values(test)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Great we don't have missing values."},{"metadata":{},"cell_type":"markdown","source":"# Logic\n\nThis extructure is not easy to understand, let's explore the first line of our train set to understand the logic of this dataset."},{"metadata":{"trusted":true},"cell_type":"code","source":"question_text_0 = train.loc[0, 'question_text']\nquestion_text_0","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"I believe this is the main question you need to respond."},{"metadata":{"trusted":true},"cell_type":"code","source":"document_text_0 = train.loc[0, 'document_text'].split()\n\" \".join(document_text_0[:800])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"> So in the first column we have a huge wikipedia text. This is where we need to find the answer for the previous question"},{"metadata":{"trusted":true},"cell_type":"code","source":"long_answer_candidates_0 = train.loc[0, 'long_answer_candidates']\nlong_answer_candidates_0[0:10]","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"This are all the possibles long answers ranges. In other words they give you the start indices and last indices of all the possibles long answers in the document text columns that could answer the question."},{"metadata":{"trusted":true},"cell_type":"code","source":"annotations_0 = train['annotations'][0][0]\nannotations_0","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* This is our target variable. In this case this is telling us that our long answer starts in indices 1952 and end at indices 2019.\n* Also, we have a short answer that starts at indices 1960 and end at indices 1969.\n* In this example we dont have a yes or no answer\n* If you check the submission file we have 692 rows, this means that for each row in the test set we have to predict the short and long answer\n* Sometime long and short answer are not available, in this case it's possible that we have a Yes or No answer for the short answer.\n\nLets check the entire logic of the first line of our train set"},{"metadata":{"trusted":true},"cell_type":"code","source":"print('Our question is : ', question_text_0)\nprint('Our short answer is : ', \" \".join(document_text_0[annotations_0['short_answers'][0]['start_token']:annotations_0['short_answers'][0]['end_token']]))\nprint('Our long answer is : ', \" \".join(document_text_0[annotations_0['long_answer']['start_token']:annotations_0['long_answer']['end_token']]))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* Now that we understand the main logic let's check the distributions of our target variable. Remeber that our test it's going to expand to 692 rows because for each question we need to answer the long and short answer.\n\n# Target Variable Exploration"},{"metadata":{"trusted":true},"cell_type":"code","source":"yes_no_answer = []\nfor i in range(len(train)):\n    yes_no_answer.append(train['annotations'][i][0]['yes_no_answer'])\nyes_no_answer = pd.DataFrame({'yes_no_answer': yes_no_answer})\n    \ndef bar_plot(df, column, title, width, height, n, get_count = True):\n    if get_count == True:\n        cnt_srs = df[column].value_counts(normalize = True)[:n]\n    else:\n        cnt_srs = df\n        \n    trace = go.Bar(\n        x = cnt_srs.index,\n        y = cnt_srs.values,\n        marker = dict(\n            color = '#1E90FF',\n        ),\n    )\n\n    layout = go.Layout(\n        title = go.layout.Title(\n            text = title,\n            x = 0.5\n        ),\n        font = dict(size = 14),\n        width = width,\n        height = height,\n    )\n\n    data = [trace]\n    fig = go.Figure(data = data, layout = layout)\n    py.iplot(fig, filename = 'bar_plot')\nbar_plot(yes_no_answer, 'yes_no_answer', 'Yes No Answer Distribution', 800, 500, 3)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* 98.7% is None\n* The amount of observations that are YES and NO only sum 1.3%!"},{"metadata":{"trusted":true},"cell_type":"code","source":"# this function extract the short answers and fill a dataframe\ndef extract_target_variable(df, short = True):\n    if short:\n        short_answer = []\n        for i in range(len(df)):\n            short = df['annotations'][i][0]['short_answers']\n            if short == []:\n                yes_no = df['annotations'][i][0]['yes_no_answer']\n                if yes_no == 'NO' or yes_no == 'YES':\n                    short_answer.append(yes_no)\n                else:\n                    short_answer.append('EMPTY')\n            else:\n                short = short[0]\n                st = short['start_token']\n                et = short['end_token']\n                short_answer.append(f'{st}'+':'+f'{et}')\n        short_answer = pd.DataFrame({'short_answer': short_answer})\n        return short_answer\n    else:\n        long_answer = []\n        for i in range(len(df)):\n            long = df['annotations'][i][0]['long_answer']\n            if long['start_token'] == -1:\n                long_answer.append('EMPTY')\n            else:\n                st = long['start_token']\n                et = long['end_token']\n                long_answer.append(f'{st}'+':'+f'{et}')\n        long_answer = pd.DataFrame({'long_answer': long_answer})\n        return long_answer\n        \nshort_answer = extract_target_variable(train)\nshort_answer.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"short_answer['type'] = short_answer['short_answer'].copy()\nshort_answer.loc[(short_answer['short_answer']!='EMPTY') & (short_answer['short_answer']!='YES') & (short_answer['short_answer']!='NO'), 'type'] =  'TEXT'\nbar_plot(short_answer, 'type', 'Short Answer Distribution', 800, 500, 10)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Short Answer Results\n\n* We have 63.47% of the observations with a empty text\n* We have 35.23% of the observations with a start and end token result\n* We have the same distribution for YES and NO from the previous plot"},{"metadata":{"trusted":true},"cell_type":"code","source":"long_answer = extract_target_variable(train, False)\nlong_answer.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"long_answer['type'] = long_answer['long_answer'].copy()\nlong_answer.loc[(long_answer['long_answer']!='EMPTY'), 'type'] =  'TEXT'\nbar_plot(long_answer, 'type', 'Long Answer Distribution', 800, 500, 10)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Long Answer Results\n* We have 50.16% of the observations empty\n* We have 49.84% of the observarions with a start and end token result"},{"metadata":{},"cell_type":"markdown","source":"# Question Explorations\n\nLet's explore our question_text column which tell us the question that we want to answer with a segment of the document text\n\n* Count the number of words and check distribution\n* Most common words"},{"metadata":{"trusted":true},"cell_type":"code","source":"def count_word_frequency(series, top = 0, bot = 20):\n    cv = CountVectorizer()   \n    cv_fit = cv.fit_transform(series)    \n    word_list = cv.get_feature_names(); \n    count_list = cv_fit.toarray().sum(axis=0)\n    frequency = pd.DataFrame({'Word': word_list, 'Frequency': count_list})\n    frequency.sort_values(['Frequency'], ascending = False, inplace = True)\n    frequency['Percentage'] = frequency['Frequency']/frequency['Frequency'].sum()\n    frequency.drop('Frequency', inplace = True, axis = 1)\n    frequency['Percentage'] = frequency['Percentage'].round(3)\n    frequency = frequency.iloc[top:bot]\n    frequency.set_index('Word', inplace = True)\n    bar_plot(pd.Series(frequency['Percentage']), 'Percentage', 'Question Text Word Frequency Distribution', 800, 500, 20, False)\n    return frequency\n    \nfrequency = count_word_frequency(train['question_text'])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"markdown","source":"* \"the\" word corresponds to 9.5% of the words in question_text column\n\n> Let´s check the next 20 words to check if we have some common topic."},{"metadata":{"trusted":true},"cell_type":"code","source":"frequency = count_word_frequency(train['question_text'], 20, 40)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"So we have some interesting words like many, world, played, name, song, movie and a lot more.\n\nLet´s check the test set and see if we have the same behaviour."},{"metadata":{"trusted":true},"cell_type":"code","source":"frequency = count_word_frequency(test['question_text'])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* Top words are repeated"},{"metadata":{"trusted":true},"cell_type":"code","source":"frequency = count_word_frequency(test['question_text'], 20, 40)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* We can see some differences between train and test."},{"metadata":{},"cell_type":"markdown","source":"# Document Text Exploration\n\nWe are only going to analyze the test set because this documents are very big."},{"metadata":{"trusted":true},"cell_type":"code","source":"frequency = count_word_frequency(test['document_text'])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* We need to clean this column to have a better idea.\n* Leaving this part for the future notebooks because i believe it´s not easy to clean it."},{"metadata":{},"cell_type":"markdown","source":"# Document URL\n\nThis is the url of the document text. Maybee we can use this for something but i will not analyze this variable because i believe we will not get any insight from it."},{"metadata":{},"cell_type":"markdown","source":"# Preprocess and Model\n\n* In this section we are going to build a baseline model\n* First, we need to create a preprocess function to pass the data and get a training and testing set with the correct format.\n* The main idea of my preprocessing is making a train were each long answer candidate text is going to be a row, in other words we are going to use the indices of the annotation to extract from the document text the answer. Next we can use the extracted segment and the question as features and label a ground truth variable y with 1 and 0.\n* Short answer have 4 possible answers, we are going to transform it to a binary classification problem were YES and NO are going to be empty answers.\n* We have a nice sample dataset to try this preprocessing function, let's start coding."},{"metadata":{"trusted":true},"cell_type":"code","source":"def build_train_test_long(df, train = True):\n    final_long_answer_frame = pd.DataFrame()\n    if train == True:\n        # get long answer\n        long_answer = extract_target_variable(df, False)\n        \n        # iterate over each row to get the possible answers\n        for index, row in df.iterrows():\n            start_end_tokens = []\n            questions = []\n            responds = []\n            for i in row['long_answer_candidates']:\n                start_token = i['start_token']\n                end_token = i['end_token']\n                start_end_token = str(i['start_token']) + ':' + str(i['end_token'])\n                question = row['question_text']\n                respond = \" \".join(row['document_text'].split()[start_token : end_token])\n                start_end_tokens.append(start_end_token)\n                questions.append(question)\n                responds.append(respond)\n\n            long_answer_frame = pd.DataFrame({'question': questions, 'respond': responds, 'start_end_token': start_end_tokens})\n            long_answer_frame['answer'] = long_answer.iloc[index][0]\n            long_answer_frame['target'] = long_answer_frame['start_end_token'] == long_answer_frame['answer']\n            long_answer_frame['target'] = long_answer_frame['target'].astype('int16')\n            long_answer_frame.drop(['answer'], inplace = True, axis = 1)\n            final_long_answer_frame = pd.concat([final_long_answer_frame, long_answer_frame])\n        return final_long_answer_frame\n    else:\n         # iterate over each row to get the possible answers\n        for index, row in df.iterrows():\n            start_end_tokens = []\n            questions = []\n            responds = []\n            for i in row['long_answer_candidates']:\n                start_token = i['start_token']\n                end_token = i['end_token']\n                start_end_token = str(i['start_token']) + ':' + str(i['end_token'])\n                question = row['question_text']\n                respond = \" \".join(row['document_text'].split()[start_token : end_token])\n                start_end_tokens.append(start_end_token)\n                questions.append(question)\n                responds.append(respond)\n\n            long_answer_frame = pd.DataFrame({'question': questions, 'respond': responds, 'start_end_token': start_end_tokens})\n            final_long_answer_frame = pd.concat([final_long_answer_frame, long_answer_frame])\n        return final_long_answer_frame\n        \n\n\ndef build_train_test_short(df, train = True):\n    \n    final_short_answer_frame = pd.DataFrame()\n    \n    if train == True:\n        # get short answer\n        short_answer = extract_target_variable(df, True)\n\n        # iterate over each row to get the possible answer\n        for index, row in df.iterrows():\n            start_tokens = []\n            end_tokens = []\n            start_end_tokens = []\n            questions = []\n            responds = []\n            for i in row['long_answer_candidates']:\n                start_token = i['start_token']\n                end_token = i['end_token']\n                start_end_token = str(i['start_token']) + ':' + str(i['end_token'])\n                question = row['question_text']\n                respond = \" \".join(row['document_text'].split()[int(start_token) : int(end_token)])\n                start_tokens.append(start_token)\n                end_tokens.append(end_token)\n                start_end_tokens.append(start_end_token)\n                questions.append(question)\n                responds.append(respond)\n\n            short_answer_frame = pd.DataFrame({'question': questions, 'respond': responds, 'start_token': start_tokens, 'end_token': end_tokens, 'start_end_token': start_end_tokens})\n            short_answer_frame['answer'] = short_answer.iloc[index][0]\n            short_answer_frame['start_token_an'] = short_answer_frame['answer'].apply(lambda x: x.split(':')[0] if ':' in x else 0)\n            short_answer_frame['end_token_an'] = short_answer_frame['answer'].apply(lambda x: x.split(':')[1] if ':' in x else 0)\n            short_answer_frame['start_token_an'] = short_answer_frame['start_token_an'].astype(int)\n            short_answer_frame['end_token_an'] = short_answer_frame['end_token_an'].astype(int)\n            short_answer_frame['target'] = 0\n            short_answer_frame.loc[(short_answer_frame['start_token_an'] >= short_answer_frame['start_token']) & (short_answer_frame['end_token_an'] <= short_answer_frame['end_token']), 'target'] = 1\n            short_answer_frame.drop(['answer', 'start_token', 'end_token', 'start_token_an', 'end_token_an'], inplace = True, axis = 1)\n            final_short_answer_frame = pd.concat([final_short_answer_frame, short_answer_frame])\n        return final_short_answer_frame\n    else:\n        # iterate over each row to get the possible answer\n        for index, row in df.iterrows():\n            start_end_tokens = []\n            questions = []\n            responds = []\n            for i in row['long_answer_candidates']:\n                start_token = i['start_token']\n                end_token = i['end_token']\n                start_end_token = str(i['start_token']) + ':' + str(i['end_token'])\n                question = row['question_text']\n                respond = \" \".join(row['document_text'].split()[int(start_token) : int(end_token)])\n                start_end_tokens.append(start_end_token)\n                questions.append(question)\n                responds.append(respond)\n\n            short_answer_frame = pd.DataFrame({'question': questions, 'respond': responds, 'start_end_token': start_end_tokens})\n            final_short_answer_frame = pd.concat([final_short_answer_frame, short_answer_frame])\n        return final_short_answer_frame","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sh = build_train_test_long(train.head())\nsh.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sh[sh['target']==1]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"long_answer.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* This is a sample of our training set for long answers.\n* We have 4 text answer for the first 5 question\n* For long answers we will get a probability for each question response combination. We will filter this result by each question to check if there is a hight probability for one of the answers. \n* For short answer we are going to do the same, but is harder because we only will know the long tokens. We need to figure how we can extract the short token indices.\n* The dataframe is very large so maybee we will need to remake this function so they work with a batching process.\n* We can make 2 models to resolve this problem (long answer and short answer)"},{"metadata":{},"cell_type":"markdown","source":"# Building the model in a different script to handle the memory better"}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":1}