{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"from https://www.kaggle.com/code/jhoward/iterate-like-a-grandmaster/notebook","metadata":{}},{"cell_type":"code","source":"from pathlib import Path\nimport os\n\niskaggle = os.environ.get('KAGGLE_KERNEL_RUN_TYPE', '')\n\nif iskaggle:\n    !pip install -Uqq fastai\n\nelse :\n    import zipfile, kaggle\n    path = Path('us-patent-phrase-to-phrase-matching')\n    kaggle.api.competition_download_cli(str(path))\n    zipfile.ZipFile(f'{path}.zip').extractall(path)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T04:40:02.153207Z","iopub.execute_input":"2022-07-31T04:40:02.153568Z","iopub.status.idle":"2022-07-31T04:40:17.437953Z","shell.execute_reply.started":"2022-07-31T04:40:02.153537Z","shell.execute_reply":"2022-07-31T04:40:17.436406Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n\nA lot of the basic imports you'll want (np, pd, plt, etc) are provided by fastai, so let's grab them in one line:\n","metadata":{}},{"cell_type":"code","source":"from fastai.imports import *","metadata":{"execution":{"iopub.status.busy":"2022-07-31T04:40:17.444076Z","iopub.execute_input":"2022-07-31T04:40:17.446300Z","iopub.status.idle":"2022-07-31T04:40:17.646360Z","shell.execute_reply.started":"2022-07-31T04:40:17.446255Z","shell.execute_reply":"2022-07-31T04:40:17.645222Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Import and EDA","metadata":{}},{"cell_type":"code","source":"if iskaggle: path = Path('../input/us-patent-phrase-to-phrase-matching')\npath.ls()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T04:40:17.652033Z","iopub.execute_input":"2022-07-31T04:40:17.654649Z","iopub.status.idle":"2022-07-31T04:40:17.669618Z","shell.execute_reply.started":"2022-07-31T04:40:17.654608Z","shell.execute_reply":"2022-07-31T04:40:17.668417Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Lets look at training set","metadata":{}},{"cell_type":"code","source":"df = pd.read_csv(path/'train.csv')\ndf","metadata":{"execution":{"iopub.status.busy":"2022-07-31T04:40:17.676486Z","iopub.execute_input":"2022-07-31T04:40:17.679112Z","iopub.status.idle":"2022-07-31T04:40:17.811679Z","shell.execute_reply.started":"2022-07-31T04:40:17.679075Z","shell.execute_reply":"2022-07-31T04:40:17.810703Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"this is the test set","metadata":{}},{"cell_type":"code","source":"eval_df = pd.read_csv(path/'test.csv')\nlen(eval_df)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T04:40:17.813105Z","iopub.execute_input":"2022-07-31T04:40:17.813455Z","iopub.status.idle":"2022-07-31T04:40:17.825842Z","shell.execute_reply.started":"2022-07-31T04:40:17.813419Z","shell.execute_reply":"2022-07-31T04:40:17.824669Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"eval_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T04:40:17.827131Z","iopub.execute_input":"2022-07-31T04:40:17.827597Z","iopub.status.idle":"2022-07-31T04:40:17.840541Z","shell.execute_reply.started":"2022-07-31T04:40:17.827559Z","shell.execute_reply":"2022-07-31T04:40:17.839480Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"looking at values of various fields","metadata":{}},{"cell_type":"code","source":"df.target.value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T04:40:17.842020Z","iopub.execute_input":"2022-07-31T04:40:17.842262Z","iopub.status.idle":"2022-07-31T04:40:17.873148Z","shell.execute_reply.started":"2022-07-31T04:40:17.842240Z","shell.execute_reply":"2022-07-31T04:40:17.872127Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n\nWe see that there's nearly as many unique targets as items in the training set, so they're nearly but not quite unique. Most importantly, we can see that these generally contain very few words (1-4 words in the above sample).\n\nLet's check anchor:\n","metadata":{}},{"cell_type":"code","source":"df.anchor.value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T04:40:17.874701Z","iopub.execute_input":"2022-07-31T04:40:17.875135Z","iopub.status.idle":"2022-07-31T04:40:17.889927Z","shell.execute_reply.started":"2022-07-31T04:40:17.875098Z","shell.execute_reply":"2022-07-31T04:40:17.888282Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n\nWe can see here that there's far fewer unique values (just 733) and that again they're very short (2-4 words in this sample).\n\nNow we'll do context\n","metadata":{}},{"cell_type":"code","source":"df.context.value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T04:40:17.892846Z","iopub.execute_input":"2022-07-31T04:40:17.893139Z","iopub.status.idle":"2022-07-31T04:40:17.906142Z","shell.execute_reply.started":"2022-07-31T04:40:17.893116Z","shell.execute_reply":"2022-07-31T04:40:17.905063Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"These are just short codes. Some of them have very few examples (18 in the smallest case) The first character is the section the patent was filed under -- let's create a column for that and look at the distribution","metadata":{}},{"cell_type":"code","source":"df['section'] = df.context.str[0]\ndf.section.value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T04:40:17.910868Z","iopub.execute_input":"2022-07-31T04:40:17.911391Z","iopub.status.idle":"2022-07-31T04:40:17.951637Z","shell.execute_reply.started":"2022-07-31T04:40:17.911366Z","shell.execute_reply":"2022-07-31T04:40:17.950740Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n\nIt seems likely that these sections might be useful, since they've got quite a bit more data in each.\n\nFinally, we'll take a look at a histogram of the scores:\n","metadata":{}},{"cell_type":"code","source":"df.score.hist()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T04:40:17.953227Z","iopub.execute_input":"2022-07-31T04:40:17.953873Z","iopub.status.idle":"2022-07-31T04:40:18.170846Z","shell.execute_reply.started":"2022-07-31T04:40:17.953838Z","shell.execute_reply":"2022-07-31T04:40:18.169576Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df[df.score==1]","metadata":{"execution":{"iopub.status.busy":"2022-07-31T04:40:18.172081Z","iopub.execute_input":"2022-07-31T04:40:18.172700Z","iopub.status.idle":"2022-07-31T04:40:18.198104Z","shell.execute_reply.started":"2022-07-31T04:40:18.172664Z","shell.execute_reply":"2022-07-31T04:40:18.197202Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n\nWe can see from this that these are just minor rewordings of the same concept, and isn't likely to be specific to context. Any pretrained model should be pretty good at finding these already.\n","metadata":{}},{"cell_type":"markdown","source":"# training","metadata":{}},{"cell_type":"code","source":"from torch.utils.data import DataLoader\nimport warnings, transformers, logging, torch\nfrom transformers import TrainingArguments, Trainer\nfrom transformers import AutoModelForSequenceClassification, AutoTokenizer","metadata":{"execution":{"iopub.status.busy":"2022-07-31T04:40:18.199452Z","iopub.execute_input":"2022-07-31T04:40:18.200012Z","iopub.status.idle":"2022-07-31T04:40:26.399629Z","shell.execute_reply.started":"2022-07-31T04:40:18.199973Z","shell.execute_reply":"2022-07-31T04:40:26.398587Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if iskaggle:\n    !pip install -q datasets\nimport datasets\nfrom datasets import load_dataset, Dataset, DatasetDict","metadata":{"execution":{"iopub.status.busy":"2022-07-31T04:40:26.402434Z","iopub.execute_input":"2022-07-31T04:40:26.403835Z","iopub.status.idle":"2022-07-31T04:40:36.425403Z","shell.execute_reply.started":"2022-07-31T04:40:26.403789Z","shell.execute_reply":"2022-07-31T04:40:36.424286Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n\nHuggingFace Transformers tends to be rather enthusiastic about spitting out lots of warnings, so let's quieten it down for our sanity:\n","metadata":{}},{"cell_type":"code","source":"warnings.simplefilter('ignore')\nlogging.disable(logging.WARNING)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T04:40:36.427160Z","iopub.execute_input":"2022-07-31T04:40:36.428098Z","iopub.status.idle":"2022-07-31T04:40:36.434975Z","shell.execute_reply.started":"2022-07-31T04:40:36.428065Z","shell.execute_reply":"2022-07-31T04:40:36.433913Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n\nI tried to find a model that I could train reasonably at home in under two minutes, but got reasonable accuracy from. I found that deberta-v3-small fits the bill, so let's use it:\n","metadata":{}},{"cell_type":"code","source":"model_nm = 'microsoft/deberta-v3-small'","metadata":{"execution":{"iopub.status.busy":"2022-07-31T04:40:36.436445Z","iopub.execute_input":"2022-07-31T04:40:36.437430Z","iopub.status.idle":"2022-07-31T04:40:36.445354Z","shell.execute_reply.started":"2022-07-31T04:40:36.437400Z","shell.execute_reply":"2022-07-31T04:40:36.444413Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can now create a tokenizer for this model. Note that pretrained models assume that text is tokenized in a particular way. In order to ensure that your tokenizer matches your model, use the AutoTokenizer, passing in your model name.","metadata":{}},{"cell_type":"code","source":"tokz = AutoTokenizer.from_pretrained(model_nm)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T04:40:36.448343Z","iopub.execute_input":"2022-07-31T04:40:36.450170Z","iopub.status.idle":"2022-07-31T04:40:42.128385Z","shell.execute_reply.started":"2022-07-31T04:40:36.450036Z","shell.execute_reply":"2022-07-31T04:40:42.127531Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We'll need to combine the context, anchor, and target together somehow. There's not much research as to the best way to do this, so we may need to iterate a bit. To start with, we'll just combine them all into a single string. The model will need to know where each section starts, so we can use the special separator token to tell it:","metadata":{}},{"cell_type":"code","source":"sep = tokz.sep_token\nsep","metadata":{"execution":{"iopub.status.busy":"2022-07-31T04:40:42.129842Z","iopub.execute_input":"2022-07-31T04:40:42.130186Z","iopub.status.idle":"2022-07-31T04:40:42.138278Z","shell.execute_reply.started":"2022-07-31T04:40:42.130148Z","shell.execute_reply":"2022-07-31T04:40:42.137428Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df['inputs'] = df.context + sep + df.anchor + sep + df.target","metadata":{"execution":{"iopub.status.busy":"2022-07-31T04:40:42.139773Z","iopub.execute_input":"2022-07-31T04:40:42.140644Z","iopub.status.idle":"2022-07-31T04:40:42.164608Z","shell.execute_reply.started":"2022-07-31T04:40:42.140605Z","shell.execute_reply":"2022-07-31T04:40:42.163768Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"looking at df","metadata":{}},{"cell_type":"code","source":"df.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T05:06:35.126632Z","iopub.execute_input":"2022-07-31T05:06:35.127433Z","iopub.status.idle":"2022-07-31T05:06:35.150234Z","shell.execute_reply.started":"2022-07-31T05:06:35.127384Z","shell.execute_reply":"2022-07-31T05:06:35.149297Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Generally we'll get best performance if we convert pandas DataFrames into HuggingFace Datasets, so we'll convert them over, and also rename the score column to what Transformers expects for the dependent variable, which is label:","metadata":{}},{"cell_type":"code","source":"ds = Dataset.from_pandas(df).rename_column('score', 'label')\neval_ds = Dataset.from_pandas(eval_df)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T04:40:42.165881Z","iopub.execute_input":"2022-07-31T04:40:42.166199Z","iopub.status.idle":"2022-07-31T04:40:42.205054Z","shell.execute_reply.started":"2022-07-31T04:40:42.166165Z","shell.execute_reply":"2022-07-31T04:40:42.204234Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ds[0]","metadata":{"execution":{"iopub.status.busy":"2022-07-31T05:12:47.161417Z","iopub.execute_input":"2022-07-31T05:12:47.162289Z","iopub.status.idle":"2022-07-31T05:12:47.169360Z","shell.execute_reply.started":"2022-07-31T05:12:47.162236Z","shell.execute_reply":"2022-07-31T05:12:47.168419Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n\nTo tokenize the data, we'll create a function (since that's what Dataset.map will need):\n","metadata":{}},{"cell_type":"code","source":"def tok_func(x): \n    return tokz(x[\"inputs\"])","metadata":{"execution":{"iopub.status.busy":"2022-07-31T04:40:42.206217Z","iopub.execute_input":"2022-07-31T04:40:42.206445Z","iopub.status.idle":"2022-07-31T04:40:42.212575Z","shell.execute_reply.started":"2022-07-31T04:40:42.206413Z","shell.execute_reply":"2022-07-31T04:40:42.211472Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's try tokenizing one input and see how it looks","metadata":{}},{"cell_type":"code","source":"tok_func(ds[0])","metadata":{"execution":{"iopub.status.busy":"2022-07-31T04:40:42.214154Z","iopub.execute_input":"2022-07-31T04:40:42.214414Z","iopub.status.idle":"2022-07-31T04:40:42.230460Z","shell.execute_reply.started":"2022-07-31T04:40:42.214380Z","shell.execute_reply":"2022-07-31T04:40:42.229648Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n\nThe only bit we care about at the moment is input_ids. We can see in the tokens that it starts with a special token 1 (which represents the start of text), and then has our three fields separated by the separator token 2. We can check the indices of the special token IDs like so:\n","metadata":{}},{"cell_type":"code","source":"tokz.all_special_tokens","metadata":{"execution":{"iopub.status.busy":"2022-07-31T04:40:42.232023Z","iopub.execute_input":"2022-07-31T04:40:42.232819Z","iopub.status.idle":"2022-07-31T04:40:42.239703Z","shell.execute_reply.started":"2022-07-31T04:40:42.232780Z","shell.execute_reply":"2022-07-31T04:40:42.238716Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can now tokenize the input. We'll use batching to speed it up, and remove the columns we no longer need:","metadata":{}},{"cell_type":"code","source":"inps = \"anchor\", \"target\", \"context\"\ntok_ds = ds.map(tok_func, batched=True, remove_columns=inps+('inputs', 'id', 'section'))","metadata":{"execution":{"iopub.status.busy":"2022-07-31T04:40:42.241233Z","iopub.execute_input":"2022-07-31T04:40:42.241700Z","iopub.status.idle":"2022-07-31T04:40:44.114907Z","shell.execute_reply.started":"2022-07-31T04:40:42.241665Z","shell.execute_reply":"2022-07-31T04:40:44.113981Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tok_ds[0]","metadata":{"execution":{"iopub.status.busy":"2022-07-31T04:40:44.116320Z","iopub.execute_input":"2022-07-31T04:40:44.117204Z","iopub.status.idle":"2022-07-31T04:40:44.127613Z","shell.execute_reply.started":"2022-07-31T04:40:44.117163Z","shell.execute_reply":"2022-07-31T04:40:44.126667Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# creating a validation set","metadata":{}},{"cell_type":"markdown","source":"\n\nAccording to this post, the private test anchors do not overlap with the training set. So let's do the same thing for our validation set.\n\nFirst, create a randomly shuffled list of anchors:\n","metadata":{}},{"cell_type":"code","source":"anchors = df.anchor.unique()\nnp.random.seed(42)\nnp.random.shuffle(anchors)\nanchors[:5]","metadata":{"execution":{"iopub.status.busy":"2022-07-31T04:40:44.129026Z","iopub.execute_input":"2022-07-31T04:40:44.129673Z","iopub.status.idle":"2022-07-31T04:40:44.155827Z","shell.execute_reply.started":"2022-07-31T04:40:44.129634Z","shell.execute_reply":"2022-07-31T04:40:44.154824Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n\nNow we can pick some proportion (e.g 25%) of these anchors to go in the validation set:\n","metadata":{}},{"cell_type":"code","source":"val_prop = 0.25\nval_sz = int(len(anchors)*val_prop)\nval_anchors = anchors[:val_sz]","metadata":{"execution":{"iopub.status.busy":"2022-07-31T04:40:44.157473Z","iopub.execute_input":"2022-07-31T04:40:44.157936Z","iopub.status.idle":"2022-07-31T04:40:44.163302Z","shell.execute_reply.started":"2022-07-31T04:40:44.157898Z","shell.execute_reply":"2022-07-31T04:40:44.162246Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n\nNow we can get a list of which rows match val_anchors, and get their indices:\n","metadata":{}},{"cell_type":"code","source":"is_val = np.isin(df.anchor, val_anchors)\nidxs = np.arange(len(df))\nval_idxs = idxs[is_val]\ntrn_idxs = idxs[~is_val]\nlen(val_idxs), len(trn_idxs)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T04:40:44.169885Z","iopub.execute_input":"2022-07-31T04:40:44.171735Z","iopub.status.idle":"2022-07-31T04:40:44.392613Z","shell.execute_reply.started":"2022-07-31T04:40:44.171702Z","shell.execute_reply":"2022-07-31T04:40:44.391457Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Our training and validation Datasets can now be selected, and put into a DatasetDict ready for training:","metadata":{}},{"cell_type":"code","source":"dds = DatasetDict({\"train\": tok_ds.select(trn_idxs), \n                   \"test\": tok_ds.select(val_idxs)})","metadata":{"execution":{"iopub.status.busy":"2022-07-31T04:40:44.394079Z","iopub.execute_input":"2022-07-31T04:40:44.394602Z","iopub.status.idle":"2022-07-31T04:40:44.416315Z","shell.execute_reply.started":"2022-07-31T04:40:44.394533Z","shell.execute_reply":"2022-07-31T04:40:44.415513Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n\nBTW, a lot of people do more complex stuff for creating their validation set, but with a dataset this large there's not much point. As you can see, the mean scores in the two groups are very similar despite just doing a random shuffle:\n","metadata":{}},{"cell_type":"code","source":"df.iloc[trn_idxs].score.mean(), df.iloc[val_idxs].score.mean()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T04:40:44.418613Z","iopub.execute_input":"2022-07-31T04:40:44.419257Z","iopub.status.idle":"2022-07-31T04:40:44.438124Z","shell.execute_reply.started":"2022-07-31T04:40:44.419203Z","shell.execute_reply":"2022-07-31T04:40:44.436970Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# initial model","metadata":{}},{"cell_type":"markdown","source":"\n\nLet's now train our model! We'll need to specify a metric, which is the correlation coefficient provided by numpy (we need to return a dictionary since that's how Transformers knows what label to use):\n","metadata":{}},{"cell_type":"code","source":"def corr(eval_pred): return {'pearson': np.corrcoef(*eval_pred)[0][1]}","metadata":{"execution":{"iopub.status.busy":"2022-07-31T04:40:44.439828Z","iopub.execute_input":"2022-07-31T04:40:44.440160Z","iopub.status.idle":"2022-07-31T04:40:44.447568Z","shell.execute_reply.started":"2022-07-31T04:40:44.440123Z","shell.execute_reply":"2022-07-31T04:40:44.446624Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n\nWe pick a learning rate and batch size that fits our GPU, and pick a reasonable weight decay and small number of epochs:\n","metadata":{}},{"cell_type":"code","source":"lr, bs = 8e-5, 128\nwd, epochs = 0.01, 4","metadata":{"execution":{"iopub.status.busy":"2022-07-31T04:40:44.449271Z","iopub.execute_input":"2022-07-31T04:40:44.449560Z","iopub.status.idle":"2022-07-31T04:40:44.455811Z","shell.execute_reply.started":"2022-07-31T04:40:44.449525Z","shell.execute_reply":"2022-07-31T04:40:44.454637Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n\nThree epochs might not sound like much, but you'll see once we train that most of the progress can be made in that time, so this is good for experimentation.\n\nTransformers uses the TrainingArguments class to set up arguments. We'll use a cosine scheduler with warmup, since at fast.ai we've found that's pretty reliable. We'll use fp16 since it's much faster on modern GPUs, and saves some memory. We evaluate using double-sized batches, since no gradients are stored so we can do twice as many rows at a time.\n","metadata":{}},{"cell_type":"code","source":"def get_trainer(dds):\n    args = TrainingArguments('outputs', learning_rate=lr, warmup_ratio=0.1, lr_scheduler_type='cosine', fp16=True,\n        evaluation_strategy=\"epoch\", per_device_train_batch_size=bs, per_device_eval_batch_size=bs*2,\n        num_train_epochs=epochs, weight_decay=wd, report_to='none')\n    model = AutoModelForSequenceClassification.from_pretrained(model_nm, num_labels=1)\n    return Trainer(model, args, train_dataset=dds['train'], eval_dataset=dds['test'],\n                   tokenizer=tokz, compute_metrics=corr)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T04:40:44.457437Z","iopub.execute_input":"2022-07-31T04:40:44.457776Z","iopub.status.idle":"2022-07-31T04:40:44.466467Z","shell.execute_reply.started":"2022-07-31T04:40:44.457674Z","shell.execute_reply":"2022-07-31T04:40:44.465609Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"args = TrainingArguments('outputs', learning_rate=lr, warmup_ratio=0.1, lr_scheduler_type='cosine', fp16=True,evaluation_strategy=\"epoch\", per_device_train_batch_size=bs, per_device_eval_batch_size=bs*2,num_train_epochs=epochs, weight_decay=wd, report_to='none')","metadata":{"execution":{"iopub.status.busy":"2022-07-31T04:40:44.468078Z","iopub.execute_input":"2022-07-31T04:40:44.468344Z","iopub.status.idle":"2022-07-31T04:40:44.546711Z","shell.execute_reply.started":"2022-07-31T04:40:44.468307Z","shell.execute_reply":"2022-07-31T04:40:44.545842Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model = AutoModelForSequenceClassification.from_pretrained(model_nm, num_labels=1)\ntrainer = Trainer(model, args, train_dataset=dds['train'], eval_dataset=dds['test'], tokenizer=tokz, compute_metrics=corr)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T04:40:44.548241Z","iopub.execute_input":"2022-07-31T04:40:44.548481Z","iopub.status.idle":"2022-07-31T04:40:57.877597Z","shell.execute_reply.started":"2022-07-31T04:40:44.548446Z","shell.execute_reply":"2022-07-31T04:40:57.876583Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"trainer.train()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T04:40:57.880556Z","iopub.execute_input":"2022-07-31T04:40:57.880832Z","iopub.status.idle":"2022-07-31T04:45:05.621764Z","shell.execute_reply.started":"2022-07-31T04:40:57.880796Z","shell.execute_reply":"2022-07-31T04:45:05.620804Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# improving the model","metadata":{}},{"cell_type":"markdown","source":"\n\nWe now want to start iterating to improve this. To do that, we need to know whether the model gives stable results. I tried training it 3 times from scratch, and got a range of outcomes from 0.808-0.810. This is stable enough to make a start - if we're not finding improvements that are visible within this range, then they're not very significant! Later on, if and when we feel confident that we've got the basics right, we can use cross validation and more epochs of training.\n\nIteration speed is critical, so we need to quickly be able to try different data processing and trainer parameters. So let's create a function to quickly apply tokenization and create our DatasetDict:\n","metadata":{}},{"cell_type":"code","source":"\n\ndef get_dds(df):\n    ds = Dataset.from_pandas(df).rename_column('score', 'label')\n    tok_ds = ds.map(tok_func, batched=True, remove_columns=inps+('inputs','id','section'))\n    return DatasetDict({\"train\":tok_ds.select(trn_idxs), \"test\": tok_ds.select(val_idxs)})\n\n","metadata":{"execution":{"iopub.status.busy":"2022-07-31T04:45:05.624596Z","iopub.execute_input":"2022-07-31T04:45:05.624799Z","iopub.status.idle":"2022-07-31T04:45:05.630699Z","shell.execute_reply.started":"2022-07-31T04:45:05.624774Z","shell.execute_reply":"2022-07-31T04:45:05.629694Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n\n...and also a function to create a Trainer:\n","metadata":{}},{"cell_type":"code","source":"def get_model(): return AutoModelForSequenceClassification.from_pretrained(model_nm, num_labels=1)\n\ndef get_trainer(dds, model=None):\n    if model is None: model = get_model()\n    args = TrainingArguments('outputs', learning_rate=lr, warmup_ratio=0.1, lr_scheduler_type='cosine', fp16=True,\n        evaluation_strategy=\"epoch\", per_device_train_batch_size=bs, per_device_eval_batch_size=bs*2,\n        num_train_epochs=epochs, weight_decay=wd, report_to='none')\n    return Trainer(model, args, train_dataset=dds['train'], eval_dataset=dds['test'],\n                   tokenizer=tokz, compute_metrics=corr)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T04:45:05.632246Z","iopub.execute_input":"2022-07-31T04:45:05.632665Z","iopub.status.idle":"2022-07-31T04:45:05.647152Z","shell.execute_reply.started":"2022-07-31T04:45:05.632628Z","shell.execute_reply":"2022-07-31T04:45:05.646266Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n\nLet's now try out some ideas...\n\nPerhaps using the special separator character isn't a good idea, and we should use something we create instead. Let's see if that makes things better. First we'll change the separator and create the DatasetDict:\n","metadata":{}},{"cell_type":"code","source":"sep = \" [s] \"\ndf['inputs'] = df.context + sep + df.anchor + sep + df.target\ndds = get_dds(df)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T04:45:05.650681Z","iopub.execute_input":"2022-07-31T04:45:05.650912Z","iopub.status.idle":"2022-07-31T04:45:07.583078Z","shell.execute_reply.started":"2022-07-31T04:45:05.650888Z","shell.execute_reply":"2022-07-31T04:45:07.582303Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n\n...and create and train a model.\n","metadata":{}},{"cell_type":"code","source":"get_trainer(dds).train()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T04:45:07.584381Z","iopub.execute_input":"2022-07-31T04:45:07.584900Z","iopub.status.idle":"2022-07-31T04:49:29.873850Z","shell.execute_reply.started":"2022-07-31T04:45:07.584859Z","shell.execute_reply":"2022-07-31T04:49:29.872875Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n\nThat's looking quite a bit better, so we'll keep that change.\n\nOften changing to lowercase is helpful. Let's see if that helps too:\n","metadata":{}},{"cell_type":"code","source":"df['inputs'] = df.inputs.str.lower()\ndds = get_dds(df)\nget_trainer(dds).train()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T04:49:29.875490Z","iopub.execute_input":"2022-07-31T04:49:29.875764Z","iopub.status.idle":"2022-07-31T04:53:54.569398Z","shell.execute_reply.started":"2022-07-31T04:49:29.875728Z","shell.execute_reply":"2022-07-31T04:53:54.568562Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n\nThat one is less clear. We'll keep that change too since most times I run it, it's a little better.\n","metadata":{}},{"cell_type":"markdown","source":"# creating our own special tokens","metadata":{}},{"cell_type":"code","source":"df['sectok'] = '[' + df.section + ']'\nsectoks = list(df.sectok.unique())\ntokz.add_special_tokens({'additional_special_tokens': sectoks})","metadata":{"execution":{"iopub.status.busy":"2022-07-31T04:53:54.573879Z","iopub.execute_input":"2022-07-31T04:53:54.574411Z","iopub.status.idle":"2022-07-31T04:53:54.594943Z","shell.execute_reply.started":"2022-07-31T04:53:54.574372Z","shell.execute_reply":"2022-07-31T04:53:54.594157Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"concatenating the section token to the start of our inputs","metadata":{}},{"cell_type":"code","source":"df['inputs'] = df.sectok + sep + df.context + sep + df.anchor.str.lower() + sep + df.target\ndds = get_dds(df)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T04:53:54.596430Z","iopub.execute_input":"2022-07-31T04:53:54.596889Z","iopub.status.idle":"2022-07-31T04:53:57.143226Z","shell.execute_reply.started":"2022-07-31T04:53:54.596853Z","shell.execute_reply":"2022-07-31T04:53:57.142414Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n\nSince we've added more tokens, we need to resize the embedding matrix in the model:\n","metadata":{}},{"cell_type":"code","source":"\n\nmodel = get_model()\nmodel.resize_token_embeddings(len(tokz))\n\n","metadata":{"execution":{"iopub.status.busy":"2022-07-31T04:53:57.144869Z","iopub.execute_input":"2022-07-31T04:53:57.145123Z","iopub.status.idle":"2022-07-31T04:54:01.269915Z","shell.execute_reply.started":"2022-07-31T04:53:57.145088Z","shell.execute_reply":"2022-07-31T04:54:01.268891Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we are ready to train","metadata":{}},{"cell_type":"code","source":"trainer = get_trainer(dds, model=model)\ntrainer.train()","metadata":{"execution":{"iopub.status.busy":"2022-07-31T04:54:01.271510Z","iopub.execute_input":"2022-07-31T04:54:01.271770Z","iopub.status.idle":"2022-07-31T04:58:39.959200Z","shell.execute_reply.started":"2022-07-31T04:54:01.271735Z","shell.execute_reply":"2022-07-31T04:58:39.958352Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Jeremy presents with many more ideas","metadata":{}},{"cell_type":"markdown","source":"\n\nIt looks like we've made another bit of an improvement!\n\nThere's plenty more things you could try. Here's some thoughts:\n\n    Try a model pretrained on legal vocabulary. E.g. how about BERT for patents?\n    You'd likely get better results by using a sentence similarity model. Did you know that there's a patent similarity model you could try?\n    You could also fine-tune any HuggingFace model using the full patent database (which is provided in BigQuery), before applying it to this dataset\n    Replace the patent context field with the description of that context provided by the patent office\n    ...and try out your own ideas too!\n\nBefore submitting a model, retrain it on the full dataset, rather than just the 75% training subset we've used here. Create a function like the ones above to make that easy for you!\"\n","metadata":{}},{"cell_type":"markdown","source":"# cross validation","metadata":{}},{"cell_type":"code","source":"n_folds =4","metadata":{"execution":{"iopub.status.busy":"2022-07-31T04:58:39.960648Z","iopub.execute_input":"2022-07-31T04:58:39.961123Z","iopub.status.idle":"2022-07-31T04:58:39.965962Z","shell.execute_reply.started":"2022-07-31T04:58:39.961085Z","shell.execute_reply":"2022-07-31T04:58:39.964848Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n\nOnce you've gotten the low hanging fruit, you might want to use cross-validation to see the impact of minor changes. This time we'll use StratifiedGroupKFold, partly just to show a different approach to before, and partly because it will give us slightly better balanced datasets.\n","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import StratifiedGroupKFold\ncv = StratifiedGroupKFold(n_splits=n_folds)","metadata":{"execution":{"iopub.status.busy":"2022-07-31T04:58:39.967339Z","iopub.execute_input":"2022-07-31T04:58:39.967861Z","iopub.status.idle":"2022-07-31T04:58:39.988378Z","shell.execute_reply.started":"2022-07-31T04:58:39.967826Z","shell.execute_reply":"2022-07-31T04:58:39.987567Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n\nHere's how to split the data frame into n_folds groups, with non-overlapping anchors and matched scores, after randomly shuffling the rows:\n","metadata":{}},{"cell_type":"code","source":"\n\ndf = df.sample(frac=1, random_state=42)\nscores = (df.score*100).astype(int)\nfolds = list(cv.split(idxs, scores, df.anchor))\nfolds\n\n","metadata":{"execution":{"iopub.status.busy":"2022-07-31T04:58:39.990602Z","iopub.execute_input":"2022-07-31T04:58:39.991074Z","iopub.status.idle":"2022-07-31T04:58:40.357181Z","shell.execute_reply.started":"2022-07-31T04:58:39.991038Z","shell.execute_reply":"2022-07-31T04:58:40.356236Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n\nWe can now create a little function to split into training and validation sets based on a fold:\n","metadata":{}},{"cell_type":"code","source":"\n\ndef get_fold(folds, fold_num):\n    trn,val = folds[fold_num]\n    return DatasetDict({\"train\":tok_ds.select(trn), \"test\": tok_ds.select(val)})\n\n","metadata":{"execution":{"iopub.status.busy":"2022-07-31T04:58:40.358673Z","iopub.execute_input":"2022-07-31T04:58:40.359152Z","iopub.status.idle":"2022-07-31T04:58:40.365533Z","shell.execute_reply.started":"2022-07-31T04:58:40.359112Z","shell.execute_reply":"2022-07-31T04:58:40.364423Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Lets try it out","metadata":{}},{"cell_type":"code","source":"dds = get_fold(folds, 0)\ndds","metadata":{"execution":{"iopub.status.busy":"2022-07-31T04:58:40.367315Z","iopub.execute_input":"2022-07-31T04:58:40.367586Z","iopub.status.idle":"2022-07-31T04:58:40.394573Z","shell.execute_reply.started":"2022-07-31T04:58:40.367552Z","shell.execute_reply":"2022-07-31T04:58:40.393763Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n\nWe can now pass this into get_trainer as we did before. If we have, say, 4 folds, then doing that for each fold will give us 4 models, and 4 sets of predictions and metrics. You could ensemble the 4 models to get a stronger model, and can also average the 4 metrics to get a more accurate assessment of your model. Here's how to get the final epoch metrics from a trainer:\n","metadata":{}},{"cell_type":"code","source":"metrics = [o['eval_pearson'] for o in trainer.state.log_history if 'eval_pearson' in o]\nmetrics[-1]","metadata":{"execution":{"iopub.status.busy":"2022-07-31T04:58:40.397305Z","iopub.execute_input":"2022-07-31T04:58:40.397561Z","iopub.status.idle":"2022-07-31T04:58:40.406663Z","shell.execute_reply.started":"2022-07-31T04:58:40.397535Z","shell.execute_reply":"2022-07-31T04:58:40.405535Z"},"trusted":true},"execution_count":null,"outputs":[]}]}