{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"!pip install svgling","metadata":{"execution":{"iopub.status.busy":"2022-07-07T11:22:00.505301Z","iopub.execute_input":"2022-07-07T11:22:00.506003Z","iopub.status.idle":"2022-07-07T11:22:14.209039Z","shell.execute_reply.started":"2022-07-07T11:22:00.505883Z","shell.execute_reply":"2022-07-07T11:22:14.207687Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"heading\">\n   <h1><span style=\"color: Black\">Keyword Extraction using spaCy Named Entity Recognition</span></h1>\n</div>","metadata":{}},{"cell_type":"markdown","source":"\n<div class=\"heading\">\n   <h1><span style=\"color: Black\">Basic Imports</span></h1>\n</div>","metadata":{}},{"cell_type":"code","source":"import re \nimport os\nimport numpy as np\nimport pandas as pd\nfrom nltk import word_tokenize, pos_tag, RegexpParser\n\nfrom nltk.draw.tree import TreeView\nfrom IPython.display import Image\nimport svgling\n\nfrom nltk.chunk import conlltags2tree, tree2conlltags,ne_chunk\nfrom pprint import pprint\n\nimport sqlite3\nimport spacy\nfrom spacy import displacy\n\nimport torch\nfrom transformers import AutoModelForTokenClassification,AutoTokenizer, get_scheduler, pipeline\nfrom torch.optim import AdamW\nfrom tqdm.notebook import tqdm\nfrom sklearn.metrics import f1_score, accuracy_score\n\nfrom sklearn.model_selection import train_test_split","metadata":{"execution":{"iopub.status.busy":"2022-07-07T11:47:07.542744Z","iopub.execute_input":"2022-07-07T11:47:07.543111Z","iopub.status.idle":"2022-07-07T11:47:07.550069Z","shell.execute_reply.started":"2022-07-07T11:47:07.543081Z","shell.execute_reply":"2022-07-07T11:47:07.549078Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"nlp_start_df = pd.read_csv('../input/nlp-getting-started/train.csv')\n\nex = nlp_start_df.loc[100]['text']\n\nsent = pos_tag(word_tokenize(ex))\nprint(sent)","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-07-07T11:22:25.475723Z","iopub.execute_input":"2022-07-07T11:22:25.476357Z","iopub.status.idle":"2022-07-07T11:22:25.656299Z","shell.execute_reply.started":"2022-07-07T11:22:25.476318Z","shell.execute_reply":"2022-07-07T11:22:25.655283Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"content\">\nFrom the example above, we have that the word <b>'Experts' is 'NNS' (noun plural), word 'France' is 'NNP' (proper noun) and 'French' is 'JJ' (adjective)</b>. The whole list of abbreviations can be found <a href=\"https://www.guru99.com/pos-tagging-chunking-nltk.html\">here</a>. After this step, we can start with noun phrase chunking to named entities\n</div>\n<a id =topic4> </a>\n<div class=\"heading\">\n   <h1><span style=\"color: Black\">Chunking</span></h1>\n</div>\n<div class=\"content\">\n\n<b>Chunking in NLP is a process of grouping small pieces of information into large units.</b> The primary use of Chunking is making groups of \"noun phrases.\" It is used to add structure to the sentence by following POS tagging combined with regular expressions. The resulted group of words are called \"chunks.\"\n\nThere are no pre-defined rules for Chunking, but we can made according to our needs. Thus, if we want to chunk <b>only 'NN'</b> tags, we need to use pattern <pre><code>`mychunk:{&lt;NN>}`</code></pre> but if we need to chunk all types of tags which <b>start with 'NN'</b>, we'll use <pre><code>`mychunk:{&lt;NN.*>}`.</code></pre> More about regex patterns can be found <a href=\"https://www.w3schools.com/python/python_regex.asp\">here</a>\n</div>","metadata":{}},{"cell_type":"code","source":"# chunk all adjacence nouns\npatterns= \"\"\"mychunk:{<NN.*>+}\"\"\"\nchunker = RegexpParser(patterns)\noutput = chunker.parse(sent)\nprint(\"After Chunking\",output)\nsvgling.draw_tree(output)","metadata":{"execution":{"iopub.status.busy":"2022-07-07T11:22:25.658900Z","iopub.execute_input":"2022-07-07T11:22:25.659571Z","iopub.status.idle":"2022-07-07T11:22:25.706848Z","shell.execute_reply.started":"2022-07-07T11:22:25.659516Z","shell.execute_reply":"2022-07-07T11:22:25.705852Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"iob_tagged = tree2conlltags(output)\nprint(iob_tagged)","metadata":{"execution":{"iopub.status.busy":"2022-07-07T11:22:25.708285Z","iopub.execute_input":"2022-07-07T11:22:25.710326Z","iopub.status.idle":"2022-07-07T11:22:25.715841Z","shell.execute_reply.started":"2022-07-07T11:22:25.710290Z","shell.execute_reply":"2022-07-07T11:22:25.714806Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id =topic6> </a>\n<div class=\"heading\">\n   <h1><span style=\"color: Black\">Extracting Named Entities</span></h1>\n</div>\n<div class=\"content\">\nRecognizing a <b>named entity</b> is a specific kind of chunk extraction that uses entity tags along with chunk tags. Common entity tags include <b>PERSON, LOCATION, and ORGANIZATION</b>. NLTK has already a pre-trained named entity chunker which can be used using ne_chunk() method in the nltk.chunk module.\n</div>","metadata":{}},{"cell_type":"code","source":"def extract_ne(trees, labels):\n    \n    ne_list = []\n    for tree in ne_res:\n        if hasattr(tree, 'label'):\n            if tree.label() in labels:\n                ne_list.append(tree)\n    return ne_list\n            \nne_res = ne_chunk(pos_tag(word_tokenize(ex)))\nlabels = ['ORGANIZATION']\n\nprint(extract_ne(ne_res, labels))","metadata":{"execution":{"iopub.status.busy":"2022-07-07T11:22:25.717536Z","iopub.execute_input":"2022-07-07T11:22:25.718201Z","iopub.status.idle":"2022-07-07T11:22:25.859084Z","shell.execute_reply.started":"2022-07-07T11:22:25.718165Z","shell.execute_reply":"2022-07-07T11:22:25.857888Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id =topic7> </a>\n<div class=\"heading\">\n   <h1><span style=\"color: Black\">spaCy Named Entity Recognition</span></h1>\n</div>\n<div class=\"content\">\n<b>spaCy features an extremely fast statistical entity recognition system, that assigns labels to contiguous spans of tokens</b>. The default trained pipelines can identify a variety of named and numeric entities, including companies, locations, organizations and products.\n</div>","metadata":{}},{"cell_type":"code","source":"connection = sqlite3.connect('../input/wikibooks-dataset/wikibooks.sqlite')\ndf_wiki_books = pd.read_sql_query(\"SELECT * FROM en\", connection)\ndf_wiki_books.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-07T11:22:25.861572Z","iopub.execute_input":"2022-07-07T11:22:25.862207Z","iopub.status.idle":"2022-07-07T11:23:12.841564Z","shell.execute_reply.started":"2022-07-07T11:22:25.862167Z","shell.execute_reply":"2022-07-07T11:23:12.840573Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"nlp = spacy.load(\"en_core_web_sm\")","metadata":{"execution":{"iopub.status.busy":"2022-07-07T11:23:12.842979Z","iopub.execute_input":"2022-07-07T11:23:12.844794Z","iopub.status.idle":"2022-07-07T11:23:14.140200Z","shell.execute_reply.started":"2022-07-07T11:23:12.844752Z","shell.execute_reply":"2022-07-07T11:23:14.139230Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"wiki_ex = df_wiki_books.iloc[21]['body_text']\ndoc = nlp(wiki_ex)\nprint(doc)","metadata":{"execution":{"iopub.status.busy":"2022-07-07T11:23:14.141798Z","iopub.execute_input":"2022-07-07T11:23:14.142167Z","iopub.status.idle":"2022-07-07T11:23:14.577018Z","shell.execute_reply.started":"2022-07-07T11:23:14.142131Z","shell.execute_reply":"2022-07-07T11:23:14.574729Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('All entity types that spacy recognised from the document above')\nset([ent.label_ for ent in doc.ents])","metadata":{"execution":{"iopub.status.busy":"2022-07-07T11:23:14.581086Z","iopub.execute_input":"2022-07-07T11:23:14.581366Z","iopub.status.idle":"2022-07-07T11:23:14.590352Z","shell.execute_reply.started":"2022-07-07T11:23:14.581340Z","shell.execute_reply":"2022-07-07T11:23:14.589279Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Persons from the document above')\nprint(set([ent for ent in doc.ents if ent.label_ == 'PERSON']))\nprint('Organizations from the document above')\nprint(set([ent for ent in doc.ents if ent.label_ == 'ORG']))","metadata":{"execution":{"iopub.status.busy":"2022-07-07T11:23:14.592060Z","iopub.execute_input":"2022-07-07T11:23:14.592495Z","iopub.status.idle":"2022-07-07T11:23:14.602093Z","shell.execute_reply.started":"2022-07-07T11:23:14.592447Z","shell.execute_reply":"2022-07-07T11:23:14.600827Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"displacy.render(doc, style=\"ent\", jupyter=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-07T11:23:14.603800Z","iopub.execute_input":"2022-07-07T11:23:14.604220Z","iopub.status.idle":"2022-07-07T11:23:14.624387Z","shell.execute_reply.started":"2022-07-07T11:23:14.604187Z","shell.execute_reply":"2022-07-07T11:23:14.623555Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"heading\">\n   <h1><span style=\"color: Black\">Training Custom Model for Keywords Extraction</span></h1>\n</div>","metadata":{}},{"cell_type":"code","source":"nlp_start_df = pd.read_csv('../input/nlp-getting-started/train.csv')\n# taking one example sentence \nex = nlp_start_df.loc[129]['text']\n\ngenerator = pipeline(\"ner\",\n                     model=\"dslim/bert-base-NER\",)\ngenerator(ex)","metadata":{"execution":{"iopub.status.busy":"2022-07-07T11:23:14.625420Z","iopub.execute_input":"2022-07-07T11:23:14.625875Z","iopub.status.idle":"2022-07-07T11:23:43.960450Z","shell.execute_reply.started":"2022-07-07T11:23:14.625839Z","shell.execute_reply":"2022-07-07T11:23:43.959459Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_dir = '../input/entity-annotated-corpus'\ndf = pd.read_csv(f'{data_dir}/ner_dataset.csv',encoding= 'unicode_escape')\ndf['Sentence #'] = df['Sentence #'].ffill()\ndf_gr = df.groupby('Sentence #').agg(lambda x: list(x))\n\ndf_gr.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-07T11:23:43.962056Z","iopub.execute_input":"2022-07-07T11:23:43.962410Z","iopub.status.idle":"2022-07-07T11:23:46.762438Z","shell.execute_reply.started":"2022-07-07T11:23:43.962374Z","shell.execute_reply":"2022-07-07T11:23:46.761464Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"content\">\nIn our case, we'll need columns `Word` (tokenized sentence) and `Tag` (entities in the sentence). Also, in order to fine-tune the model, we'll need to have the same entities as the model is trained on. Let's print entities from our data set and from the pretrained model.\n</div>","metadata":{}},{"cell_type":"code","source":"tags = []\nfor tag in df_gr['Tag'].to_list():\n    tags.extend(tag)\nprint('Entities in our data set')\nprint(set(tags))","metadata":{"execution":{"iopub.status.busy":"2022-07-07T11:23:46.763760Z","iopub.execute_input":"2022-07-07T11:23:46.764551Z","iopub.status.idle":"2022-07-07T11:23:46.794782Z","shell.execute_reply.started":"2022-07-07T11:23:46.764494Z","shell.execute_reply":"2022-07-07T11:23:46.793751Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model_checkpoint_default = \"dslim/bert-base-NER\"\nmodel = AutoModelForTokenClassification.from_pretrained(model_checkpoint_default)\n\nprint('Entities from the pretrained model')\nmodel.config.id2label","metadata":{"execution":{"iopub.status.busy":"2022-07-07T11:23:46.796186Z","iopub.execute_input":"2022-07-07T11:23:46.796537Z","iopub.status.idle":"2022-07-07T11:23:48.848952Z","shell.execute_reply.started":"2022-07-07T11:23:46.796487Z","shell.execute_reply":"2022-07-07T11:23:48.847832Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"entity_mapping = {\n'O':0,'B-per':3, 'I-per':4, 'B-org':5, 'I-org':6,'B-geo':7, 'I-geo':8,\n'B-art':1, 'B-eve':1 , 'B-gpe':1, 'B-nat':1, 'B-tim':1,\n'I-art':1, 'I-eve':1 , 'I-gpe':1, 'I-nat':1, 'I-tim':1,\n}","metadata":{"execution":{"iopub.status.busy":"2022-07-07T11:23:48.850613Z","iopub.execute_input":"2022-07-07T11:23:48.850987Z","iopub.status.idle":"2022-07-07T11:23:48.857700Z","shell.execute_reply.started":"2022-07-07T11:23:48.850948Z","shell.execute_reply":"2022-07-07T11:23:48.856417Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class Keyword_Dataset_formatter:\n    def __init__(self, df):\n        # input is annotated data frame\n        self.texts = df['Word'].to_list()\n        self.tags = df['Tag'].to_list()\n    \n    def __len__(self):\n        return len(self.texts)\n    \n    def __getitem__(self, item):\n        text = self.texts[item]\n        tags = self.tags[item]\n        \n        ids = []\n        target_tag =[]\n        \n        # tokenize words and define tags accordingly\n        # running -> [run, ##ning]\n        # tags - ['O', 'O']\n        for i, s in enumerate(text):\n            inputs = tokenizer.encode(s, add_special_tokens=False)\n            input_len = len(inputs)\n            ids.extend(inputs)\n            target_tag.extend([entity_mapping[tags[i]]] * input_len)\n        \n        # truncate\n        ids = ids[:MAX_LEN - 2]\n        target_tag = target_tag[:MAX_LEN - 2]\n        \n        # add special tokens\n        ids = [101] + ids + [102]\n        target_tag = [0] + target_tag + [0]\n        mask = [1] * len(ids)\n        token_type_ids = [0] * len(ids)\n        \n        # construct padding\n        padding_len = MAX_LEN - len(ids)\n        ids = ids + ([0] * padding_len)\n        mask = mask + ([0] * padding_len)\n        token_type_ids = token_type_ids + ([0] * padding_len)\n        target_tag = target_tag + ([0] * padding_len)\n        \n        return {'input_ids': torch.tensor(ids, dtype=torch.long),\n                'attention_mask': torch.tensor(mask, dtype=torch.long),\n                'token_type_ids': torch.tensor(token_type_ids, dtype=torch.long),\n                'labels': torch.tensor(target_tag, dtype=torch.long)\n               }","metadata":{"execution":{"iopub.status.busy":"2022-07-07T11:23:48.859134Z","iopub.execute_input":"2022-07-07T11:23:48.859995Z","iopub.status.idle":"2022-07-07T11:23:48.875432Z","shell.execute_reply.started":"2022-07-07T11:23:48.859956Z","shell.execute_reply":"2022-07-07T11:23:48.874431Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\ndf_train, df_val = train_test_split(df_gr, test_size=0.2, random_state=42)\ndf_val, df_test = train_test_split(df_val, test_size=0.5, random_state=42)\n\nmodel_checkpoint = \"dslim/bert-base-NER\"\ntokenizer = AutoTokenizer.from_pretrained(model_checkpoint)\n\nMAX_LEN = 128\n\ndata_train = Keyword_Dataset_formatter(df_train)\ndata_val = Keyword_Dataset_formatter(df_val)\ndata_test = Keyword_Dataset_formatter(df_test)\n\n# initialize DataLoader used to return batches for training/validation\nloader_train = torch.utils.data.DataLoader(\n    data_train, batch_size=32, num_workers=2\n)\n\nloader_val = torch.utils.data.DataLoader(\n    data_val, batch_size=32, num_workers=2\n)","metadata":{"execution":{"iopub.status.busy":"2022-07-07T11:23:48.876701Z","iopub.execute_input":"2022-07-07T11:23:48.877108Z","iopub.status.idle":"2022-07-07T11:23:52.312160Z","shell.execute_reply.started":"2022-07-07T11:23:48.877073Z","shell.execute_reply":"2022-07-07T11:23:52.311220Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"os.environ[\"TOKENIZERS_PARALLELISM\"] = \"false\"","metadata":{"execution":{"iopub.status.busy":"2022-07-07T11:23:52.313525Z","iopub.execute_input":"2022-07-07T11:23:52.313874Z","iopub.status.idle":"2022-07-07T11:23:52.319175Z","shell.execute_reply.started":"2022-07-07T11:23:52.313838Z","shell.execute_reply":"2022-07-07T11:23:52.317993Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"param_optimizer = list(model.classifier.named_parameters())\noptimizer_grouped_parameters = [{\"params\": [p for n, p in param_optimizer]}]\noptimizer = AdamW(\n    optimizer_grouped_parameters,\n    lr=3e-5,\n    eps=1e-12\n)\n## full finetuning\n#optimizer = AdamW(model.parameters())\n\ndevice = torch.device(\"cuda\") if torch.cuda.is_available() else torch.device(\"cpu\")\nmodel.to(device)\n\nnum_epochs = 3\nnum_training_steps = num_epochs * len(loader_train)\nlr_scheduler = get_scheduler(\n    \"linear\",\n    optimizer=optimizer,\n    num_warmup_steps=0,\n    num_training_steps=num_training_steps\n)\n\n\nprogress_bar = tqdm(range(num_training_steps))\nfor epoch in range(num_epochs):\n    model.train()\n    final_loss = 0\n    predictions , true_labels = [], []\n    for batch in loader_train:\n        batch = {k: v.to(device) for k, v in batch.items()}\n        outputs = model(**batch)\n        loss = outputs.loss\n        loss.backward()\n        true_labels.extend(batch['labels'].detach().cpu().numpy().ravel())\n        predictions.extend(np.argmax(outputs[1].detach().cpu().numpy(), axis=2).ravel())\n        \n        optimizer.step()\n        lr_scheduler.step()\n        optimizer.zero_grad()\n        progress_bar.update(1)\n        final_loss+=loss.item()\n        \n    print(f'Training loss: {final_loss/len(loader_train)}')\n    print('Training F1: {}'.format(f1_score(predictions, true_labels, average='macro')))\n    print(f'Training acc: {accuracy_score(predictions, true_labels)}')\n    print('*'*20)\n    \n    model.eval()\n    final_loss = 0\n    predictions , true_labels = [], []\n    for batch in loader_val:\n        batch = {k: v.to(device) for k, v in batch.items()}\n        outputs = model(**batch)\n        final_loss+=outputs.loss.item()\n        true_labels.extend(batch['labels'].detach().cpu().numpy().ravel())\n        predictions.extend(np.argmax(outputs[1].detach().cpu().numpy(), axis=2).ravel())\n    print(f'Validation loss: {final_loss/len(loader_val)}')\n    print('Vallidation F1: {}'.format(f1_score(predictions, true_labels, average='macro')))\n    print(f'Validaton acc: {accuracy_score(predictions, true_labels)}')\n    print('*'*20)","metadata":{"execution":{"iopub.status.busy":"2022-07-07T11:23:52.320951Z","iopub.execute_input":"2022-07-07T11:23:52.321554Z","iopub.status.idle":"2022-07-07T11:46:31.489901Z","shell.execute_reply.started":"2022-07-07T11:23:52.321518Z","shell.execute_reply":"2022-07-07T11:46:31.488548Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_sentence = \"\"\"\nDoug and I are praying for the dozens of people who have been hospitalized and those who were lost today in Highland Park Illinois.\nThis shooting is an unmistakable reminder that more should be done to address gun violence in our country.\n\"\"\"\n#Vice President Kamala Harris Tweet 5th July 2022\n\ntokenized_sentence = tokenizer.encode(test_sentence)\ninput_ids = torch.tensor([tokenized_sentence]).cuda()\nwith torch.no_grad():\n    output = model(input_ids)\nlabel_indices = np.argmax(output[0].to('cpu').numpy(), axis=2)\n\ntokens = tokenizer.convert_ids_to_tokens(input_ids.to('cpu').numpy()[0])\nnew_tokens, new_labels = [], []\n\nfor token, label_idx in zip(tokens, label_indices[0]):\n    if token.startswith(\"##\"):\n        new_tokens[-1] = new_tokens[-1] + token[2:]\n    else:\n        new_labels.append(label_idx)\n        new_tokens.append(token)","metadata":{"execution":{"iopub.status.busy":"2022-07-07T11:46:31.491903Z","iopub.execute_input":"2022-07-07T11:46:31.492272Z","iopub.status.idle":"2022-07-07T11:46:31.514536Z","shell.execute_reply.started":"2022-07-07T11:46:31.492229Z","shell.execute_reply":"2022-07-07T11:46:31.513648Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dict1={'B-PER':'B-PERSON','B-LOC':'B-LOCATION','O':'Other','I-LOC':'I-LOCATION',}\ntags=[]\nwords=[]\nfor token, label in zip(new_tokens, new_labels):\n    tags.append(dict1[model.config.id2label[label]])\n    words.append(token)\n    #print(token,model.config.id2label[label],)","metadata":{"execution":{"iopub.status.busy":"2022-07-07T11:46:31.515898Z","iopub.execute_input":"2022-07-07T11:46:31.516454Z","iopub.status.idle":"2022-07-07T11:46:31.527272Z","shell.execute_reply.started":"2022-07-07T11:46:31.516414Z","shell.execute_reply":"2022-07-07T11:46:31.526313Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"lst=[]\ntotal_words=[]\ntemp=0\nfor key,i in enumerate(words[1:-1]):\n    #print(i)\n    if i not in total_words:\n        re_string = r\"\\b({})\\b\".format(i)\n        start=re.search(re_string, test_sentence).start()\n        end=re.search(re_string, test_sentence).end()\n        dict1={}\n        dict1[\"start\"]=start+temp\n        dict1[\"end\"]=end\n        dict1[\"label\"]=tags[1:-1][key]\n        dict1[\"Word\"]=i\n        total_words.append(i)\n\n        lst.append(dict1)\n        \n    else:\n        start=start+len(words[1:-1][key-1])\n        end=end+len(i)+1\n\n        dict1={}\n        dict1[\"start\"]=start\n        dict1[\"end\"]=end\n        dict1[\"label\"]=tags[1:-1][key]+'2'\n        dict1[\"Word\"]=i\n        total_words.append(i)\n\n        lst.append(dict1)","metadata":{"execution":{"iopub.status.busy":"2022-07-07T11:47:20.993363Z","iopub.execute_input":"2022-07-07T11:47:20.994298Z","iopub.status.idle":"2022-07-07T11:47:21.008530Z","shell.execute_reply.started":"2022-07-07T11:47:20.994259Z","shell.execute_reply":"2022-07-07T11:47:21.007403Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ex = [{\"text\": test_sentence,\n       \"ents\": lst,\n       \"title\": None}]\ncols = {'B-PERSON': '#dad1f6','I-LOCATION': '#adcfad','OTHER2': '#fbbf9a',\\\n        'B-LOCATION': '#bdf2fa','I-GENE_OR_GENOME': '#eea69e','OTHER': \"linear-gradient(90deg, #aa9cfc, #fc9ce7)\"}\noptions = {\"colors\": cols,}\n\nhtml = displacy.render(ex, style=\"ent\",options=options, manual=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-07T11:47:24.940006Z","iopub.execute_input":"2022-07-07T11:47:24.940361Z","iopub.status.idle":"2022-07-07T11:47:24.949125Z","shell.execute_reply.started":"2022-07-07T11:47:24.940331Z","shell.execute_reply":"2022-07-07T11:47:24.948112Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id =topic11> </a>\n<div class=\"heading\">\n   <h1><span style=\"color: black\">References and Usefull Resources</span></h1>\n</div>","metadata":{}},{"cell_type":"markdown","source":"* https://spacy.io/usage/linguistic-features#named-entities\n* https://www.kaggle.com/abhishek/entity-extraction-model-using-bert-pytorch\n* https://huggingface.co/course/chapter3/4?fw=pt\n* https://www.depends-on-the-definition.com/named-entity-recognition-with-bert/\n* https://www.lighttag.io/blog/sequence-labeling-with-transformers/example\n* https://en.wikipedia.org/wiki/Named-entity_recognition\n* https://towardsdatascience.com/named-entity-recognition-with-nltk-and-spacy-8c4a7d88e7da\n* https://www.guru99.com/pos-tagging-chunking-nltk.html\n* https://www.geeksforgeeks.org/nlp-extracting-named-entities/\n* https://colab.research.google.com/github/huggingface/notebooks/blob/master/examples/token_classification.ipynb#scrollTo=545PP3o8IrJV\n* https://discuss.huggingface.co/search?q=token%20classification","metadata":{}}]}