{"cells":[{"metadata":{"_uuid":"984e151005f06e7b416470b5967491b29cc1d654"},"cell_type":"markdown","source":"Every once in a while I mess around with [spaCy](https://spacy.io/) to see what it can do. It comes with a rich set of features, including it's own pretrained language models. Some of these models include word embeddings like the ones we're given. A spaCy model might be useful here as a way to bring in additional vectors and dictionaries. Let's see what we have."},{"metadata":{"_uuid":"2d3bf02b397fc92b269f76728da2706a7d7503f3"},"cell_type":"markdown","source":"## spaCy Vectors\n\nFirst let's look at spaCy's \"large model\". The documentation says the model uses GloVe vectors trained on Common Crawl. We are already given the 300d vectors as a text file. I'll compare a vector from spaCy with a vector in the text file to see if there's a difference."},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true,"scrolled":true},"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport spacy as sp\n\nnlp_lg = sp.load('en_core_web_lg')\nnlp_lg","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"2c43e14f2d6a8156dbaba80e4356de19764af363"},"cell_type":"code","source":"# get spacy vector\nlgword = nlp_lg(\"and\")\nlgvec =   \",\".join(lgword.vector[0:10].round(5).astype(str))\n\n# get glove vector\nglv = pd.read_csv('../input/embeddings/glove.840B.300d/glove.840B.300d.txt', header=None, sep=' ', skiprows=2, nrows=5, index_col=[0])\nglvec = glv.loc['and', 0:10].round(5).astype(str).str.cat(sep=' ')\n\nprint(lgword.vector.shape[0], \"\\n\",\n      lgvec, \"\\n\",\n      glv.shape[1], \"\\n\",\n      glvec)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"38961fb23bd85907d8505f7f166ff47dd28d4cd1"},"cell_type":"markdown","source":"Vectors are the same for the word \"and\" as well as other words I checked. Oh well, no new information here. \n\nLet's check the small model, which \"only includes context-sensitive tensors\". The docs say that the small models don't work as well. Maybe they can be helpful anyway as an additional source of information. "},{"metadata":{"trusted":true,"_uuid":"c2ab4006b10ac04ec61fcec67c369127eb8cf2b6"},"cell_type":"code","source":"nlp_sm = sp.load('en_core_web_sm')\nsmword = nlp_sm(\"and\")\nsmvec = \",\".join(smword.vector[0:10].round(5).astype(str))\n\nprint(smword.vector.shape[0], \"\\n\",\n       smvec)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"19e9299abfe879e11bc32206f72fd01970f3ca60"},"cell_type":"markdown","source":"The GloVe vector and spaCy vector (or rank1 tensor if you insist) are indeed different. The model may be a useful addition to other vectors.\n\nspaCy will also calculate vectors for an entire question. The model tokenizes the string according to its own rules, gets vectors for each word, and averages them to get a single vector. "},{"metadata":{"trusted":true,"_uuid":"6d36ec1ecc2d8edcdd95c2229008bf86828c03d8"},"cell_type":"code","source":"e = nlp_lg('Why are aliens so smart?')\n\nprint(e.vector.shape, \"\\n\",\n       e.vector[0:10])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"805af672ca42b87a21de907c982be94fecf6ed91"},"cell_type":"markdown","source":"## Language Features\n\nspaCy has a host of other language features. You can use a built-in similarity function to compare questions. If I remember correctly, it's a shorthand function for cosine similarity."},{"metadata":{"trusted":true,"_uuid":"1cb7953820285c8d2cae272b0cbc708689573408"},"cell_type":"code","source":"c = nlp_sm('What capital city is the prettiest?') \nd = nlp_sm('Which country has the nicest people?')\ne = nlp_sm('Why are aliens so smart?')\n\nprint(\"\\n\", c.similarity(d),\n        c.similarity(e))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"abefebf801a628243ae50fb352028e329b3fbd66"},"cell_type":"markdown","source":"The model can also lemmatize, assign parts of speech, find dependencies and otherwise annotate text."},{"metadata":{"trusted":true,"_uuid":"e487911c282e39cca9a89e7cbf53b372b54f8d30"},"cell_type":"code","source":"df = pd.DataFrame({\"text\": [tokens.text for tokens in d], \n                   \"lemmatized\": [tokens.lemma_ for tokens in d],\n                   \"part of speech\": [tokens.pos_ for tokens in d],\n                  \"stop word\": [tokens.is_stop for tokens in d]})\ndisplay(df)                 \nsp.displacy.render(d, style='dep', jupyter=True, options={'compact':60})","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c1fba0a7b811984007e88749863e3e9b3ff5ce06"},"cell_type":"markdown","source":"## A Simple Model\nHere's a simple model to get average vectors for each question and train a logistic regression model. The vectors are the same as the GloVe vectors we're given, except there are fewer words available. \n\nCalculating vectors for each question is time consuming. It's 4-5 times faster to get vectors for each unique token and manually average them."},{"metadata":{"trusted":true,"_uuid":"df68eba0141fdbc108374a5354d77ca84a55eefc","_kg_hide-output":false,"scrolled":true},"cell_type":"code","source":"#%% import\nimport time\nimport numpy as np\nimport pandas as pd\nimport spacy as sp\nnlp_lg = sp.load('en_core_web_lg')\nfrom sklearn.model_selection import StratifiedKFold, train_test_split\nfrom sklearn.metrics import f1_score\nfrom sklearn.linear_model import LogisticRegression\nfrom tqdm import tqdm\n\n\n# get train data\ntrain = pd.read_csv('../input/train.csv', nrows=30_000)  #limiting the data for time's sake\ntrain['question_text'] = train.question_text.str.replace('?', ' ?')\ntrain['question_text'] = train.question_text.str.replace('.', ' .')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"e727abae8e368a1912f28016212e10b7b0038a17"},"cell_type":"code","source":"tstacked = pd.DataFrame(train.question_text.str.split(expand=True).stack(), \n                columns=['token'])\n\ntlist = tstacked.token.unique().tolist()\nvlist = [nlp_lg(str).vector for str in tqdm(tlist)]\nlookup = dict(zip(tlist, vlist))\n\ntstacked['vec'] = tstacked.token.map(lookup)\n\ncolnames = ['t'+str(i) for i in range(300)]\ntstacked[colnames] = pd.DataFrame(tstacked.vec.values.tolist(), \n                            index=tstacked.index)\ntstacked.drop(['token', 'vec'], axis=1, inplace=True)\n\ndel tlist\ndel vlist\ndel lookup\ntagg = tstacked.groupby(level=0).apply(np.mean)\ndel tstacked\n\nX_vecs = tagg.values\ny = train.target.values\ndel tagg","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-output":true,"trusted":true,"_uuid":"0bdc9da2d80bed6fbe9564d71e1702ccd3dd1567"},"cell_type":"code","source":"# Logistic Regression\nskf = StratifiedKFold(n_splits=5, shuffle=True, random_state=911)\ntrain_pred = np.zeros(train.shape[0])\nfor train_idx, val_idx in skf.split(X_vecs, y):\n    X_train, y_train  = X_vecs[train_idx], y[train_idx]\n    X_val, y_val = X_vecs[val_idx], y[val_idx]\n    model = LogisticRegression(solver='saga', class_weight='balanced', \n                                    C=0.5, max_iter=250, verbose=1, n_jobs=-1) #seed not set\n    model.fit(X_train, y_train)\n    val_pred = model.predict_proba(X_val)\n    train_pred[val_idx] = val_pred[:,1]\n    \n\nprint(\"finding best threshold\")\nbest_thresh = 0.0\nbest_score = 0.0\nfor thresh in np.arange(0, 1, 0.01):\n    score = f1_score(y, train_pred > thresh)\n    if score > best_score:\n        best_thresh = thresh\n        best_score = score\nprint(best_thresh, best_score)","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true,"_uuid":"8f4b687804d4cdd3314298300478b269a4194449"},"cell_type":"code","source":"print(best_thresh, best_score)","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-output":true,"trusted":true,"_uuid":"9be6ea9d5af0751c5ebc16f755ee3e949f00306b"},"cell_type":"code","source":"# predict on test set\ntest = pd.read_csv('../input/test.csv', index_col=['qid'])\ntest.head()\nX_test = test.question_text.tolist()\nX_testvecs = np.array([nlp_lg(text).vector for text in tqdm(X_test)])\n\ntrounds = 3\npreds_test = np.zeros(len(X_test))\nfor i in range(trounds):\n    model = LogisticRegression(solver='saga', class_weight='balanced', \n                                    C=0.5, max_iter=250, verbose=1, n_jobs=-1, random_state=40*i)\n    model.fit(X_vecs, y)\n    preds_test += lgr.predict_proba(X_testvecs)[:, 1] / trounds\n\n    \n# submit\nsub = pd.read_csv('../input/sample_submission.csv', index_col=['qid'])\nsub['prediction'] = preds_test > best_thresh\nsub.to_csv('submission.csv')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"73b04ddf352de7d9e3777b6a97f61a5b0cd50699"},"cell_type":"markdown","source":"This is a basic model trained on part of the data. So far the results have not been as good as logistic regression with tf-idf features. I think using the other annotations (parts of speech, etc.) as meta-features might be the best way to use spaCy.\n\nAlternately, you can get vectors for each word in a question and assemble them for a Keras model. See https://www.kaggle.com/enerrio/scary-nlp-with-spacy-and-keras for an example.\n\n\n\n## spaCy's CNN\nspaCy also has it's own CNN for text classification. I haven't dug into it very much, but it seems to work at a basic level. Here is an example of how to format the data and train a classifier from scratch. You can also run the code (with modifications) on a GPU for better speed. "},{"metadata":{"trusted":true,"_uuid":"a6d7ea1b0fe2b7cae503e317d507761963459cd0"},"cell_type":"code","source":"#%% import\nimport numpy as np\nimport pandas as pd\nimport spacy as sp\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import f1_score\n\n\n# get data and make json format for spacy\ntrain = pd.read_csv('../input/train.csv', nrows=10_000)  ## using part of the data again\ntexts = train.question_text.tolist()\ncats = train.target.apply(lambda t: {'cats': {'Insincere': t == 1}}).tolist()\ntrain_texts, dev_texts, train_cats, dev_cats = train_test_split(texts, cats, \n        test_size=0.2, random_state=90)\ntrain_data = list(zip(train_texts, train_cats))\nprint(\"Example format \\n\", train_data[0:10])\n\n\n#%% set up the pipeline\nnlp_bl = sp.blank('en') \nnlp_bl.vocab.vectors.name = 'spacy_pretrained_vectors'\ntextcat = nlp_bl.create_pipe('textcat')\nnlp_bl.add_pipe(textcat, last=True)\ntextcat.add_label('Insincere')\n\n\n# train\nn_iter = 10\nother_pipes = [pipe for pipe in nlp_bl.pipe_names if pipe != 'textcat']\nwith nlp_bl.disable_pipes(*other_pipes):  #only train textcat\n    optimizer = nlp_bl.begin_training()\n    print(\"Training the model...\")\n    for i in range(n_iter):\n        losses = {}\n        batches = sp.util.minibatch(train_data, size=sp.util.compounding(4., 32., 1.001))\n        for batch in batches:\n            texts, annotations = zip(*batch)\n            nlp_bl.update(texts, annotations, sgd=optimizer, drop=0.2,\n                        losses=losses)\n        print(\"iter {} loss: {:4f}\".format(i, losses['textcat']))\n\n        \n# evaluate model\npreds = []\ndocs = (nlp_bl(text) for text in dev_texts)\nfor doc in docs:\n    pred = doc.cats['Insincere']\n    preds.append(pred)\n    \ntruths = [val['Insincere'] for val in [dc['cats'] for dc in dev_cats]]\n\n#%% find best threshold\nbest_thresh = 0.0\nbest_score = 0.0\nfor thresh in np.arange(0, 1, 0.01):\n    score = f1_score(truths, preds > thresh)\n    if score > best_score:\n        best_thresh = thresh\n        best_score = score\nprint(best_thresh, best_score)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"e8f932dc1316e0f1be333efcdaa49d7564355cfe"},"cell_type":"markdown","source":"Again, this model needs to run longer on more data to seee what it can do. Hope to see some clever uses of spaCy in other kernels!"}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}