{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"在这个核心中，我想说明我是如何在构建深度学习NLP模型时提出有意义的预处理的。\n\n我从两条黄金法则开始:\n\n1.  **当你有预先训练好的嵌入时，不要使用标准的预处理步骤，如词干提取或停词删除** \n\n你们中的一些人在进行基于单词计数的特征提取(例如TFIDF)时可能会使用标准的预处理步骤，例如删除停止词、词干提取等。\n\n原因很简单:你丢失了有价值的信息，这将有助于你的神经网络解决问题。\n\n2. **让你的词汇尽可能接近嵌入**\n\n我会把重点放在这个笔记本上，如何做到这一点。举个例子，我用GoogleNews预训练的嵌入，这个选择没有更深层次的原因。","metadata":{"_uuid":"5fb07eae7b445e3bf358222065133a144b7c4ade"}},{"cell_type":"markdown","source":"我们从一个简洁的小技巧开始，它使我们能够在对pandas Dataframe应用函数时看到进度条","metadata":{"_uuid":"ec55ed96a9cab0ac75b22f9710c0393b17cfefd2"}},{"cell_type":"code","source":"import pandas as pd\nfrom tqdm import tqdm\ntqdm.pandas()\n","metadata":{"_uuid":"b378958a9606ac48fe0dc54e24bed4cd503e0ac7","execution":{"iopub.status.busy":"2023-05-15T04:06:55.678864Z","iopub.execute_input":"2023-05-15T04:06:55.679589Z","iopub.status.idle":"2023-05-15T04:06:55.715386Z","shell.execute_reply.started":"2023-05-15T04:06:55.679527Z","shell.execute_reply":"2023-05-15T04:06:55.714527Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Lets load our data","metadata":{"_uuid":"6e432bb170a329c79f58f78527bbbe4e857b2c41"}},{"cell_type":"code","source":"train = pd.read_csv(\"../input/train.csv\")\ntest = pd.read_csv(\"../input/test.csv\")\nprint(\"Train shape : \",train.shape)\nprint(\"Test shape : \",test.shape)","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-05-15T04:06:55.716904Z","iopub.execute_input":"2023-05-15T04:06:55.717428Z","iopub.status.idle":"2023-05-15T04:07:01.811766Z","shell.execute_reply.started":"2023-05-15T04:06:55.717209Z","shell.execute_reply":"2023-05-15T04:07:01.810706Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"我将使用下面的函数来跟踪我们的训练词汇表，它遍历我们所有的文本并计算包含单词的出现次数。","metadata":{"_uuid":"59fd43e3727bd265df08d2d8254230c655c51f9d"}},{"cell_type":"code","source":"def build_vocab(sentences, verbose =  True):\n    \"\"\"\n    :参数句:单词列表的列表\n    :返回:单词及其计数的字典\n    \"\"\"\n    vocab = {}\n    for sentence in tqdm(sentences, disable = (not verbose)):\n        for word in sentence:\n            try:\n                vocab[word] += 1\n            except KeyError:\n                vocab[word] = 1\n    return vocab","metadata":{"_uuid":"3e050f2fa9668c765466d766cfd79720cf6cc819","execution":{"iopub.status.busy":"2023-05-15T04:07:01.813478Z","iopub.execute_input":"2023-05-15T04:07:01.813866Z","iopub.status.idle":"2023-05-15T04:07:01.820897Z","shell.execute_reply.started":"2023-05-15T04:07:01.813794Z","shell.execute_reply":"2023-05-15T04:07:01.819104Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"因此，让我们填充词汇表并显示前5个元素及其计数。注意，现在我们可以使用progresess_apply来查看进度条","metadata":{"_uuid":"a453d10106aea7f6bf3f9bd952f2e02dd7a7472c"}},{"cell_type":"code","source":"sentences = train[\"question_text\"].progress_apply(lambda x: x.split()).values\nvocab = build_vocab(sentences)\nprint({k: vocab[k] for k in list(vocab)[:5]})","metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","execution":{"iopub.status.busy":"2023-05-15T04:07:01.822828Z","iopub.execute_input":"2023-05-15T04:07:01.823435Z","iopub.status.idle":"2023-05-15T04:07:17.644801Z","shell.execute_reply.started":"2023-05-15T04:07:01.823139Z","shell.execute_reply":"2023-05-15T04:07:17.643579Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"接下来，我们导入稍后要在模型中使用的嵌入。我在这里使用GoogleNews作为例证。","metadata":{"_uuid":"c5c1c2c96aecd40e969f13913e850a36d8e3a13d"}},{"cell_type":"code","source":"from gensim.models import KeyedVectors\n\nnews_path = '../input/quora-insincere-questions-classification/embeddings.zip/GoogleNews-vectors-negative300/GoogleNews-vectors-negative300.bin'\nembeddings_index = KeyedVectors.load_word2vec_format(news_path, binary=True)","metadata":{"_uuid":"807e734e0ce617480c824f8bf26f0672f38397d5","execution":{"iopub.status.busy":"2023-05-15T04:11:00.870729Z","iopub.execute_input":"2023-05-15T04:11:00.871089Z","iopub.status.idle":"2023-05-15T04:11:00.905332Z","shell.execute_reply.started":"2023-05-15T04:11:00.871023Z","shell.execute_reply":"2023-05-15T04:11:00.903502Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"接下来，我定义一个函数来检查词汇表和嵌入之间的交集。它将输出一个超出词汇表(oov)的单词列表，我们可以用它来改进我们的预处理","metadata":{"_uuid":"b5f9048e53a5cb91bdaa518d33fda756f1038562"}},{"cell_type":"code","source":"import operator \n\ndef check_coverage(vocab,embeddings_index):\n    a = {}\n    oov = {}\n    k = 0\n    i = 0\n    for word in tqdm(vocab):\n        try:\n            a[word] = embeddings_index[word]\n            k += vocab[word]\n        except:\n\n            oov[word] = vocab[word]\n            i += vocab[word]\n            pass\n\n    print('Found embeddings for {:.2%} of vocab'.format(len(a) / len(vocab)))\n    print('Found embeddings for  {:.2%} of all text'.format(k / (k + i)))\n    sorted_x = sorted(oov.items(), key=operator.itemgetter(1))[::-1]\n\n    return sorted_x","metadata":{"_uuid":"f35a7213fc9a7e80a7c210d11b3a8094d3a8e07e","execution":{"iopub.status.busy":"2023-05-15T04:07:19.040145Z","iopub.status.idle":"2023-05-15T04:07:19.040640Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"oov = check_coverage(vocab,embeddings_index)","metadata":{"_uuid":"42c2b82740ac82f5678a46b7c82b5525616c1304","execution":{"iopub.status.busy":"2023-05-15T04:07:19.041529Z","iopub.status.idle":"2023-05-15T04:07:19.042074Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"哎哟，只有24%的词汇会有嵌入，这使得21%的数据或多或少毫无用处。所以让我们看看并开始改进。为此，我们可以很容易地查看排名前50位的单词。","metadata":{"_uuid":"9ff8e507e571e671ed4249a95fb0eb339098446e"}},{"cell_type":"code","source":"oov[:10]","metadata":{"_uuid":"de0a78175815413d204d9b7f70b5facba97a7b58","execution":{"iopub.status.busy":"2023-05-15T04:07:19.042933Z","iopub.status.idle":"2023-05-15T04:07:19.043415Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"排在第一位的是“to”。为什么?原因很简单，因为在训练GoogleNews嵌入时，“to”被删除了。我们稍后会解决这个问题，因为现在我们要注意标点的分裂，因为这似乎也是一个问题。但是我们该怎么处理标点符号呢?我们是想删除它还是把它当作一个符号呢?我会说:这要看情况。如果令牌有嵌入，保留它，如果没有，我们就不再需要它了。让我们检查一下:","metadata":{"_uuid":"dba213d3a9fefce81a3ab41e2ba95846f87da997"}},{"cell_type":"code","source":"'?' in embeddings_index","metadata":{"_uuid":"69129cad1f08fda5912051d273db012f9adb081a","execution":{"iopub.status.busy":"2023-05-15T04:07:19.043992Z","iopub.status.idle":"2023-05-15T04:07:19.044372Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"'&' in embeddings_index","metadata":{"_uuid":"63deed2445710aef37e504a4d252b5ef02598f21","execution":{"iopub.status.busy":"2023-05-15T04:07:19.044881Z","iopub.status.idle":"2023-05-15T04:07:19.045243Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"有趣。虽然“&”在谷歌新闻嵌入中，但“?”不在。因此，我们基本上定义了一个函数来分离“&”并删除其他标点符号。","metadata":{"_uuid":"361cbf32713d89b11d62c444cc87968460c454b8"}},{"cell_type":"code","source":"def clean_text(x):\n\n    x = str(x)\n    for punct in \"/-'\":\n        x = x.replace(punct, ' ')\n    for punct in '&':\n        x = x.replace(punct, f' {punct} ')\n    for punct in '?!.,\"#$%\\'()*+-/:;<=>@[\\\\]^_`{|}~' + '“”’':\n        x = x.replace(punct, '')\n    return x","metadata":{"_uuid":"afd570d1160b826a84c4c7af76950fd7a30e4471","execution":{"iopub.status.busy":"2023-05-15T04:07:19.045839Z","iopub.status.idle":"2023-05-15T04:07:19.046620Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train[\"question_text\"] = train[\"question_text\"].progress_apply(lambda x: clean_text(x))\nsentences = train[\"question_text\"].apply(lambda x: x.split())\nvocab = build_vocab(sentences)","metadata":{"_uuid":"b47c945e6f02c2a9f72aedd72da4681695cb20dd","execution":{"iopub.status.busy":"2023-05-15T04:07:19.047329Z","iopub.status.idle":"2023-05-15T04:07:19.047795Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"oov = check_coverage(vocab,embeddings_index)\n","metadata":{"_uuid":"9f90189a592e93f563d84d68a02ff48cc08315c0","execution":{"iopub.status.busy":"2023-05-15T04:07:19.048485Z","iopub.status.idle":"2023-05-15T04:07:19.048919Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"好了!仅通过处理标点符号，我们就能将嵌入率从24%提高到57%。好了，我们来检查一下这50个单词。","metadata":{"_uuid":"a06bdc24e539ea6d3fd17afd637c978fcaf36e7f"}},{"cell_type":"code","source":"oov[:10]","metadata":{"_uuid":"5dff2b2ed016358b734b49e4d23d22cce746f43a","execution":{"iopub.status.busy":"2023-05-15T04:07:19.049619Z","iopub.status.idle":"2023-05-15T04:07:19.049982Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Hmm seems like numbers also are a problem. Lets check the top 10 embeddings to get a clue.","metadata":{"_uuid":"55f7d6d8960ec03f0cf73e7e97d317065bf940dd"}},{"cell_type":"code","source":"for i in range(10):\n    print(embeddings_index.index2entity[i])","metadata":{"_uuid":"385ba7b4747ca28e1fea5d25c6a6edbeb9532cc6","execution":{"iopub.status.busy":"2023-05-15T04:07:19.050702Z","iopub.status.idle":"2023-05-15T04:07:19.051069Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"嗯，为什么“##”在里面?很简单，因为作为一个重处理，所有大于9的数字都被哈希替换了。例如，15变成##，123变成##，或者15.80€变成##。因此，让我们模拟这个预处理步骤来进一步提高我们的嵌入覆盖率","metadata":{"_uuid":"c2340ec62a0f679ffffe94bfafc323ecf517f4c9"}},{"cell_type":"code","source":"import re\n\ndef clean_numbers(x):\n\n    x = re.sub('[0-9]{5,}', '#####', x)\n    x = re.sub('[0-9]{4}', '####', x)\n    x = re.sub('[0-9]{3}', '###', x)\n    x = re.sub('[0-9]{2}', '##', x)\n    return x","metadata":{"_uuid":"bbc48ed45237a1096e29fac87c52d258b38fa1ee","execution":{"iopub.status.busy":"2023-05-15T04:07:19.051806Z","iopub.status.idle":"2023-05-15T04:07:19.052242Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train[\"question_text\"] = train[\"question_text\"].progress_apply(lambda x: clean_numbers(x))\nsentences = train[\"question_text\"].progress_apply(lambda x: x.split())\nvocab = build_vocab(sentences)","metadata":{"_uuid":"48da9d68b5d8911247a17a927649d33877570e0e","execution":{"iopub.status.busy":"2023-05-15T04:07:19.053027Z","iopub.status.idle":"2023-05-15T04:07:19.053931Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"oov = check_coverage(vocab,embeddings_index)","metadata":{"_uuid":"fb6a821cfdb49be15023fe4579c063394eaedbb9","execution":{"iopub.status.busy":"2023-05-15T04:07:19.054623Z","iopub.status.idle":"2023-05-15T04:07:19.055267Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"好了!又增加了3%。现在和处理标点符号一样多，但每一点都有帮助。让我们检查一下单词","metadata":{"_uuid":"38618cca73324e598a8610c2ac8648e6aa7eb42b"}},{"cell_type":"code","source":"oov[:20]","metadata":{"_uuid":"14638ae1defa12014ed7b6be63f3a3bbebb9ed3e","execution":{"iopub.status.busy":"2023-05-15T04:07:19.055930Z","iopub.status.idle":"2023-05-15T04:07:19.056526Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"好了，现在我们来处理在使用美式/英式词汇时常见的拼写错误，并将一些“现代”单词替换为“社交媒体”。为了完成这项任务，我使用了我不久前在stack overflow上找到的一个多正则表达式脚本。此外，我们将简单地删除单词“a”，“to”，“and”和“of”，因为在训练GoogleNews嵌入时，这些词显然已经被下采样了。\n","metadata":{"_uuid":"17ca5651a141e21572e4a5abb23a7b368b5337ba"}},{"cell_type":"code","source":"def _get_mispell(mispell_dict):\n    mispell_re = re.compile('(%s)' % '|'.join(mispell_dict.keys()))\n    return mispell_dict, mispell_re\n\n\nmispell_dict = {'colour':'color',\n                'centre':'center',\n                'didnt':'did not',\n                'doesnt':'does not',\n                'isnt':'is not',\n                'shouldnt':'should not',\n                'favourite':'favorite',\n                'travelling':'traveling',\n                'counselling':'counseling',\n                'theatre':'theater',\n                'cancelled':'canceled',\n                'labour':'labor',\n                'organisation':'organization',\n                'wwii':'world war 2',\n                'citicise':'criticize',\n                'instagram': 'social medium',\n                'whatsapp': 'social medium',\n                'snapchat': 'social medium'\n\n                }\nmispellings, mispellings_re = _get_mispell(mispell_dict)\n\ndef replace_typical_misspell(text):\n    def replace(match):\n        return mispellings[match.group(0)]\n\n    return mispellings_re.sub(replace, text)","metadata":{"_uuid":"f1c3084cd7503c9148b681de897e0d38d7a90741","execution":{"iopub.status.busy":"2023-05-15T04:07:19.057351Z","iopub.status.idle":"2023-05-15T04:07:19.057964Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train[\"question_text\"] = train[\"question_text\"].progress_apply(lambda x: replace_typical_misspell(x))\nsentences = train[\"question_text\"].progress_apply(lambda x: x.split())\nto_remove = ['a','to','of','and']\nsentences = [[word for word in sentence if not word in to_remove] for sentence in tqdm(sentences)]\nvocab = build_vocab(sentences)","metadata":{"_uuid":"c887a0a6f7498724a2c377e9e29fac683ea3a59b","execution":{"iopub.status.busy":"2023-05-15T04:07:19.058695Z","iopub.status.idle":"2023-05-15T04:07:19.059268Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"oov = check_coverage(vocab,embeddings_index)","metadata":{"_uuid":"63aba43671245f1ebc87540a8bd49a174791c378","execution":{"iopub.status.busy":"2023-05-15T04:07:19.060059Z","iopub.status.idle":"2023-05-15T04:07:19.060597Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We see that although we improved on the amount of embeddings found for all our text from 89% to 99%. Lets check the oov words again ","metadata":{"_uuid":"a5543990c3d10f9c77d58840a66bbcad11f233b5"}},{"cell_type":"code","source":"oov[:20]","metadata":{"_uuid":"21f6da9413c717da789f1068387ac83e3bfd85ae","execution":{"iopub.status.busy":"2023-05-15T04:07:19.061249Z","iopub.status.idle":"2023-05-15T04:07:19.061812Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Looks good. No obvious oov words there we could quickly fix.\nThank you for reading and happy kaggling","metadata":{"_uuid":"fe32acac0c1280d58ef66ac1611f63bbb98f0aec"}},{"cell_type":"code","source":"","metadata":{"_uuid":"73df67fc23dd4e8d999d6c0cbb3f70cb952747fb","trusted":true},"execution_count":null,"outputs":[]}]}