{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<h1> Top 5% Solution - CountVectorizer vs Tfidf Vectorizer vs spaCy vs BERT </h1>","metadata":{}},{"cell_type":"markdown","source":"In this notebook, I am going to detail my approach to the Predicting Disaster Tweets NLP competition with all of the models and pipelines I tested. This entry landed me on the top 5% of the leaderboard (when you exclude obviously leaked entries) at the time of this writing.\n\nThis beginner competition is a great opportunity to learn how to approach an NLP problem, which requires a completely different train of thought from numeric tabular data.\n\nSo if you are also getting started with NLP, I recommend that you fork this notebook and/or code alongside it so you can get familiar with how text problems are tackled.\n\n**Now, here are the models and pipelines combinations I tested:**\n\n\n<li>Count Vectorizer + Logistic Regression\n<li>Count Vectorizer + Naive Bayes\n<li>Count Vectorizer + XGB\n<li>Tfidf + Logistic Regression\n<li>Tfidf + Naive Bayes\n<li>Tfidf + XGB\n<li>spaCy + Logistic Regression\n<li>spaCy + Naive Bayes\n<li>spaCy + XGB\n<li>BERT\n\n    \nAnd without further ado, let's get started...\n\n\n---","metadata":{}},{"cell_type":"markdown","source":"<h2> Importing Libraries, Functions and Datasets </h2>","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport re\nsns.set_style('darkgrid')\n\nfrom wordcloud import WordCloud, STOPWORDS\nfrom sklearn.model_selection import KFold\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.pipeline import Pipeline\nfrom sklearn.metrics import f1_score\nfrom sklearn.compose import ColumnTransformer\nfrom sklearn.base import BaseEstimator, TransformerMixin\nfrom sklearn import feature_extraction\nfrom sklearn.naive_bayes import GaussianNB\nfrom xgboost import XGBClassifier\nimport warnings\nwarnings.filterwarnings('ignore')\n\n\nfrom keras import Input\nfrom keras.layers import Dense, Dropout\nimport tensorflow as tf\nfrom tensorflow.keras.optimizers import Adam\n\nfrom transformers import BertTokenizer, TFBertModel","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class Cross_valid:\n    def __init__(self,train_data,n_splits):\n        self.best_score = 0\n        self.best_model = list()\n        self.best_model_std = 0\n        self.feature_cols = None\n        \n        train_data['kfold'] = -1\n\n        kf = KFold(n_splits=n_splits,shuffle=True)\n        for fold, (train_idx, valid_idx) in enumerate(kf.split(train_data)):\n            train_data.loc[valid_idx,'kfold'] = fold\n        \n        self.data = train_data\n        self.n_splits = n_splits\n        \n    \n    def run_model(self, pipeline_steps, model, feature_cols, target_col, valid_pipeline = False):\n        data = self.data.copy()\n        current_model_scores = np.array([])\n        print(f'Now running the {str(model)} model...\\n')\n        for fold in range(self.n_splits):\n            x_train = data[data['kfold'] != fold][feature_cols].copy()\n            x_valid = data[data['kfold'] == fold][feature_cols].copy()\n            \n            y_train = data[data['kfold'] != fold][target_col].copy()\n            y_valid = data[data['kfold'] == fold][target_col].copy()\n            \n            x_train_pipelined = Pipeline(steps=pipeline_steps).fit_transform(x_train)\n            x_valid_pipelined = Pipeline(steps=pipeline_steps).transform(x_valid)\n                \n            \n\n            \n            model.fit(x_train_pipelined,y_train)\n            prediction_valid = model.predict(x_valid_pipelined)\n            \n            current_fold_score = f1_score(y_valid,prediction_valid)\n            current_model_scores = np.append(current_model_scores,current_fold_score)\n            \n            print(f'Fold {fold} validation score: {current_fold_score}')\n        \n        \n        avg_score = current_model_scores.mean()\n        std = current_model_scores.std()\n        print(f'Finished running {str(model)} model...\\nAverage score: {round(avg_score,5)} \\nStandard deviation: {round(std,5)}')\n        \n        if avg_score > self.best_score:\n            self.best_score = avg_score\n            self.best_model_std = std\n            self.best_model = [pipeline_steps,model]\n            self.feature_cols = feature_cols\n            print(f'This is the new baseline model! New benchmark value of {round(self.best_score,5)}')\n        else:\n            print(f'This model failed to beat the benchmark score of {round(self.best_score,5)}...')\n\n            \nclass DenseTransformer(TransformerMixin):\n\n    def fit(self, X, y=None, **fit_params):\n        return self\n\n    def transform(self, X, y=None, **fit_params):\n        return X.todense()\n\n    \nclass Reshaper(TransformerMixin):\n\n    def fit(self, X, y=None, **fit_params):\n        return self\n\n    def transform(self, X, y=None, **fit_params):\n        X = X.to_numpy()\n        n_size = X[0].shape[0]\n        X = X.reshape(-1,1)\n        X = np.concatenate(np.concatenate(X, axis = 0), axis = 0).reshape(-1, n_size)\n        return X\n\n    \ndef word_cloud(data,text_column,title=None):\n\n    comment_words = ''\n    stopwords = set(STOPWORDS)\n\n    # iterate through the csv file\n    for val in data[text_column]:\n\n        # typecaste each val to string\n        val = str(val)\n\n        # split the value\n        tokens = val.split()\n\n        # Converts each token into lowercase\n        for i in range(len(tokens)):\n            tokens[i] = tokens[i].lower()\n\n        comment_words += \" \".join(tokens)+\" \"\n\n    wordcloud = WordCloud(width = 1500, height = 800,\n                    background_color ='white',\n                    stopwords = stopwords,\n                    min_font_size = 10).generate(comment_words)\n\n    # plot the WordCloud image\t\t\t\t\t\n    plt.figure(figsize = (15, 6), facecolor = None)\n    plt.imshow(wordcloud)\n    plt.axis(\"off\")\n    plt.tight_layout(pad = 5)\n    plt.suptitle(title,fontsize=25,fontweight=100)\n\n    plt.show()\n    \n    \ndef clean_text(text, tokenizer, stopwords):\n    \"\"\"Pre-process text and generate tokens\n\n    Args:\n        text: Text to tokenize.\n\n    Returns:\n        Tokenized text.\n    \"\"\"\n    text = str(text).lower()  # Lowercase words\n    text = re.sub(r\"\\[(.*?)\\]\", \"\", text)  # Remove [+XYZ chars] in content\n    text = re.sub(r\"\\s+\", \" \", text)  # Remove multiple spaces in content\n    text = re.sub(r\"\\w+…|…\", \"\", text)  # Remove ellipsis (and last word)\n    text = re.sub(r\"(?<=\\w)-(?=\\w)\", \" \", text)  # Replace dash between words\n    text = re.sub(\n        f\"[{re.escape(string.punctuation)}]\", \"\", text\n    )  # Remove punctuation\n\n    tokens = tokenizer(text)  # Get tokens from text\n    tokens = [t for t in tokens if not t in stopwords]  # Remove stopwords\n    tokens = [\"\" if t.isdigit() else t for t in tokens]  # Remove digits\n    tokens = [t for t in tokens if len(t) > 1]  # Remove short tokens\n    return tokens\n\ndef f1(pred, y):\n    score = f1_score(y,pred)\n    return 'f1', score\n\nclass Text_clearer(TransformerMixin):\n    def fit(self, X, y=None, **fit_params):\n        return self\n\n    def transform(self, X, y=None, **fit_params):\n        X = X.apply(word_tokenize)\n        return X\n    \n\n# Used the bert encoder and bert model found in the link below with very little changes to them\n# https://www.kaggle.com/code/dhruv1234/huggingface-tfbertmodel/notebook?scriptVersionId=34225936\n    \n\ndef bert_encode(data,maximum_length) :\n  input_ids = []\n  attention_masks = []\n  \n\n  for i in range(len(data.text)):\n      encoded = tokenizer.encode_plus(\n        \n        data.text[i],\n        add_special_tokens=True,\n        max_length=maximum_length,\n        pad_to_max_length=True,\n        \n        return_attention_mask=True,\n        \n      )\n      \n      input_ids.append(encoded['input_ids'])\n      attention_masks.append(encoded['attention_mask'])\n  return np.array(input_ids),np.array(attention_masks)\n\n\n\n\n\ndef create_model(bert_model):\n  input_ids = tf.keras.Input(shape=(60,),dtype='int32')\n  attention_masks = tf.keras.Input(shape=(60,),dtype='int32')\n  \n  output = bert_model([input_ids,attention_masks])\n  output = output[1]\n  output = tf.keras.layers.Dense(32,activation='relu')(output)\n  output = tf.keras.layers.Dropout(0.2)(output)\n\n  output = tf.keras.layers.Dense(1,activation='sigmoid')(output)\n  model = tf.keras.models.Model(inputs = [input_ids,attention_masks],outputs = output)\n  model.compile(Adam(lr=6e-6), loss='binary_crossentropy', metrics=['accuracy'])\n  return model","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test = pd.read_csv('../input/nlp-getting-started/test.csv')\ntest.head(4)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train = pd.read_csv('../input/nlp-getting-started/train.csv')\ntrain.head(4)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_sub = pd.read_csv('../input/nlp-getting-started/sample_submission.csv')\nsample_sub.head(4)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h2> Exploratory Data Analysis </h2>","metadata":{}},{"cell_type":"code","source":"print('Null values by column on the train dataset:\\n')\nfor column, null_pct in ((train.isna().sum())/(train.shape[0])).items():\n    print(f'{round(null_pct*100,2)}% of values on the {column} column are null')","metadata":{"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Null values by column on the test dataset:\\n')\nfor column, null_pct in ((test.isna().sum())/(test.shape[0])).items():\n    print(f'{round(null_pct*100,2)}% of values on the {column} column are null')","metadata":{"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Disaster tweets occurence of keywords:')\ntrain[train['target'] == 1]['keyword'].value_counts()","metadata":{"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Non-disaster tweets occurence of keywords:')\ntrain[train['target'] == 0]['keyword'].value_counts()","metadata":{"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.catplot(data=train,kind='count',x='target',aspect=1.8)\nplt.suptitle('Number of Tweets by Target',fontsize=15,fontweight=100)","metadata":{"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"word_cloud(train[train['target'] == 0],'text','Non-disater Tweets Wordcloud')","metadata":{"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"word_cloud(train[train['target'] == 1],'text','Disaster Tweets Wordcloud')","metadata":{"_kg_hide-input":false,"_kg_hide-output":true,"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h2> Pre-processing <\\h2>","metadata":{}},{"cell_type":"markdown","source":"In this step, I am simply loading spaCy pre-trained doc vectors. Since spaCy is already pre-trained, this way of transforming the data will not result in data leakage.\n\n**This is really the only pre-processing I did outside of the pipelines.**","metadata":{}},{"cell_type":"code","source":"import en_core_web_lg\nnlp = en_core_web_lg.load()\n\n\ndef get_vector(x):\n    doc = nlp(x)\n    vec = doc.vector\n    return vec\n\ntrain['doc_vec'] = train['text'].apply(lambda x: get_vector(x))\ntest['doc_vec'] = test['text'].apply(lambda x: get_vector(x))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h2> Testing Models </h2>\n\n<h3> Count Vectorizer + Logistic Regression </h3>","metadata":{}},{"cell_type":"code","source":"cv_pipeline = [('cv',feature_extraction.text.CountVectorizer(stop_words='english',max_features=1000,ngram_range=(1,3))),\n ('todense',DenseTransformer())]\n\nnb = GaussianNB()\nlr = LogisticRegression()\n\nprint('Creating 5-fold cross-validation object...')\nfive_fold = Cross_valid(train,5)\nfive_fold.run_model(cv_pipeline,lr,feature_cols = 'text',target_col = 'target')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h3> Count Vectorizer + Naive Bayes </h3>","metadata":{}},{"cell_type":"code","source":"five_fold.run_model(cv_pipeline,nb,feature_cols = 'text',target_col = 'target')","metadata":{"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h3> Tfidf + Logistic Regression </h3>","metadata":{}},{"cell_type":"code","source":"tfidf_pipeline = [('cv',feature_extraction.text.TfidfVectorizer(stop_words='english',max_features=1000,ngram_range=(1,3))),\n               ('todense',DenseTransformer())]\n\n\nfive_fold.run_model(tfidf_pipeline,lr,feature_cols = 'text',target_col = 'target')","metadata":{"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h3> Tfidf + Naive bayes </h3>","metadata":{}},{"cell_type":"code","source":"\nfive_fold.run_model(tfidf_pipeline,nb,feature_cols = 'text',target_col = 'target')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h3> Tfidf + XGBoost </h3>","metadata":{}},{"cell_type":"code","source":"xgb = XGBClassifier(n_jobs=4,maximize=True, feval=f1)\n\n\nfive_fold.run_model(tfidf_pipeline,xgb,feature_cols = 'text',target_col = 'target')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h3> CountVectorizer + XGBoost </h3>","metadata":{}},{"cell_type":"code","source":"five_fold.run_model(cv_pipeline,xgb,feature_cols = 'text',target_col = 'target')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h2> spaCy Pipeline + Logistic Regression <\\h2>","metadata":{}},{"cell_type":"code","source":"spacy_pipeline = [('reshape',Reshaper())]\n\nfive_fold.run_model(spacy_pipeline,lr,'doc_vec','target')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h2> spaCy Pipeline + Naive Bayes <\\h2>","metadata":{}},{"cell_type":"code","source":"spacy_pipeline = [('reshape',Reshaper())]\n\nfive_fold.run_model(spacy_pipeline,nb,'doc_vec','target')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h2> spaCy Pipeline + XGBoost <\\h2>","metadata":{}},{"cell_type":"code","source":"spacy_pipeline = [('reshape',Reshaper())]\n\nfive_fold.run_model(spacy_pipeline,xgb,'doc_vec','target')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h3> BERT </h3>\n\nI did not cross-validate the BERT model because it takes way too long to fit and predict and would probably give me 'out of memory' errors... But as you will be able to see shortly, it outperforms all of the other models.","metadata":{}},{"cell_type":"code","source":"# BERT has dozens of pre-trained models, I took the smallest one in order to prevent out of memory errors.\n# However, you can surely try more complex pre-trained models, which will probably give you an even better score\n\ntokenizer = BertTokenizer.from_pretrained('bert-base-uncased')\n\ntrain_input_ids, train_attention_masks = bert_encode(train,60)\ntest_input_ids, test_attention_masks = bert_encode(test,60)\n\nbert_model = TFBertModel.from_pretrained('bert-base-uncased',output_attentions=True)\nmodel = create_model(bert_model)\nhistory = model.fit([train_input_ids,train_attention_masks],train.target,validation_split=0.2, epochs=3,batch_size=50,workers=4,)\ny_pred = model.predict([test_input_ids,test_attention_masks])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h2> Final Model Training <\\h2>","metadata":{}},{"cell_type":"markdown","source":"And the winning pipeline for this problem was the BERT model, clearly outperforming all of the others.\n\nHope you have learned a thing or two by reading this notebook.\n\nThanks for your time and attention.\n\nGoodbye.","metadata":{}},{"cell_type":"code","source":"y_pred = np.round(y_pred).astype(int)\ny_pred = pd.DataFrame(y_pred)\nsubmission = pd.read_csv('/kaggle/input/nlp-getting-started/sample_submission.csv')\noutput = pd.DataFrame({'id':submission.id,'target':y_pred[0]})\noutput.to_csv('subm1.csv',index=False)\nprint('Done!')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}