{"cells":[{"metadata":{"_uuid":"34c2dd93223452a4204531810872d3a4ad98884b"},"cell_type":"markdown","source":"DATA MINING PROJECT\n\nEbubekir Durukal"},{"metadata":{"_uuid":"423448e1b8fbb790629323e9308c6f4dd2f36ef4"},"cell_type":"markdown","source":"\nQuora Insincere Questions Classification is an\nongoing kaggle competition. This project aims to\ncome up with the best algorithm to detect insincere\n(i.e. toxic, troll) questions on Quora. Quora is a\nsocial media platform where people from all around\nthe globe ask questions and answer existing\nquestions. As always, some people want to troll the\nwebsite by asking things that are not real questions.\nThese statements may include harassments and\nracist expressions. An automated system to\neliminate these questions is crucial for Quora."},{"metadata":{"_uuid":"60f249e6e71a5e32c9d3248959f16dfb5e25fb3b"},"cell_type":"markdown","source":"İmporting necessary packages:"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"_kg_hide-output":false},"cell_type":"code","source":"import os\nimport time\nfrom sklearn import tree\nfrom sklearn.metrics import accuracy_score\nfrom sklearn.naive_bayes import GaussianNB\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nfrom tqdm import tqdm\nimport string\nimport math\nfrom sklearn.model_selection import train_test_split\nfrom sklearn import metrics\nfrom nltk.corpus import stopwords\nimport seaborn as sns\nimport matplotlib.pylab as pylab\nimport matplotlib.pyplot as plt\nfrom keras.preprocessing.text import Tokenizer\nfrom keras.preprocessing.sequence import pad_sequences\nfrom keras.layers import Dense, Input, LSTM, Embedding, Dropout, Activation, CuDNNGRU, Conv1D\nfrom keras.layers import Bidirectional, GlobalMaxPool1D\nfrom keras.models import Model\nfrom keras import initializers, regularizers, constraints, optimizers, layers","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-output":false,"_uuid":"aa4aad03a0377fb806aa4c42df17f3deb69b944e"},"cell_type":"markdown","source":"Reading the dataset and printing insights"},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true,"_kg_hide-input":false,"_kg_hide-output":false},"cell_type":"code","source":"train = pd.read_csv(\"../input/train.csv\")\ntest = pd.read_csv(\"../input/test.csv\")\nprint('describe the data: \\n')\nprint(train.describe())\nprint('info about the train data:\\n')\nprint(train.info())\nprint('shape of the train data:\\n')\nprint(train.shape)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"6c7c04b60d369c659db0d1a0930ebb5c380a9333"},"cell_type":"markdown","source":"For Exploratory Data Analysis, training and test\nfiles were read using pandas package. Then, for\ndata cleaning, I checked the dataset for null values.\nThere were no null valued elements in the dataset."},{"metadata":{"trusted":true,"_uuid":"28ac7fe8c736d3ccb7aeee0e17783d2c43471ee3"},"cell_type":"code","source":"train.isnull().sum()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a8a7b877db3230a2b017a6bbcc1429a554af1aed"},"cell_type":"markdown","source":"Plotting data imbalance:"},{"metadata":{"trusted":true,"_uuid":"8223357767d63bb0e5048b5b39d3dd80da19609a","_kg_hide-output":true},"cell_type":"code","source":"print(train.where(train ['target']==1).count())\ntrain[\"target\"].value_counts()\neng_stopwords = set(stopwords.words(\"english\"))\nprint(len(eng_stopwords))\nax=sns.countplot(x='target',hue=\"target\", data=train  ,linewidth=5,edgecolor=sns.color_palette(\"dark\", 3))\nplt.title('Is data set imbalance?');\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5740549db1475375bc4d7c77cfdca375171a9589"},"cell_type":"markdown","source":"I applied Data Augmentation for better\nresults. Before this, the dataset included only\n‘question text’. But after this, its attributes\ncontained 'num_words', 'num_unique_words',\n'num_chars', 'num_stopwords', 'num_punctuations',\n'num_words_upper', 'num_words_title',\n'mean_word_len'"},{"metadata":{"_uuid":"b1b1255d2b6bf9d2cfba7fbab70502d7d3cb2ac2"},"cell_type":"markdown","source":"After having a better understanding of the data,\nData was splitted into train and validation sets. 10%\nof the data is separated as validation set and models\ndid not see this data during training process."},{"metadata":{"trusted":true,"_uuid":"002c9e72c0c7871470bdf639cb6df1fd9cbb5658"},"cell_type":"code","source":"## split to train and val\ntrain, val = train_test_split(train, test_size=0.1, random_state=2018)\n\n## some config values \nembed_size = 300 # how big is each word vector\nmax_features = 50000 # how many unique words to use (i.e num rows in embedding vector)\nmaxlen = 100 # max number of words in a question to use\n\n## fill up the missing values\ntrain_X = train[\"question_text\"].fillna(\"_na_\").values\nval_X = val[\"question_text\"].fillna(\"_na_\").values\ntest_X = test[\"question_text\"].fillna(\"_na_\").values\n\n## Tokenize the sentences\ntokenizer = Tokenizer(num_words=max_features)\ntokenizer.fit_on_texts(list(train_X))\ntrain_X = tokenizer.texts_to_sequences(train_X)\nval_X = tokenizer.texts_to_sequences(val_X)\ntest_X = tokenizer.texts_to_sequences(test_X)\n\n## Pad the sentences \ntrain_X = pad_sequences(train_X, maxlen=maxlen)\nval_X = pad_sequences(val_X, maxlen=maxlen)\ntest_X = pad_sequences(test_X, maxlen=maxlen)\n\n## Get the target values\ntrain_y = train['target'].values\nval_y = val['target'].values","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"79857e868bf263619477fcc6fe0d3b8fc42e2703"},"cell_type":"markdown","source":"Embeddings for better algorithm:"},{"metadata":{"trusted":true,"_uuid":"d4f80b3ee9cc83f8ffd66028c7ed8d62ffdb60b1"},"cell_type":"code","source":"EMBEDDING_FILE = '../input/embeddings/glove.840B.300d/glove.840B.300d.txt'\ndef get_coefs(word,*arr): return word, np.asarray(arr, dtype='float32')\nembeddings_index = dict(get_coefs(*o.split(\" \")) for o in open(EMBEDDING_FILE))\n\nall_embs = np.stack(embeddings_index.values())\nemb_mean,emb_std = all_embs.mean(), all_embs.std()\nembed_size = all_embs.shape[1]\n\nword_index = tokenizer.word_index\nnb_words = min(max_features, len(word_index))\nembedding_matrix = np.random.normal(emb_mean, emb_std, (nb_words, embed_size))\nfor word, i in word_index.items():\n    if i >= max_features: continue\n    embedding_vector = embeddings_index.get(word)\n    if embedding_vector is not None: embedding_matrix[i] = embedding_vector\n        \n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"a73a063cd3184e65ec607dea1e5ac95980f0cd22"},"cell_type":"markdown","source":"Defining models and  training them"},{"metadata":{"trusted":true,"_uuid":"a493b73a74c6fccec640c83a569d328ca31d9172"},"cell_type":"code","source":"\n\ninp = Input(shape=(maxlen,))\nx = Embedding(max_features, embed_size, weights=[embedding_matrix])(inp)\nx = Bidirectional(CuDNNGRU(64, return_sequences=True))(x)\nx = GlobalMaxPool1D()(x)\nx = Dense(16, activation=\"relu\")(x)\nx = Dropout(0.1)(x)\nx = Dense(1, activation=\"sigmoid\")(x)\nmodel = Model(inputs=inp, outputs=x)\nmodel.compile(loss='binary_crossentropy', optimizer='adam', metrics=['accuracy'])\nprint(model.summary())","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d6ec390e093df44cbc52ca6b9394505d58643f1b"},"cell_type":"code","source":"from sklearn import tree\n\ntree = tree.DecisionTreeClassifier()\ngauss = GaussianNB()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"71043b13942b7222c97accc4a98596d2cbde2feb"},"cell_type":"code","source":"#model.fit(train_X, y_train, batch_size=512, epochs=2, validation_data=(X_test, y_test))\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"e1825040949f7cefa761571b3d1a91f421e6606c"},"cell_type":"code","source":"tree.fit(train_X, train_y)\ngauss.fit(train_X,train_y)\n\ny_pred=tree.predict(val_X)\ny_pred_1=gauss.predict(val_X)\n#y_pred_2=model.predict(val_X)\n\nprint('decision tree accuracy is: ',accuracy_score(val_y, y_pred))\nprint('naive bayes accuracy is: ',accuracy_score(val_y, y_pred_1))\n#print('nn accuracy is: ',accuracy_score(val_y, y_pred_2))\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"a4a9effbd16de83a6522984b9cebda47562655e2"},"cell_type":"code","source":"pred_test_y = model.predict([test_X], batch_size=1024, verbose=1)\npred_test_y = (pred_test_y>0.35).astype(int)\nout_df = pd.DataFrame({\"qid\":test[\"qid\"].values})\nout_df['prediction'] = pred_test_y\nout_df.to_csv(\"submission.csv\", index=False)\n\n            \nprint(out_df.tail())","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"bbfe9b5036488a75f1d5a48f65af4fc7bc4a074c"},"cell_type":"code","source":"train[\"num_words\"] = train[\"question_text\"].apply(lambda x: len(str(x).split()))\ntest[\"num_words\"] = test[\"question_text\"].apply(lambda x: len(str(x).split()))\n\ntrain[\"num_unique_words\"] = train[\"question_text\"].apply(lambda x: len(set(str(x).split())))\ntest[\"num_unique_words\"] = test[\"question_text\"].apply(lambda x: len(set(str(x).split())))\n\ntrain[\"num_chars\"] = train[\"question_text\"].apply(lambda x: len(str(x)))\ntest[\"num_chars\"] = test[\"question_text\"].apply(lambda x: len(str(x)))\n\ntrain[\"num_stopwords\"] = train[\"question_text\"].apply(lambda x: len([w for w in str(x).lower().split() if w in eng_stopwords]))\ntest[\"num_stopwords\"] = test[\"question_text\"].apply(lambda x: len([w for w in str(x).lower().split() if w in eng_stopwords]))\n\ntrain[\"num_punctuations\"] =train['question_text'].apply(lambda x: len([c for c in str(x) if c in string.punctuation]) )\ntest[\"num_punctuations\"] =test['question_text'].apply(lambda x: len([c for c in str(x) if c in string.punctuation]) )\n\ntrain[\"mean_word_len\"] = train[\"question_text\"].apply(lambda x: np.mean([len(w) for w in str(x).split()]))\ntest[\"mean_word_len\"] = test[\"question_text\"].apply(lambda x: np.mean([len(w) for w in str(x).split()]))\n\nsns.violinplot(data=train,x=\"target\", y=\"num_words\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d37fd29101ffc5514a220b893b97c992be5e2a8c"},"cell_type":"code","source":"train.hist(figsize=(15,20))\nplt.figure()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"04b86922315ae3b0c51c476c259c4a6827b8e8cf"},"cell_type":"markdown","source":"Conclusions:\n\nFor KDD process, methods like PCA cannot be\nused since this dataset does not deliver a lot of\nattributes. I understood this better with data\nvisualization. \n\nDataset was checked for null values.\n\nFor this reason, data augmentation\ntechnique was used. For data transformation, question text was turned into integer values. \n\nData selection was done using train test splitting method.\n\nIt turned out that the best\naccuracy was obtained when zeroR used. But it was\nbecause the dataset was quite imbalanced. It is not\nsuggested to use zeroR in real world applications.\nKeras neural networks used a considerable amount\nof time even though it used GPU.\n\nThe results for the models are measured using cross\nvalidation evaluation metric. According to sklearn\naccuracy results:\n\nAccuracy for decision tree: 0.894136551246101\nAccuracy for naive bayes: 0.9301280233571736\nAccuracy for zeroR: 0.9373536238276778\nAccuracy for neural nets: 0.782782378238732\n\nAfter obtaining accuracy results, I submitted the\nresults to the competition. The results took 1497th\nplace among 2652 participants."},{"metadata":{"_uuid":"f8bdf4e3e869d0f42e17310f2ad3909410f52e59"},"cell_type":"markdown","source":"References:\nhttps://www.kaggle.com/sudalairajkumar/a-look-at-different-embeddings"}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}