{"cells":[{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load in \n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\nimport os\nimport json\nimport string\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n%matplotlib inline\n\nfrom sklearn import metrics\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.feature_extraction.text import TfidfVectorizer, CountVectorizer\nfrom sklearn.decomposition import TruncatedSVD\nfrom sklearn.model_selection import train_test_split\nimport lightgbm as lgb\n\npd.options.mode.chained_assignment = None\npd.options.display.max_columns = 999\n\nimport re\nimport os\nimport gc\nfrom collections import Counter\nprint(os.listdir(\"../input\"))\n\n# Any results you write to the current directory are saved as output.","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"print(os.listdir(\"../input/embeddings\"))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ebdcbd9460d92a992c23e87a8a0c67b5ca93e14e"},"cell_type":"markdown","source":"## Intro\n\nThis is just a brief notebook that does some exploration and modeling.  The objective is to predict whether a question asked on Quora is sincere or not. \nAn insincere question is said to have an intent to make a statement rather than look for helpful answers. Some characteristics that can signify that a question is insincere:\n\n- Has a non-neutral tone\n    - Has an exaggerated tone to underscore a point about a group of people\n    - Is rhetorical and meant to imply a statement about a group of people\n- Is disparaging or inflammatory\n    - Suggests a discriminatory idea against a protected class of people, or seeks confirmation of a stereotype\n    - Makes disparaging attacks/insults against a specific person or group of people\n    - Based on an outlandish premise about a group of people\n- Disparages against a characteristic that is not fixable and not measurable\n    - Isn't grounded in reality\n    - Based on false information, or contains absurd assumptions\n    - Uses sexual content (incest, bestiality, pedophilia) for shock value, and not to seek genuine answers\n\n<a id=\"toc\"></a>\n### Contents\n\n[Brief Exploration](#bexp)\n\n[Word Counts](#wc)\n\n[TF-IDF and Count Vecs](#tf)\n\n[LogReg Classifier with Naive Bayes and Thresholding](#lrc)"},{"metadata":{"_uuid":"2a72a50ce2c3bce39232a35a847dc5d8e035934e"},"cell_type":"markdown","source":"<a id=\"bexp\"></a>\n## Brief Exploration\n\n[Table of contents](#toc)"},{"metadata":{"trusted":true,"_uuid":"36a1efef1c8fbd7ca617135417395d299abf7a3e"},"cell_type":"code","source":"train = pd.read_csv(\"../input/train.csv\").set_index('qid')\ntest = pd.read_csv(\"../input/test.csv\").set_index('qid')\nprint(\"Train shape : \", train.shape)\nprint(\"Test shape : \", test.shape)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"efc9ca41c46bf8816ae6218e118be20815aa90f6"},"cell_type":"code","source":"train.head(10)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"6a952af242259f764bee023c32951f620fd2f5bd"},"cell_type":"markdown","source":"The train data set is pretty straight forward. It has the question text and the target readily available.\n\nSome intitial thoughts:\n- This is probably not a good instance to remove stop words as it could change the semantics.\n- Perhaps we can create low dimensional instances where we do remove them as well."},{"metadata":{"trusted":true,"_uuid":"a26045711e72ff3afe27b8f2950afaa0e0327ac7"},"cell_type":"code","source":"counts = Counter(train.target)\nprint(\"Insincere ratio : {:.3f}%\".format(counts[1]/len(train)*100))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"a5e0839f1e21f66df9c7b6c6d0024010f97ad807"},"cell_type":"markdown","source":"So there ratio is 6.187% of the questions are insincere. "},{"metadata":{"trusted":true,"_uuid":"d0f89419e669525f319419d336e34529f496507a"},"cell_type":"code","source":"fig, ax=plt.subplots(1,1,figsize=(12,6))\nsns.countplot(x=\"target\", data=train, ax=ax)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"3c436fea51b974c7c672d6bb06fa349c8385bdcd"},"cell_type":"code","source":"import nltk\nfrom nltk.corpus import stopwords\nstops = stopwords.words('english')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"e9a0e506e3b8403b75bb0f9f0a0ff82068b94e19"},"cell_type":"code","source":"## simple meta features\ndef meta_features(df):\n    df['num_words'] = df[\"question_text\"].apply(lambda x: len(str(x).split()))\n    df['num_nonstopwords'] = df[\"question_text\"].apply(lambda x: len([w for w in str(x).lower().split() if w in stops]))\n    df['num_punctuation'] = df['question_text'].apply(lambda x: len([c for c in str(x) if c in string.punctuation]) )\n    df['num_upper'] = df[\"question_text\"].apply(lambda x: len([w for w in str(x).split() if w.isupper()]))\n    df[\"mean_word_len\"] = df[\"question_text\"].apply(lambda x: np.mean([len(w) for w in str(x).split()]))\n    return df","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d26f2f768ab82950caab45fc7f732b2e1a1da940"},"cell_type":"code","source":"train = meta_features(train)\ntest = meta_features(test)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"0a7927b2c56e61e4dbb2e5257a6c1e9a1dccc717"},"cell_type":"code","source":"pd.options.display.float_format = '{:.4f}'.format\ntrain.drop('target',axis=1).describe()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"80f2fe4b65d2978a4c69aa141c19863f95e1f500"},"cell_type":"code","source":"test.describe()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"cca0010fbd79fd455cf934f703b7c8c6edb1e599"},"cell_type":"markdown","source":"If were comparing the train and tests sets with meta data, we see that there are some similarities.\n"},{"metadata":{"_uuid":"0ff1c073684809a1a9b02c1fb18f7597a3360f8c"},"cell_type":"markdown","source":"<a id=\"wc\"></a>\n## Word Counts\n\n[Table of contents](#toc)"},{"metadata":{"trusted":true,"_uuid":"e3e8505bc4c269b5b2bc8a8ad250a0c7cfbb7eb3"},"cell_type":"code","source":"def text_cleaner(text, sw=False):\n    text = re.sub(r\"[^A-Za-z0-9^\\s]\", \"\", text).lower()\n    text = re.sub(r'\\.{3}',\" \", text)#.split()\n    text = re.sub(r'\\x85',\" \", text).split()\n    if sw == True:\n        text = [w for w in text if not w in stops]    \n    text = \" \".join(text)\n    return(text)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"42eb4a75efaaecec9f622bcc5318333dacf82084"},"cell_type":"markdown","source":"Cleaned examples...."},{"metadata":{"trusted":true,"_uuid":"b55b4088c34373d561fddccf8a38f55d8f06dbab"},"cell_type":"code","source":"print(\"For Target == False ...\")\nprint(\"Normal ...\")\nprint(train.question_text.iloc[0])\nprint(\"Cleaned ...\")\nprint(text_cleaner(train.question_text.iloc[0]))\nprint(\"Clean no stopwords ...\")\nprint(text_cleaner(train.question_text.iloc[0], sw=True))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"81c1327689afeb944a25d6ebe30f0d54219fa34b"},"cell_type":"code","source":"# cnt = 0\n# for i in train.target:\n#     if i == True:\n#         print(cnt)\n#         break\n#     cnt+=1","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"6103e469e2942cf163c2b0bb80c098dd86a0c91a"},"cell_type":"code","source":"print(\"For Target == True ...\")\nprint(\"Normal ...\")\nprint(train.question_text.iloc[22])\nprint(\"Cleaned ...\")\nprint(text_cleaner(train.question_text.iloc[22]))\nprint(\"Clean no stopwords ...\")\nprint(text_cleaner(train.question_text.iloc[22], sw=True))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"01ff46e810b0d56bb0d6beacaa491535e0ac03bc"},"cell_type":"code","source":"train_text = []\nfor it in train['question_text']:\n    newT = text_cleaner(it, sw=True)\n    train_text.append(newT)\n    \ntest_text = []\nfor it in test['question_text']:\n    newT = text_cleaner(it, sw=True)\n    test_text.append(newT)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c374494b5bd05aee4cfd5957dce20bf336819ae6"},"cell_type":"code","source":"train['cleaned_text'] = train_text\ntest['cleaned_text'] = test_text","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"76e3728dd4b5dc805d0ad2b6b1f5dfaa32e3d43a"},"cell_type":"code","source":"insi = train[train[\"target\"] == 1]\nsinc = train[train[\"target\"] == 0]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c742ea8f6b9f119632c07b61fa2905c8be1bd9d1"},"cell_type":"code","source":"insi_tops = Counter(str(insi.cleaned_text).split()).most_common()[:30]\nlabs, vals = zip(*insi_tops)\nidx = np.arange(len(labs))\nwid=0.6\nfig, ax=plt.subplots(1,1,figsize=(14,8))\nax=plt.bar(idx, vals, wid, color='r')\nax=plt.xticks(idx - wid/8, labs, rotation=45, size=14)\nplt.title('Top 30 Counts of Most-Common Words Among Insincere Text');","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"4b2b784ae0a5793b487b3270e76b362b0fa7c4dc"},"cell_type":"code","source":"sinc_tops = Counter(str(sinc.cleaned_text).split()).most_common()[:30]\nlabs, vals = zip(*sinc_tops)\nidx = np.arange(len(labs))\nwid=0.6\nfig, ax=plt.subplots(1,1,figsize=(14,8))\nax=plt.bar(idx, vals, wid, color='g')\nax=plt.xticks(idx - wid/8, labs, rotation=45, size=14)\nplt.title('Top 30 Counts of Most-Common Words Among Sincere Text');","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"60b70756fb78150dbde9984073e7746bb1c21e5b"},"cell_type":"markdown","source":"__The issue above are those elipses.  I need to come back to this and figure out how to remove those without messing up the cleaner.__\n\n-  Some words can be found in both instances.\n- Insincere generally _looks_ like it might be negative."},{"metadata":{"trusted":true,"_uuid":"e9c7a90a76065023ed282a2ee770430d80de276d"},"cell_type":"code","source":"try:\n    del insi, sinc, train_text, test_text\nexcept:\n    pass\n\ngc.collect()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"9920c0a9d10cd7bf1a0767e967ed859e2696a8b4"},"cell_type":"markdown","source":"<a id=\"tf\"></a>\n## TF-IDF anc Count Vecs\n\n[Table of contents](#toc)"},{"metadata":{"trusted":true,"_uuid":"5e9e44ad9bbb8d9a9d45a57304c94c2e8f344056"},"cell_type":"code","source":"## I pulled this little goodie from https://www.kaggle.com/ryanzhang/tfidf-naivebayes-logreg-baseline\n## Very cool\nTOKENIZER = re.compile(f'([{string.punctuation}“”¨«»®´·º½¾¿¡§£₤‘’])')\ndef tokenize(s):\n    return TOKENIZER.sub(r' \\1 ', s).split()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"9700ee532701e940b783718a2a99e74b7d24f8f3"},"cell_type":"code","source":"tfidf = TfidfVectorizer(ngram_range=(1,4), tokenizer=tokenize, min_df=3,\n                        max_df=0.82,\n#                         max_features = 20000,\n                        strip_accents='unicode', use_idf=True,\n                        smooth_idf=True, sublinear_tf=True)\ntfidf.fit(train.question_text)\n\ncvec = CountVectorizer(ngram_range=(1,4), tokenizer=tokenize, min_df=3,\n                       max_df=0.82, \n#                        max_features=20000,\n                       strip_accents='unicode')\ncvec.fit(train.question_text)\n\n## still did the best with all features at max_df=0.33 for a score of 0.62","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"4f86590b103fe5bd6ae8310543856869893f45d3"},"cell_type":"code","source":"from scipy.sparse import hstack\n\ndef build_doc_feats_sparse(temp):\n    txt = temp.question_text\n    print(\"...tfidf\")\n    tf = tfidf.transform(txt)\n    print('...cvec')\n    cv = cvec.transform(txt)\n    print('...hstack')\n    x = hstack([tf,cv])\n    return x\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"464232ae97b9404ff2d4a0d449994eddb4022654"},"cell_type":"markdown","source":"__Building the features using TF-IDF, Count Vec, and Latent Sem. Analysis.__"},{"metadata":{"trusted":true,"_uuid":"b9b320655eb31d1b6328d6fda860ac76d9ce42f7"},"cell_type":"code","source":"print(\" ....train feats\")\nx_train = build_doc_feats_sparse(train)\nprint(\" ....test feats\")\nx_test = build_doc_feats_sparse(test)\ny_train = train['target']","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"6cd7459cdd0e8f86d13aee13f5d050b56ff34d86"},"cell_type":"code","source":"gc.collect()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"99a8d6f8796cb93694a28c7d325a809ba0af3c7c"},"cell_type":"code","source":"print(\"Shape of Train and Test:\")\nprint(x_train.shape, y_train.shape)\nprint(x_test.shape)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"a122a289260347335efbab968fb8f08487c27baa"},"cell_type":"code","source":"print(type(x_train),type(x_test))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"0b1dc8398c813b59a080c4f58504423be3722b29"},"cell_type":"markdown","source":"<a id=\"lrc\"></a>\n## LogReg Classifier with Naive Bayes and Thresholding\n\n[Table of contents](#toc)"},{"metadata":{"trusted":true,"_uuid":"649b67653bc68d37be99bbea9354064a08690fe9"},"cell_type":"code","source":"try:\n    del train, test\nexcept:\n    pass\ngc.collect()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d80f90b54f068873fd729b30d2eed31016ec2442"},"cell_type":"code","source":"## smaller values of C support stronger regularization\nlr_mod = LogisticRegression(solver='lbfgs', \n                            class_weight='balanced',\n                            C=0.5,\n                            max_iter=50,\n                            random_state=42, n_jobs=-1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"43744c54236c6a9b1f4e105917a08e9b4131cc06"},"cell_type":"code","source":"## from : https://www.kaggle.com/hung96ad/pytorch-starter\nfrom tqdm import tqdm\ndef threshold_search(y_true, y_proba):\n    best_threshold = 0\n    best_score = 0\n    for threshold in tqdm([i * 0.01 for i in range(100)]):\n        score = metrics.f1_score(y_true=y_true, y_pred=y_proba > threshold)\n        if score > best_score:\n            best_threshold = threshold\n            best_score = score\n    search_result = {'threshold': best_threshold, 'f1': best_score}\n    return search_result","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"4676241310c884759f5f0a37f62a41d82e1ce693"},"cell_type":"code","source":"kfolds = 3\nkd = 0\nlr_preds = 0\nval_preds = 0\n\ndef pr(X, y_i, y):\n    p = X[y==y_i].sum(0)\n    return (p+1) / ((y==y_i).sum()+1)\n\nths = []\n\nfor i in range(kfolds):\n    print('In kfold:',str(i+1))\n    xt,xv,yt,yv = train_test_split(x_train, y_train, test_size=0.15, random_state=(i*42))\n\n    r = np.log(pr(xt, 1, yt.ravel()) / pr(xt, 0, yt.ravel()))\n    x_nb = xt.multiply(r)\n    lr_mod.fit(x_nb, yt.ravel())\n    \n    val_prob = lr_mod.predict_proba(xv.multiply(r))\n    pred_prob = lr_mod.predict_proba(x_test.multiply(r))\n    lr_preds += pred_prob\n    val_preds += val_prob\n    kd += 1\n    print('=========================')\n    print(\"Validation Set Classification Report :\")\n    print(metrics.classification_report(yv, np.round(val_prob[:,1],0).astype(int)))\n    print(\"F1 Score : {:.4f}\".format(metrics.f1_score(yv, np.round(val_prob[:,1],0).astype(int))))\n    th_sr = threshold_search(yv, val_prob[:,1])\n    ths.append(th_sr)\n    print(\"Threshold Search F1 Result :\")\n    print(th_sr)\n    print('=========================')\n    \nlr_preds /= kfolds\nval_preds /= kfolds","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"08ac47c5e718ca831359f5b586c22d8634779e67"},"cell_type":"code","source":"gc.collect()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"8cbd0d8d78d541da417a1eb5892e5a13d1a2def6"},"cell_type":"code","source":"ths","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"a96090fbbaa3109adf1380715b0ae524afb92495"},"cell_type":"code","source":"def prob_to_bool(x, threshold=0.5):\n    #function for thesholding predictions if less than 0.5\n    x_bool = x>=threshold\n    return x_bool","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"6125957b605d0b0fda539957f830ddd297901934"},"cell_type":"code","source":"final_pred = prob_to_bool(lr_preds[:,1], \n                          threshold=0.675)\n\nprint(final_pred[:10])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c13b59744183ec44f3b5a3147fdad28381e85778"},"cell_type":"code","source":"import datetime\nnow = str(datetime.datetime.now().strftime(\"%Y-%m-%d-%H-%M\"))\nprint(now)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"8cfb588cca64c89d6c2e3c6a72acd8d196a515d6"},"cell_type":"code","source":"## submission\nsubs = pd.read_csv(\"../input/sample_submission.csv\")\nprint(\"Model - logreg 3fold - \",now)\nsubs.prediction = final_pred\nsubs.to_csv('submission.csv', index=False)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"00784066cc195409a4bdb004b8ef5e69c3d33fb5"},"cell_type":"code","source":"print(\"Estimated True : {:.1f} \".format(subs[subs == True].sum()['prediction']))\n# print(\"Estimated False: {:.1f} \".format(subs[subs != True].sum()['prediction']))\nprint(\"Original Ratio ~6.2% , Test Ratio = {:.3f}%\".format(subs[subs == True].sum()['prediction']/len(subs)*100))\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b1109a525222c9744e3a0d43131ea9e1b5a6b938"},"cell_type":"markdown","source":"The final model output had a score of about 0.62."},{"metadata":{"trusted":true,"_uuid":"f6f5d2b36c6c46893a1052550f8f74a71ac0fcab"},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}