{"cells":[{"metadata":{"_uuid":"1d522b60e0b64e61514ff5bc8bf8ce3118df9a81"},"cell_type":"markdown","source":"...while insincere questions tend to be politicially, racially, religiously, etc. charged. \n\nOther notebooks (like [this one](https://www.kaggle.com/sudalairajkumar/simple-exploration-notebook-qiqc)) looked at term frequencies of sincere and insincere questions and found that terms like 'will', 'many people', 'united states' appear frequently in both classes, so these terms are most likely not very predictive. \n\nThe goal of this notebook is to go one step futher and look at class-specific term frequencies. The class-specific frequency of a term is \n\n$tf_{c1} = \\frac{n_{c1}}{n}$,\n\nwhere $n$ is the number of times the term appears in the corpus, and $n_{c1}$ is the number of times the term appears in class 1 documents (insincere questions in our case). The notebook cited above looks at $n_{c0}$ and $n_{c1}$, while we normalize $n_{c1}$ with $n$. \n\nIt is useful to place the unique terms of the corpus on the $n$ - $tf_{c1}$ plane for two reasons:\n\n- it helps feature selection: predictive terms are common in the corpus (high $n$) and they predominantly appear in insincere questions (high $tf_{c1}$), so these terms are in the upper right corner of the $n$ - $tf_{c1}$ plane,\n- we can get a feeling of what the sincere and insincere questions tend to be about.\n\nWe see that sincere questions are about books, computer science, and asking for advice. The following lemmatized terms have $n > 1000$ and $tf_{c1} < 0.001$.\n```python\n['be mean by' 'be the scope' 'book for' 'computer science' 'do i prepare'\n 'ece' 'fresher' 'in computer' 'java' 'major accomplishment'\n 'mechanical engineering' 'the best book' 'the scope' 'us of'\n 'what inspire']\n```\n\nAnd the terms below have $n > 100$ and $tf_{c1} > 0.7$.\n``` python\n['all muslim' 'american so' 'asian woman' 'be american so'\n 'be black people' 'be castrate' 'be democrat' 'be feminist' 'be hindus'\n 'be indians' 'be liberal' 'be liberal so' 'be muslims' 'be white people'\n 'bhakts' 'black american' 'black man' 'black men' 'black people'\n 'black woman' 'bullshit' 'castrate' 'castration' 'democrat be'\n 'democrats' 'do black people' 'do democrat' 'do feminist' 'do liberal'\n 'do white people' 'fuck' 'hindu be' 'indian on' 'indian so' 'liberal so'\n 'liberal think' 'liberals' 'men so' 'moron' 'muslim be' 'muslims'\n 'quora moderator' 'sex with my' 'shithole' 'so stupid' 'than white'\n 'that muslim' 'the fuck' 'to fuck' 'to rape' 'white american'\n 'white girl' 'white men' 'white people' 'white woman' 'why be black'\n 'why be democrat' 'why be indians' 'why be liberal' 'why be muslim'\n 'why be white' 'why do black' 'why do democrat' 'why do feminist'\n 'why do liberal' 'why do muslim' 'why do white' 'why muslim' 'woman so']\n```\n\nLet's walk through the notebook. \n"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"# load packages\nimport numpy as np \nimport pandas as pd \nfrom sklearn.feature_extraction.text import TfidfVectorizer\nimport nltk\nfrom nltk.stem import PorterStemmer, WordNetLemmatizer\nfrom nltk import pos_tag\nimport matplotlib.pyplot as plt\nimport matplotlib\nmatplotlib.rcParams.update({'font.size': 14})\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"299cd5aabec23a7b0db568f595a6249b840fe607"},"cell_type":"code","source":"# tokenizer for stemming and lemmatization\ndef tokenize(text):\n    def convert_tag(tag):\n        part = {'ADJ' : 'a',\n                'ADV' : 'r',\n                'VERB' : 'v',\n                'NOUN' : 'n'}\n        if tag in part.keys():\n            return part[tag]\n        else:\n            # other parts of speech will be tagged as nouns\n            return 'n'\n\n    tokens = nltk.word_tokenize(text)\n    tokens_pos = pos_tag(tokens,tagset='universal')\n    stems = []\n    for item in tokens_pos:\n        term = item[0]\n        pos = item[1]\n        # stem/lemmatize tokens consisting of alphabetic characters only\n        if term.isalpha():\n            stems.append(WordNetLemmatizer().lemmatize(term, pos=convert_tag(pos)))\n            #stems.append(PorterStemmer().stem(item))\n    return stems\n","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"# vectorize corpus\ntrain_df = pd.read_csv(\"../input/train.csv\")\n\nprint(\"Train shape : \",train_df.shape)\nprint(train_df.columns)\n\n# target variable and some basic stats\ny_all = train_df['target']\nprint('fraction of datapoints in class 1: ',1e0*np.sum(y_all == 1)/len(y_all)) # fraction of datapoints in class 1\nprint('number of datapoints in class 1: ',np.sum(y_all == 1)) # number of datapoints in class 1\n\n# n-grams with n = 1 - 3, no stopwords, use words that appear in at least min_df documents\nvectorizer = TfidfVectorizer(ngram_range=(1,3),tokenizer=tokenize,min_df=100,\\\n                             sublinear_tf=True) \nX_all = vectorizer.fit_transform(train_df['question_text'])#.sample(100,random_state=seed))\nprint(np.shape(X_all))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c9b5e309407ae2878a76b43ece88caf0994a2cdd"},"cell_type":"code","source":"# a quick look at the unique terms in the corpus\nterms = np.array(vectorizer.get_feature_names())\nprint(terms[:100]) # the first 100 terms\nprint(terms[-100:]) # the last 100 terms\nprint(np.random.choice(terms,size=100,replace=False)) # 100 randomly selected terms","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d6df9e3916796105bd913b8bd287546cdf869810"},"cell_type":"code","source":"# n -- number of times the terms appear in docs \nterm_count = X_all.getnnz(axis=0)\n\nindices = np.where(y_all == 1)[0]\n# tf_c1 -- the fraction of times the terms appear in class 1\nfrac_in_class1 = 1e0*X_all.tocsc()[indices].getnnz(axis=0)/term_count\n\nindcs = np.where((frac_in_class1 <= 0.001) & (term_count >= 1000))[0]\nprint(terms[indcs]) # terms in sincere questions\n\nindcs = np.where((frac_in_class1 >= 0.7) & (term_count >= 100))[0]\nprint(terms[indcs]) # terms in insincere questions\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b4b293dda33fc2e915feaeb2cf45a369c4b8dd49"},"cell_type":"code","source":"# plotting: there aren't so many terms in the upper right quadrant of the plot\n\nplt.scatter(term_count,frac_in_class1)\nplt.semilogx()\nplt.xlabel('n - # times term in corpus')\nplt.ylabel('tf_c1 - class-specific frequency')\nplt.title('scatter plot')\nplt.show()\n\nxbins = 10**np.linspace(2,5,31)\nybins = np.linspace(0,1,41)\ncounts, _, _ = np.histogram2d(term_count,frac_in_class1,bins=(xbins,ybins))\ncounts[counts == 0] = 0.5 # so log10(0) is not nan\nplt.pcolormesh(xbins, ybins, np.log10(counts.T))\nplt.semilogx()\nplt.xlabel('n - # times term in corpus')\nplt.ylabel('tf_c1 - class-specific frequency')\ncbar = plt.colorbar(label='count',ticks=[0,1,2,3])\ncbar.ax.set_yticklabels([1,10,100,1000])\nplt.title('heatmap')\nplt.show()\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"6eb2c6d9f3745cf10e64dda07e08cea1d0994c8b"},"cell_type":"markdown","source":"Any comments or suggestions are greatly appreaciated especially on how I can improve the preprocessing and lemmatization."}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}