{"cells":[{"metadata":{"_uuid":"8dc4465d05db7f8443d66385d36b5e3491948371"},"cell_type":"markdown","source":"# インポート"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"_kg_hide-input":false},"cell_type":"code","source":"import os\nimport json\nimport string\nimport numpy as np\nimport pandas as pd\nfrom pandas.io.json import json_normalize\nimport matplotlib.pyplot as plt\nimport seaborn as sns\ncolor = sns.color_palette()\nfrom tqdm import tqdm\ntqdm.pandas()\n\n%matplotlib inline\n\nfrom plotly import tools\nimport plotly.offline as py\npy.init_notebook_mode(connected=True)\nimport plotly.graph_objs as go\n\nfrom sklearn import model_selection, preprocessing, metrics, ensemble, naive_bayes, linear_model\nfrom sklearn.feature_extraction.text import TfidfVectorizer, CountVectorizer\nfrom sklearn.decomposition import TruncatedSVD\nimport lightgbm as lgb\n\npd.options.mode.chained_assignment = None\npd.options.display.max_columns = 999","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c0d647a81bb5b1a13332a17fffad6c6b0e5cf787"},"cell_type":"markdown","source":"# データファイルの確認"},{"metadata":{"trusted":true,"_uuid":"6fc5aa15387e3585ac51fbc5db66565a058f1a2d","_kg_hide-input":false},"cell_type":"code","source":"!ls ../input/","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ca4069cd714e42e1c9b6b3bb4f95c95aa0fc7838"},"cell_type":"markdown","source":"## 3つのCSVと1つのディレクトリが存在している\n\n* train.csv - the training set\n* test.csv - the test set\n* sample_submission.csv - A sample submission in the correct format\n* enbeddings/ - Folder containing word embeddings."},{"metadata":{"trusted":true,"_uuid":"5adcd5cd216cf95623326abf3235b7670a2af729"},"cell_type":"code","source":"!ls ../input/embeddings/","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"e8e3bc86f52a01c4125201a736ce77343a48c11e"},"cell_type":"markdown","source":"## embeddingsディレクトリの配下には、下記ファイルが存在している\nこのコンペには外部データソースは許可されていないが、モデルで使用できるデータセットとともに、下記にあるようにいくつかの単語データが提供されている。\n\n* GoogleNews-vectors-negative300 - https://code.google.com/archive/p/word2vec/\n* glove.840B.300d - https://nlp.stanford.edu/projects/glove/\n* paragram_300_sl999 - https://cogcomp.org/page/resource_view/106\n* wiki-news-300d-1M - https://fasttext.cc/docs/en/english-vectors.html"},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"# データの読み込みとデータ形状確認\ntrain_df = pd.read_csv(\"../input/train.csv\")\ntest_df = pd.read_csv(\"../input/test.csv\")\nprint(\"Train shape : \", train_df.shape)\nprint(\"Test shape : \", test_df.shape)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"scrolled":true,"_uuid":"6473283634cdd157b12575c0b828cdff9721eadc"},"cell_type":"code","source":"# trainデータの確認\ntrain_df.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"9d9bc1ce61cbe47aecaeb02a31b11cd78068f7ac"},"cell_type":"code","source":"# テストデータの確認\ntest_df.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a5fb3e1dd7990f063b5cbb199c51a6f6fc3a9d91"},"cell_type":"markdown","source":"# 目的変数の確認\n\nまず、目的変数（今回は0 or 1）がどのぐらいの割合で存在しているか確認してみる"},{"metadata":{"trusted":true,"_uuid":"35979d0d7088cc525b961bfacbbd22b7b4b6d8bb"},"cell_type":"code","source":"# target count\ncnt_srs = train_df['target'].value_counts()\ncnt_srs.head()","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":false,"trusted":true,"_uuid":"dea34a7f45dc6656c34472f0b7a941070bc9bee6"},"cell_type":"code","source":"# 棒グラフとして出力するデータの設定\ntrace = go.Bar(\n    x=cnt_srs.index,\n    y=cnt_srs.values,\n    marker=dict(\n        color=cnt_srs.values,\n        colorscale = 'Picnic',\n        reversescale = True\n    ),\n)\n\n# レイアウト\nlayout = go.Layout(\n    title='Target Count',\n    font=dict(size=18)\n)\n\ndata = [trace]\nfig = go.Figure(data=data, layout=layout)\npy.iplot(fig, filename=\"TargetCount\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"21793142a0271ce513ea89584f2177d098f1245e"},"cell_type":"markdown","source":"グラフからも分かるように、データに偏りがある。\n\n- 0\n→問題のない質問内容\n\n- 1\n→問題のある質問内容（insincere）"},{"metadata":{"trusted":true,"_uuid":"6ddbc99054e1143d61c215b393a321a5969f9fba"},"cell_type":"code","source":"## 円グラフも表示してみる\nlabels = (np.array(cnt_srs.index))\nsizes = (np.array((cnt_srs / cnt_srs.sum())*100))\n\ntrace = go.Pie(labels=labels, values=sizes)\nlayout = go.Layout(\n    title='Target distribution',\n    font=dict(size=18),\n    width=600,\n    height=600,\n)\ndata = [trace]\nfig = go.Figure(data=data, layout=layout)\npy.iplot(fig, filename=\"usertype\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"595a4e32ed75442c2500b9b12b178e40b0a7e155"},"cell_type":"markdown","source":"問題のある質問データはわずか6%しか存在しないことが分かる。\n\n残りの約94%は、問題のない質問データとなっている。"},{"metadata":{"_uuid":"85c22a79a9bc726e25012d31870f59e1f2ee694e"},"cell_type":"markdown","source":"# wordCroud\n\n質問内容をWordCloudを使用して可視化してみる。"},{"metadata":{"_kg_hide-input":false,"trusted":true,"_uuid":"0b0fcfc1ab5b12bb9efcff2882a799bdee8d3959"},"cell_type":"code","source":"# import\nfrom wordcloud import WordCloud, STOPWORDS\n\n# 描画関数\ndef plot_wordcloud(text, mask=None, max_words=200, max_font_size=100, figure_size=(24.0,16.0), \n                   title = None, title_size=40):\n    \n    # 除外ワードの設定\n    stopwords = set(STOPWORDS)\n    \n    # 除外ワードを追加したい場合はここに記述\n    #more_stopwords = {''}\n    #stopwords = stopwords.union(more_stopwords)\n    \n    # wordcloudの生成\n    wordcloud = WordCloud(background_color='black',\n                    stopwords = stopwords,\n                    max_words = max_words,\n                    max_font_size = max_font_size, \n                    random_state = 42,\n                    width=800, \n                    height=400,\n                    mask = mask)\n    wordcloud.generate(str(text))\n    \n    plt.figure(figsize=figure_size)\n    plt.imshow(wordcloud)\n    plt.title(title, fontdict={'size': title_size, \n                               'color': 'black', \n                               'verticalalignment': 'bottom'})\n    plt.axis('off')\n    plt.tight_layout()  ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"cb46e2f3fe84cfe9e9996421ad4b6eaf1873733c"},"cell_type":"code","source":"# 描画\nplot_wordcloud(train_df[\"question_text\"], title=\"Word Cloud of Questions\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"1fe54117c937a36f48c76fcb53f424d60f1db1a4"},"cell_type":"markdown","source":"パッと見ただけでも、様々な単語が含まれていることがわかる。\nknow, good, become, think, Amazon, Quora...\n\nここから、**問題のある質問** と **問題のない質問** に関して、頻出している単語を見ていく"},{"metadata":{"_uuid":"cd40abcc266f886ac338af9ef59b9dbc9d184655"},"cell_type":"markdown","source":"# 頻出単語の可視化"},{"metadata":{"trusted":true,"_uuid":"d0b46d8f4c2ec4171708010705d4168fbe48183d"},"cell_type":"code","source":"from collections import defaultdict\ntrain1_df = train_df[train_df[\"target\"]==1] #問題あり（insincere）\ntrain0_df = train_df[train_df[\"target\"]==0] #問題なし（sincere）","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"6d072e5971307ae6b9226c7d430aa200c24e48f1"},"cell_type":"code","source":"train1_df.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"6c4d5431880debff49f4ec1b18feb7a0bcb10109"},"cell_type":"code","source":"train0_df.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"e9f4b1fb7c9ab0df9fba0e6018adbe78cd86bb20"},"cell_type":"code","source":"# テキスト分割関数\ndef generate_ngrams(text, n_gram=1):\n    token = [token for token in text.lower().split(\" \") if token != \"\" if token not in STOPWORDS] # ストップワード有り\n    #token = [token for token in text.lower().split(\" \") if token != \"\"] # ストップワード無し\n    ngrams = zip(*[token[i:] for i in range(n_gram)])\n    return [\" \".join(ngram) for ngram in ngrams]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ca8f14e509fc91eceec252da88d906199b6a062d"},"cell_type":"code","source":"# 例：分割させる文字列\ntrain0_df['question_text'][1]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"3903a3382d11114e53a202717ea58d0f4cff2942"},"cell_type":"code","source":"# 試しに分割してみる\ngenerate_ngrams(train0_df['question_text'][1])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"fae78ec839e76f707da1be8c14c36152fa667e3b"},"cell_type":"markdown","source":"ストップワードに含まれている文字（'do', 'you', 'an', 'to' など）は除外された状態でリスト化される"},{"metadata":{"trusted":true,"_uuid":"342a9085c202215dce0ca47dc09957668bdf2816"},"cell_type":"code","source":"# 棒グラフ生成関数\ndef horizontal_bar_chart(df, color):\n    trace = go.Bar(\n        y=df[\"word\"].values[::-1],\n        x=df[\"wordcount\"].values[::-1],\n        showlegend=False,\n        orientation = 'h',\n        marker=dict(\n            color=color,\n        ),\n    )\n    return trace","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"faad0e24c13083fdd6651e26b41b8dff37f67dea"},"cell_type":"markdown","source":"## 棒グラフ表示\n\n 問題のない質問"},{"metadata":{"trusted":true,"_uuid":"b1d70d7b83e8eed88a9efa2690e5c880fc61a736"},"cell_type":"code","source":"# 初期化\nfreq_dict = defaultdict(int)\n\n# 単語毎にカウント\nfor sent in train0_df[\"question_text\"]:\n    for word in generate_ngrams(sent):\n        freq_dict[word] += 1","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"3c1858e4a619df46010bb3c462f574772f3fdb9b"},"cell_type":"code","source":"# 1番目のカラムでソートしたデータフレームを生成\nfd_sorted = pd.DataFrame(sorted(freq_dict.items(), key=lambda x: x[1])[::-1])\nfd_sorted.head(10)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"3c49991059d91cacb55e7e51661837362ca60f60"},"cell_type":"code","source":"# カラム名を設定 \nfd_sorted.columns = [\"word\", \"wordcount\"]\n\n# グラフの設定\ntrace0 = horizontal_bar_chart(fd_sorted.head(50), 'blue')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"9d7372cc8edab9a4ffe1ba91047bb11198623e49"},"cell_type":"code","source":"# 同じように、問題のある質問についても設定する\nfreq_dict = defaultdict(int)\nfor sent in train1_df[\"question_text\"]:\n    for word in generate_ngrams(sent):\n        freq_dict[word] += 1\nfd_sorted = pd.DataFrame(sorted(freq_dict.items(), key=lambda x: x[1])[::-1])\nfd_sorted.columns = [\"word\", \"wordcount\"]\ntrace1 = horizontal_bar_chart(fd_sorted.head(50), 'blue')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"e75ff74421cc6a447ddc14c2f86ba6298ac0ab5f"},"cell_type":"code","source":"# プロット\nfig = tools.make_subplots(rows=1, cols=2, vertical_spacing=0.04,\n                          subplot_titles=[\"Frequent words of sincere questions\", \n                                          \"Frequent words of insincere questions\"])\nfig.append_trace(trace0, 1, 1)\nfig.append_trace(trace1, 1, 2)\nfig['layout'].update(height=1200, width=900, paper_bgcolor='rgb(233,233,233)', title=\"Word Count Plots\")\npy.iplot(fig, filename='word-plots')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"0e59af0195cf991bf30af4dfbc1d4ec1a6c760ff"},"cell_type":"markdown","source":"上記のグラフを見ると、いくつか分かることがあると思います。例えば、\n\n- `will`や`think`や`many`や`people`という一般的な単語は、sincere / insincere に関係なく、どちらのクラスにも多く出現している。\n- 上記で挙げた単語を除くと、sincereクラスでは`good`や`best`といった、比較的ポジティブな単語が目立つ。\n- 逆に、insincereクラスでは、`women`, `trump`, `white`といった単語が目立つ。\n\n`trump`ってあのトランプ大統領だろうか・・・？  \n`white`もなぜ上位に来ているか気になる。\n\n次に、コンテキストの数を2つにして分割した結果を見てみる。"},{"metadata":{"trusted":true,"_uuid":"5cd8097b14319401437d16c2f149beff8aab6203"},"cell_type":"code","source":"# sincereクラス\nfreq_dict = defaultdict(int)\nfor sent in train0_df[\"question_text\"]:\n    for word in generate_ngrams(sent,2): #引数を2にすることで、前後2つのワードで分割される\n        freq_dict[word] += 1\nfd_sorted = pd.DataFrame(sorted(freq_dict.items(), key=lambda x: x[1])[::-1])\nfd_sorted.columns = [\"word\", \"wordcount\"]\ntrace0 = horizontal_bar_chart(fd_sorted.head(50), 'orange')\n\n# insincereクラス\nfreq_dict = defaultdict(int)\nfor sent in train1_df[\"question_text\"]:\n    for word in generate_ngrams(sent,2): #引数を2にすることで、前後2つのワードで分割される\n        freq_dict[word] += 1\nfd_sorted = pd.DataFrame(sorted(freq_dict.items(), key=lambda x: x[1])[::-1])\nfd_sorted.columns = [\"word\", \"wordcount\"]\ntrace1 = horizontal_bar_chart(fd_sorted.head(50), 'orange')\n\n# プロット\nfig = tools.make_subplots(rows=1, cols=2, vertical_spacing=0.04,horizontal_spacing=0.15,\n                          subplot_titles=[\"Frequent bigrams of sincere questions\", \n                                          \"Frequent bigrams of insincere questions\"])\nfig.append_trace(trace0, 1, 1)\nfig.append_trace(trace1, 1, 2)\nfig['layout'].update(height=1200, width=900, paper_bgcolor='rgb(233,233,233)', title=\"Bigram Count Plots\")\npy.iplot(fig, filename='word-plots')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a064d49584268688d20c6c0e5c5da2b0fa5b5066"},"cell_type":"markdown","source":"`trump`はやはりトランプ大統領だった。  \n`donald trump`が insincereクラスの中ではぶっちぎりのTOPワードということが分かったが、\nsincereクラスも見てみると、こちらにも`donald trump`が 登場しているので一概に問題のある言葉と言えるわけでもなさそうだ。\n\n`white`は白人のことだった。併せて`black people（黒人）`もTOPワードとして出現している。\n\nまた、勘違いしてしまいがちだが、各図の横軸の数値のスケールは大きく異なっている。  \n（`donald trump`の例で行くと、sincareクラスには1417個存在するが、insincareクラスには1076個しか存在しない）\n\n"},{"metadata":{"trusted":true,"_uuid":"45edfc9f5e8b364c9997dc76ad760c01cb27ad4b"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d2bc5e8c505bec12be247ebd3acb0056bee561af"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c556b34decc05717ecea216186c677291b595a08"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ec9f1affad65638693b6f46a83091d110e37f5be"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"68ab279dd44dce4c53750e24471c8a9680208f0c"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"3e641e43e9b0de00a75abd83483df3c4ec04d1f7"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"e5912d51cf0c78fc1cc55e93247562e8d04622aa"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"9d8e77fcb7c184cb713776b0f132ab4df2520ddc"},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}