{"cells":[{"metadata":{"_uuid":"a8edbc21117fdb76c44ad1ac62064fe2827155cb"},"cell_type":"markdown","source":"# Data exploration\nThis kernel gives an initial understanding of data distribution **Quora Insincere Questions Classification**."},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nfrom datetime import datetime\nimport matplotlib.pyplot as plt\n%matplotlib inline\nfrom plotly.offline import download_plotlyjs, init_notebook_mode, plot, iplot\ninit_notebook_mode(connected=True)\nimport plotly.graph_objs as go","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","collapsed":true,"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":false},"cell_type":"markdown","source":"## Import the data\nLet us look at the training data first. We have around 1.3 million questions and the corresponding target labels."},{"metadata":{"trusted":true,"_uuid":"9de7bdfa2e47bf3c259db8b8f5e3742d9f99e075"},"cell_type":"code","source":"# First let us load the datasets into different Dataframes\ntrain_df = pd.read_csv('../input/train.csv')\n\n# Dimensions\nprint('Train shape:', train_df.shape)\n# Set of features we have are: date, store, and item\ndisplay(train_df.sample(10))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"1bcb91e7aa91985172325cd963c86fc0fdc12174"},"cell_type":"markdown","source":"## Sincere vs Insincere questions\nClearly there are only **6.1% insincere** questions and most of them are sincere ones."},{"metadata":{"trusted":true,"_uuid":"b96d7c64272300191eefdcdb7eeeb72da8fa7814"},"cell_type":"code","source":"trgt_cnts = train_df.groupby(['target']).target.count()\ntrgt_cnts = trgt_cnts.to_frame()\ntrgt_cnts.columns=['counts']\ntrgt_cnts.head(20)\nx_labels = ['sincere' if i==0 else 'insencere' for i in trgt_cnts.index]\ndata = [go.Bar(x=x_labels, y=trgt_cnts.counts)]\nlayout = go.Layout(title='Sincere vs Insincere Question distribution',\n    xaxis=dict(title='Sincere and Insincere questions'),\n    yaxis=dict(title='Question Counts'))\niplot(go.Figure(data=data, layout=layout))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"42a2bc38ca26d1ea1b7b6190b7cef8847d8a3da9"},"cell_type":"markdown","source":"## Character and word length distributions\nLet us look how does the character length varies in these questions."},{"metadata":{"trusted":true,"_uuid":"e89b122339e58abc370d8bab9fa485ff01f2fd77"},"cell_type":"code","source":"tmp_df = train_df.copy(deep=True)\ntmp_df['q_length'] = tmp_df.question_text.apply(lambda qt: len(qt))\ntmp_df.sort_values(['q_length'], inplace=True)\nprint('Min question length', tmp_df.q_length.min())\nprint('Max question length', tmp_df.q_length.max())\nprint('Avg question length', tmp_df.q_length.mean())\ncounts = tmp_df.groupby(['q_length'])['q_length'].count()\ncln = counts.to_frame()\ncln.columns=['counts']\n#cln.head()\ndata = [go.Bar(x=cln.index, y=cln.counts)]\nlayout = go.Layout(title='Question length distribution',\n    xaxis=dict(title='Question length'),\n    yaxis=dict(title='Question Counts'))\niplot(go.Figure(data=data, layout=layout))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"77c8a03c39fc6cafb2996ec407304fae90ddf0ac"},"cell_type":"markdown","source":"And now let us look at the word counts in these questions"},{"metadata":{"trusted":true,"_uuid":"63c6088beb2856b03b652b98673b491f2ce32db8"},"cell_type":"code","source":"tmp_df['w_length'] = tmp_df.question_text.apply(lambda qt: len(qt.split()))\ntmp_df.sort_values(['w_length'], inplace=True)\nprint('Min words', tmp_df.w_length.min())\nprint('Max words', tmp_df.w_length.max())\nprint('Avg words', tmp_df.w_length.mean())\nword_counts = tmp_df.groupby(['w_length'])['w_length'].count()\ncln = word_counts.to_frame()\ncln.columns=['counts']\n#cln.head()\ndata = [go.Bar(x=cln.index, y=cln.counts)]\nlayout = go.Layout(title='Number of words distribution',\n    xaxis=dict(title='Number of words'),\n    yaxis=dict(title='Question Counts'))\niplot(go.Figure(data=data, layout=layout))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5292ae2b341d2504fcf614420499e8410b945c28"},"cell_type":"markdown","source":"This is for initial analysis. I plan to keep adding new insights regarding this dataset.\nHope this is helpful!"}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}