{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Mô tả bài toán\n> ***Bài toán đặt ra là với các câu hỏi trên Quora có phải là câu hỏi toxic hay không (Question Classification) ?***\n* Input: Các câu hỏi dưới dạng text.\n* Output: Yes or No ?\n","metadata":{}},{"cell_type":"markdown","source":"# Công việc cần làm\n**1. Xử lý dữ liệu**\n> Dữ liệu đầu vào là text nên chúng ta cần xử lý trước khi train\n* loại bỏ dấu câu, số, stop_words, .....\n* xử lý lại nghĩa của từ\n\n**2. Training dữ liệu với mô hình Logistic Regression**\n> Đối với bài toán này, ban đầu mình sử dụng mô hình **Logistic Regression** \n>  Mô hình này giống với **Linear Regression** ở khía cạnh đầu ra là số thực, và giống với **PLA** ở việc đầu ra bị chặn (trong đoạn 0 -> 1). Mặc dù trong tên có chứa từ \"regression\", tuy vậy **Logistic Regression** thường được sử dụng nhiều cho các bài toán **classification**. Do vậy, ban đầu mình lựa chọn **Logistic Regression** cho bài toán này. Sau đó mình sẽ lựa chọn một mô hình khác để mong đạt được kết quả cao hơn.\n\n**3. Đánh giá mô hình Training LR, sử dụng mô hình khác để đạt kết quả cao hơn**\n\n**4. Xử dụng mô hình tốt nhất để giải quyết dữ liệu Test**","metadata":{}},{"cell_type":"code","source":"import pandas as pd \nimport seaborn as sns","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 1. Xử lý dữ liệu","metadata":{}},{"cell_type":"markdown","source":"**Đọc dữ liệu vào từ định dạng csv**\n\n**Dữ liệu vào gồm 3 loại**\n\n* Sample submission\n* Train\n* Test\n","metadata":{}},{"cell_type":"code","source":"df_sample_sub = pd.read_csv('/kaggle/input/quora-insincere-questions-classification/sample_submission.csv')\ndf_train = pd.read_csv('/kaggle/input/quora-insincere-questions-classification/train.csv')\ndf_test = pd.read_csv('/kaggle/input/quora-insincere-questions-classification/test.csv')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Mình xem thử từng loại dữ liệu sẽ như thế nào**","metadata":{}},{"cell_type":"code","source":"df_sample_sub.head(5)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.head(5)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Dữ liệu vào gồm các Question_text là ngôn ngữ tự nhiên. Nếu để nguyên, có vẻ vô cùng khó để train. Vì vậy chúng ta cần xử lý các Question_text.**\n> Dấu câu, số, các \"stopword\", và nhiều từ đồng nghĩa chúng ta cần xử lý lại.","metadata":{}},{"cell_type":"code","source":"df_train.info()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.target.value_counts()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Đối với tập dữ liệu Train**\n\n**Có 1225312 câu hỏi bình thường và 80810 câu hỏi Toxic\nCó thể thấy là khá là nhiều câu hỏi Toxic.**","metadata":{}},{"cell_type":"code","source":"sns.countplot(data=df_train, x='target')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Để xử lý dữ liệu vào, mình xử dụng Natural Language Toolkit**","metadata":{}},{"cell_type":"code","source":"import nltk\nimport string\nfrom nltk.tokenize import word_tokenize\nfrom nltk.corpus import stopwords\nfrom nltk.stem import WordNetLemmatizer\n\nnltk.download('stopwords')\nnltk_stopwords = stopwords.words('english')\n\nwordnet_lemmatizer = WordNetLemmatizer()\n\ndef lemSentence(sentence):\n    token_words = word_tokenize(sentence)\n    lem_sentence = []\n    for word in token_words:\n        lem_sentence.append(wordnet_lemmatizer.lemmatize(word, pos=\"v\"))\n        lem_sentence.append(\" \")\n    return \"\".join(lem_sentence)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def clean(message, lem = True):\n    # Loại bỏ dấu câu\n    message = message.translate(str.maketrans('', '', string.punctuation))\n    \n    # Loại bỏ số\n    message = message.translate(str.maketrans('', '', string.digits))\n    \n    # Loại bỏ \"stopwords\"\n    message = [word for word in word_tokenize(message) if not word.lower() in nltk_stopwords]\n    message = ' '.join(message)\n    \n    # Xử lý lại nghĩa của từ\n    if lem:\n        message = lemSentence(message)\n    \n    return message","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Clean các câu hỏi\ndf_train['question_text_cleaned'] = df_train.question_text.apply(lambda x: clean(x, True))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Có thể thấy dữ liệu sau khi xử lý đã \"sạch\" hơn so với ban đầu như bảng dưới**","metadata":{}},{"cell_type":"code","source":"df_train.head(5)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 2. Training dữ liệu với mô hình Logistic Regression\n> **Như đã trình bày ở trên, mình sử dụng LR để train**\n\n> ***Nói qua về Logistic Regression***\n>> Có lẽ các bạn đều đã biết về hai mô hình tuyến tính (linear models) **Linear Regression** và **Perceptron Learning Algorithm (PLA)**. Trong **Linear Regression**, người ta sử dụng để dự đoán output y, mô hình này phù hợp để dự đoán một giá trị thực của đầu ra không bị chặn trên và chặn dưới. Còn trong **PLA**, đầu ra chỉ nhận một trong hai giá trị 1 hoặc −1, phù hợp với các bài toán **Binary Classification**. \n\n>> Tuy vậy, mình lại không sử dụng **Linear Regression**, hay thậm chí là **PLA**, mặc dù **PLA** rất phù hợp với bài toán của mình. Mà mình lựa chọn **Logistic Regression**. Mô hình này giống với **Linear Regression** ở khía cạnh đầu ra là số thực, và giống với **PLA** ở việc đầu ra bị chặn (trong đoạn 0 -> 1). Mặc dù trong tên có chứa từ \"**regression**\", tuy vậy **Logistic Regression** thường được sử dụng nhiều cho các bài toán **classification**. Đánh giá mô hình mình thấy nó rất là linh hoạt va dễ sử dụng. Do vậy, mình lựa chọn **Logistic Regression** cho bài toán này.\n","metadata":{}},{"cell_type":"code","source":"from sklearn.linear_model import LogisticRegression\nfrom sklearn.pipeline import Pipeline\nfrom sklearn.metrics import accuracy_score, f1_score\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.feature_extraction.text import CountVectorizer\n\ncount_vectorizer = CountVectorizer()\nmodel = LogisticRegression(C=1, random_state=0)\n\nvectorize_model_pipeline = Pipeline([\n    ('count_vectorizer', count_vectorizer),\n    ('model', model)\n])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Mặc dù đã xử lý dữ liệu ở phần phía trên, tuy vậy, dữ liệu đầu vào của mình vẫn là dạng Text (nó sạch hơn một chút thôi).\nĐể sử dụng dữ liệu văn bản cho mô hình dự đoán, văn bản phải được phân tích cú pháp để loại bỏ một số từ nhất định - quá trình này được gọi là mã hóa . Sau đó, những từ này cần được mã hóa dưới dạng số nguyên hoặc giá trị dấu phẩy động, để sử dụng làm đầu vào trong thuật toán học máy. Quá trình này được gọi là trích xuất đặc trưng (hoặc vectơ hóa) .**\n\n> Scikit-learning CountVectorizerđược sử dụng để chuyển đổi một bộ sưu tập các tài liệu văn bản thành một vectơ có số lượng thuật ngữ / mã thông báo. Nó cũng cho phép xử lý trước dữ liệu văn bản trước khi tạo biểu diễn vectơ. Chức năng này làm cho nó trở thành một mô-đun biểu diễn tính năng rất linh hoạt cho văn bản.\n\n**Hay nói tóm gọn lại là mình sử dụng CountVectorizer để chuyển dữ liệu đầu vào từ Text sang Vectơ****","metadata":{}},{"cell_type":"code","source":"X_train, X_test, y_train, y_test = train_test_split(df_train['question_text_cleaned'], df_train['target'], test_size = 0.3)\nvectorize_model_pipeline.fit(X_train, y_train)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predictions = vectorize_model_pipeline.predict(X_test)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Accuracy :', accuracy_score(y_test, predictions))\nprint('F1 score :', accuracy_score(y_test, predictions))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.metrics import classification_report\n\nprint(classification_report(y_test, predictions))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Cleaning the questions\ndf_test['question_text_cleaned'] = df_test.question_text.apply(lambda x: clean(x, True))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_test['prediction'] = vectorize_model_pipeline.predict(df_test['question_text_cleaned'])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_final = df_test[['qid','prediction']]\ndf_final.set_index('qid', inplace = True)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_final.head(5)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_final.to_csv('submission.csv')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 3. Đánh giá mô hình Logistic Regression\n**Logistic Regression là mô hình dễ sử dụng, và dễ hiểu, số điểm nó đạt được cũng tương đối cao:**\n\nPrivate Score: **0.51781**\n\nPublic Score:  **0.51129**\n\n**Logistic Regression thực ra được sử dụng nhiều trong các bài toán Classification.**\n\nMặc dù có tên là **Regression**, tức một mô hình cho fitting, **Logistic Regression** lại được sử dụng nhiều trong các bài toán **Classification**. Sau khi tìm được mô hình, việc xác định class y cho một điểm dữ liệu x được xác định bằng việc so sánh hai biểu thức xác suất:\n        **P(y = 1|x; w); P(y = 0|x; w)**\nNếu biểu thức thứ nhất lớn hơn thì ta kết luận điểm dữ liệu thuộc class 1, ngược lại thì nó thuộc class 0. Vì tổng hai biểu thức này luôn bằng 1 nên một cách gọn hơn, ta chỉ cần xác định xem **P(y = 1|x; w)** lớn hơn 0.5 hay không. Nếu có, class 1. Nếu không, class 0.\n\n***Logistic Regression đạt được số điểm khá là khả quan, tuy vậy, nó vẫn không thực sự là một điểm cao, vì vậy, mình xử dụng một mô hình khác đối với bài toán này, mong rằng số điểm nó sẽ cao hơn***","metadata":{}},{"cell_type":"markdown","source":"# 4. Training dữ liệu với mô hình khác","metadata":{}}]}