{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"#toxic question classification final exam","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-01-08T08:06:03.729506Z","iopub.execute_input":"2022-01-08T08:06:03.730020Z","iopub.status.idle":"2022-01-08T08:06:03.754566Z","shell.execute_reply.started":"2022-01-08T08:06:03.729916Z","shell.execute_reply":"2022-01-08T08:06:03.753137Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Mô tả bài toán:**\nTrong thời đại phát triển của công nghệ số, công nghệ cùng với mạng internet đã kết nối con người từ mọi lục địa trên khắp trái đất lại với nhau chỉ trong một tích tăc. Các diễn đàn cùng với đó là các nền tảng mạng xã hội ngày một trở nên quan trọng hơn. Tuy nhiên chúng như những nơi công cộng khác, không gì có thể đảm bảo sự văn minh của những người tham gia được cả. Chính vì thế, để góp phần đảm bảo một môi trường văn minh và đẹp đẽ, Quora đã áp dụng phương pháp dùng học máy để hỗ trợ nhận diện những câu hỏi toxic trong các topic!\n\nNhững câu hỏi được đặt ra trên diễn đàn này có đủ các thể loại. Đa phần các câu hỏi được yêu cầu để xử lý trong đề tài đều là những câu hỏi tiếng anh và hoặc là chúng toxic hoặc chỉ là những câu hỏi bình thường. Việc của mô hình bài toán đó là phân loại xem câu hỏi nào là câu hỏi toxic và câu hỏi nào không! \n- Input: Một câu hỏi\n- Output: Đánh giá toxic hoặc không toxic","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\n\n# natural language tool kit\nimport nltk\nfrom nltk.corpus import stopwords\nimport string\n\n#sklearn things\nfrom sklearn.model_selection import train_test_split as df_split\nfrom sklearn.naive_bayes import MultinomialNB","metadata":{"execution":{"iopub.status.busy":"2022-01-08T08:06:03.781247Z","iopub.execute_input":"2022-01-08T08:06:03.782393Z","iopub.status.idle":"2022-01-08T08:06:05.465930Z","shell.execute_reply.started":"2022-01-08T08:06:03.782344Z","shell.execute_reply":"2022-01-08T08:06:05.464905Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Miêu tả mô hình:**\nmô hình được sử dụng là Multinomial Naive Bayes\nThuật toán Multinomial Naive Bayes là phương pháp học máy dựa trên xác suất thường sử dụng trong xử lý ngôn ngữ - NLP (Natural Language Processing). Thuật toán dựa trên nguyên lý Bayes và dự đoán label của một tập hợp từ. Nó tính xác suất của mỗi label từ những feature và sau đó đưa ra cái label có tỉ lệ cao nhất.  \n\nNguyên lý Bayes, được phát triển bởi Thomas Bayes, được sử dụng để tính toán tỉ lệ xảy ra của một sự kiện dựa trên những điều kiện của những sự kiện đã biết xảy ra trước đó. Để biểu diễn nguyên lý dưới dạng toán học ta có công thức:\n\n**P(A|B) = P(A) * P(B|A)/P(B)**\nTrong đó: \n- P(A), P(B): là xác suất của A và B\n- P(B|A): là xác suất của A khi đã biết tỉ lệ xác suất của B\n\nMô hình này có những ưu nhược điểm sau:\n- Ưu điểm:\n1. Có thể dễ dàng sử dụng nếu như chỉ cần tính xác suất\n2. Có thể sử dụng trên cả những dữ liệu liên tục hay rời rạc.\n3. Thuật toán đơn giản và có thể được dùng để dự đoán trong những ứng dụng thơì gian thật.\n4. Linh hoạt và dễ dàng làm việc với những ứng dụng thời gian thực.\n\n- Nhược điểm:\n1. Độ chính xác của thuật toán thấp hơn những thuật toán khác.\n2. Thuật toán không phù hợp với những bài toán hồi quy.Thuật toán Naive Bayes chỉ sử dụng cho việc phân loại dữ liệu dạng xâu ký tự và không thể dự đoán các dữ liệu số","metadata":{}},{"cell_type":"code","source":"#get data\npath_train = '../input/quora-insincere-questions-classification/train.csv'\npath_test = '../input/quora-insincere-questions-classification/test.csv'\n\nraw_df_train = pd.read_csv(path_train)\nraw_df_train.head(10)","metadata":{"execution":{"iopub.status.busy":"2022-01-08T08:06:05.467474Z","iopub.execute_input":"2022-01-08T08:06:05.467713Z","iopub.status.idle":"2022-01-08T08:06:10.920087Z","shell.execute_reply.started":"2022-01-08T08:06:05.467684Z","shell.execute_reply":"2022-01-08T08:06:10.919106Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Phân tích dữ liệu**\nTệp dữ liệu của chúng ta bao gồm:\n1. qid(Question ID): Dạng chuỗi ký tự đánh dấu phân biệt các câu hỏi với nhau và độc nhất. \n2. Question_text: là nội dung của các câu hỏi của Quora sẽ được sử dụng để phân tích và huấn luyện mô hình. Đây là sẽ trở thành các feature của mô hình. \n3. Target: Phân biệt giữa những câu hỏi toxic và các câu hỏi bình thường. Đây là label của mô hình","metadata":{}},{"cell_type":"code","source":"raw_df_train.groupby('target').describe()","metadata":{"execution":{"iopub.status.busy":"2022-01-08T08:06:10.921408Z","iopub.execute_input":"2022-01-08T08:06:10.921666Z","iopub.status.idle":"2022-01-08T08:06:14.549448Z","shell.execute_reply.started":"2022-01-08T08:06:10.921634Z","shell.execute_reply":"2022-01-08T08:06:14.548218Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Ta có thể nhận thấy: \n- Số lượng câu hỏi toxic: 80810\n- Số lượng câu hỏi bình thường: 1225312\nSố lượng câu hỏi toxic đang bé hơn lượng câu hỏi bình thường. Giả sử trong trường hợp nếu ta sử dụng toàn bộ dataset để huấn luyện mô hình sẽ gây ra tình trạng overfit với các câu hỏi bình thường. Như thế ta sẽ sampling tập câu hỏi không toxic sau đây để đảm bảo điều này không xảy ra.","metadata":{}},{"cell_type":"code","source":"#download stop word\nnltk.download('stopwords')","metadata":{"execution":{"iopub.status.busy":"2022-01-08T08:06:14.552072Z","iopub.execute_input":"2022-01-08T08:06:14.552471Z","iopub.status.idle":"2022-01-08T08:06:34.580421Z","shell.execute_reply.started":"2022-01-08T08:06:14.552424Z","shell.execute_reply":"2022-01-08T08:06:34.579405Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Các stop word là các từ không mang quá nhiều ý nghĩa cho nội dung chính của câu và có thể gây confuse cho mô hình, vì thế ta sẽ lược bỏ\nví dụ của stopwords: a, an , the, ...","metadata":{}},{"cell_type":"code","source":"#declare a function for pre-processing sentence before feed into the model\ndef preprocess_text(text):\n    #remove punctuation\n    nopunc = [char for char in text if char not in string.punctuation]\n    nopunc = ''.join(nopunc)\n    #remove stopwords\n    clean_words = [word for word in nopunc.split() if word.lower() not in stopwords.words('english')]\n    #return a list of clean text words\n    return clean_words","metadata":{"execution":{"iopub.status.busy":"2022-01-08T08:06:34.582197Z","iopub.execute_input":"2022-01-08T08:06:34.582578Z","iopub.status.idle":"2022-01-08T08:06:34.590579Z","shell.execute_reply.started":"2022-01-08T08:06:34.582540Z","shell.execute_reply":"2022-01-08T08:06:34.589606Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Tiền xử lý dữ liệu:**\nHàm tiền xử lý dữ liệu để đưa một câu về thành dạng các feature:\n1. Một câu được đưa vào hàm sẽ bị tách rời rạc các từ trong câu ra.\n2. Tập các từ rời rạc sẽ bị lược bỏ đi các từ stop word.\n3. Sau đó trả ra một list các từ mang ý nghĩa chính của câu, chính là các feature.","metadata":{}},{"cell_type":"code","source":"from sklearn.feature_extraction.text import CountVectorizer\ntemp_text = 'I do not know but this is so damn bad, why it is so hard. wish i could quit cause i am so tired'\nprint('check 1: '+str(preprocess_text(temp_text)))\ntemp = CountVectorizer(analyzer = preprocess_text).fit_transform([temp_text])\nprint('check 2: '+str(temp))","metadata":{"execution":{"iopub.status.busy":"2022-01-08T08:06:34.591823Z","iopub.execute_input":"2022-01-08T08:06:34.592145Z","iopub.status.idle":"2022-01-08T08:06:34.624097Z","shell.execute_reply.started":"2022-01-08T08:06:34.592104Z","shell.execute_reply":"2022-01-08T08:06:34.623399Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Ví dụ: Câu \"I do not know but this is so damn bad, why it is so hard. wish i could quit cause i am so tired\"\nĐược đưa vào hàm tiền xử lý sẽ được tách rời và lược bỏ đi các từ stopword. \nKết quả trả về là một chuỗi như ở check_1\n\nSau đó ta sẽ sử dụng phương pháp One Hot Coding với các từ trên. ở đây chỉ là thử nghiệm nên có thể thấy các hệ số biểu diễn đều được set lên 1.\nTuy nhiên trong thực tế sau khi huấn luyện thì các feature được xắp xếp trở thành một chuỗi từ điển mà với mỗi một câu đưa vào thì thực chất ta đưa vào mô hình một \"chuỗi từ điển\" mà các feature trong câu được set lên 1 tại vị trí tương ứng.","metadata":{}},{"cell_type":"code","source":"#separate question by it's label\ngroup_non_toxic = raw_df_train[raw_df_train['target']==0]\ngroup_toxic = raw_df_train[raw_df_train['target']==1]\n\n#down sampling non-toxic question\ndownsampled_non_toxic = group_non_toxic.sample(group_toxic.shape[0]*10)\nprint('raw group toxic size: '+str(group_toxic.shape[0]))\nprint('raw downsampled_non_toxic size: '+str(downsampled_non_toxic.shape[0]))","metadata":{"execution":{"iopub.status.busy":"2022-01-08T08:06:34.625706Z","iopub.execute_input":"2022-01-08T08:06:34.626139Z","iopub.status.idle":"2022-01-08T08:06:35.071404Z","shell.execute_reply.started":"2022-01-08T08:06:34.626107Z","shell.execute_reply":"2022-01-08T08:06:35.070325Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#mix toxic and non-toxic quest to a new dataset\ntrain_df = pd.concat([group_toxic, downsampled_non_toxic])\nprint(train_df.shape)","metadata":{"execution":{"iopub.status.busy":"2022-01-08T08:06:35.072804Z","iopub.execute_input":"2022-01-08T08:06:35.073034Z","iopub.status.idle":"2022-01-08T08:06:35.164289Z","shell.execute_reply.started":"2022-01-08T08:06:35.073007Z","shell.execute_reply":"2022-01-08T08:06:35.163372Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Sau khi downsampling lượng câu hỏi non-toxic, ta gộp 2 tập câu hỏi lại thành 1 tập lớn và trộn đều thành một tập dataset lớn và sử dụng trong quá trình training model","metadata":{}},{"cell_type":"code","source":"CV = CountVectorizer(analyzer = preprocess_text)","metadata":{"execution":{"iopub.status.busy":"2022-01-08T08:06:35.165464Z","iopub.execute_input":"2022-01-08T08:06:35.165700Z","iopub.status.idle":"2022-01-08T08:06:35.170273Z","shell.execute_reply.started":"2022-01-08T08:06:35.165663Z","shell.execute_reply":"2022-01-08T08:06:35.169283Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Khởi tạo một bộ mã hóa One-Hot-Coding lấy từ thư viện của Sklearn","metadata":{}},{"cell_type":"code","source":"message_ = CV.fit_transform(train_df['question_text'])","metadata":{"execution":{"iopub.status.busy":"2022-01-08T08:06:35.172916Z","iopub.execute_input":"2022-01-08T08:06:35.173182Z","iopub.status.idle":"2022-01-08T08:33:26.910108Z","shell.execute_reply.started":"2022-01-08T08:06:35.173145Z","shell.execute_reply":"2022-01-08T08:33:26.908795Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Cho bộ mã hóa xử ký tập các từ có trong bộ dataset sử dụng để huấn luyện.","metadata":{}},{"cell_type":"code","source":"x_train, x_test, y_train, y_test  = df_split(message_, train_df['target'],\n                                             test_size=0.05, random_state=4, stratify=train_df['target'])","metadata":{"execution":{"iopub.status.busy":"2022-01-08T08:33:26.911786Z","iopub.execute_input":"2022-01-08T08:33:26.912058Z","iopub.status.idle":"2022-01-08T08:33:27.724545Z","shell.execute_reply.started":"2022-01-08T08:33:26.912029Z","shell.execute_reply":"2022-01-08T08:33:27.723649Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Chia tập dataset ra cho hai mục đích riêng rẽ là huấn luyện và kiểm thử. Tỉ lệ lượng dữ liệu test trên toàn tập dữ liệu là 0.05\nCó thêm điều kiện về tỉ lệ cân bằng câu hỏi toxic và non-toxic để có thể cho ra một kết quả trực quan nhất.","metadata":{}},{"cell_type":"code","source":"#model classifier\nfrom sklearn.naive_bayes import MultinomialNB\nclassifier = MultinomialNB().fit(x_train,y_train)","metadata":{"execution":{"iopub.status.busy":"2022-01-08T08:33:27.725843Z","iopub.execute_input":"2022-01-08T08:33:27.726060Z","iopub.status.idle":"2022-01-08T08:33:27.954728Z","shell.execute_reply.started":"2022-01-08T08:33:27.726032Z","shell.execute_reply":"2022-01-08T08:33:27.953612Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Khởi tạo mô hình dự đoán sử dụng thư viện sklearn.\nSau đó ta huấn luyện mô hình với tập dữ liệu train.","metadata":{}},{"cell_type":"code","source":"from sklearn.metrics import classification_report, confusion_matrix, accuracy_score\npred = classifier.predict(x_test)\nprint(classification_report(y_test, pred))\nprint('Confusion Matrix: \\n', confusion_matrix(y_test,pred))\nprint('\\nAccuracy: ', accuracy_score(y_test, pred))","metadata":{"execution":{"iopub.status.busy":"2022-01-08T08:33:27.956491Z","iopub.execute_input":"2022-01-08T08:33:27.956848Z","iopub.status.idle":"2022-01-08T08:33:28.114818Z","shell.execute_reply.started":"2022-01-08T08:33:27.956803Z","shell.execute_reply":"2022-01-08T08:33:28.113528Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Kiếm tra qua mô hình với tập dữ liệu test, ta có thể thấy độ chính xác của mô hình là 0.8619","metadata":{}},{"cell_type":"code","source":"text = 'What is the currency in Langkawi?'\nprint(text)\ntext = CV.transform([text]).toarray()\nprint(text)\n# print(text)\ntest_ = classifier.predict(text)\nprint(test_)","metadata":{"execution":{"iopub.status.busy":"2022-01-08T08:33:28.116978Z","iopub.execute_input":"2022-01-08T08:33:28.117300Z","iopub.status.idle":"2022-01-08T08:33:28.135448Z","shell.execute_reply.started":"2022-01-08T08:33:28.117265Z","shell.execute_reply":"2022-01-08T08:33:28.134593Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Kiểm thử mô hình với một câu được trích từ file test của cuộc thi.","metadata":{}},{"cell_type":"code","source":"raw_df_test = pd.read_csv(path_test)\nraw_df_test.head(10)\ntest_df = CV.transform(raw_df_test['question_text'])\npred_test = classifier.predict(test_df)\nprint(pred_test)","metadata":{"execution":{"iopub.status.busy":"2022-01-08T08:33:28.137276Z","iopub.execute_input":"2022-01-08T08:33:28.137603Z","iopub.status.idle":"2022-01-08T08:44:28.386236Z","shell.execute_reply.started":"2022-01-08T08:33:28.137560Z","shell.execute_reply":"2022-01-08T08:44:28.385483Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Chạy dự đoán với dữ liệu từ file test của cuộc thi.","metadata":{}},{"cell_type":"code","source":"sumbission = pd.read_csv('../input/quora-insincere-questions-classification/sample_submission.csv')\nsumbission['prediction'] = pred_test","metadata":{"execution":{"iopub.status.busy":"2022-01-08T08:57:17.714021Z","iopub.execute_input":"2022-01-08T08:57:17.714327Z","iopub.status.idle":"2022-01-08T08:57:18.096924Z","shell.execute_reply.started":"2022-01-08T08:57:17.714297Z","shell.execute_reply":"2022-01-08T08:57:18.095774Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sumbission.head(20)\nsumbission.to_csv(\"submission.csv\", encoding='utf-8', index=False)\n# sumbission.to_csv(r\"../input/quora-insincere-questions-classification/submission.csv\", encoding='utf-8', index=False)","metadata":{"execution":{"iopub.status.busy":"2022-01-08T08:58:48.123302Z","iopub.execute_input":"2022-01-08T08:58:48.123596Z","iopub.status.idle":"2022-01-08T08:58:49.077726Z","shell.execute_reply.started":"2022-01-08T08:58:48.123566Z","shell.execute_reply":"2022-01-08T08:58:49.076798Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Tận dụng file hướng dẫn sumbit, ta xuất ra các dự đoán của mô hình vào file \"submission.csv\"","metadata":{}}]}