{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-01-06T19:23:44.912862Z","iopub.execute_input":"2022-01-06T19:23:44.913442Z","iopub.status.idle":"2022-01-06T19:23:44.954687Z","shell.execute_reply.started":"2022-01-06T19:23:44.913318Z","shell.execute_reply":"2022-01-06T19:23:44.953758Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"markdown","source":"**--Topic: Toxic Comment Classifier--**","metadata":{}},{"cell_type":"markdown","source":"**Sử dụng các mô hình học máy để dự đoán khả năng của đoạn văn (comment_text) theo 6 trường dữ liệu :**\n* Toxic\n* Severe Toxic\n* Obscene\n* Insult\n* Threat\n* Identity Hate","metadata":{}},{"cell_type":"markdown","source":"Mục Lục\n1. Tổng quan\n2. Thực hiện\n3. Đánh giá và kết luận","metadata":{}},{"cell_type":"code","source":"# Giải nén data\n# import zipfile\n# # unzip = zipfile.ZipFile('/kaggle/input/jigsaw-toxic-comment-classification-challenge/train.csv.zip')\n# unzip = zipfile.ZipFile('/kaggle/input/jigsaw-toxic-comment-classification-challenge/train.csv.zip')\n# unzip.extractall()\n# unzip = zipfile.ZipFile('/kaggle/input/jigsaw-toxic-comment-classification-challenge/test.csv.zip')\n# unzip.extractall()\n","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:23:44.986394Z","iopub.execute_input":"2022-01-06T19:23:44.98665Z","iopub.status.idle":"2022-01-06T19:23:44.990638Z","shell.execute_reply.started":"2022-01-06T19:23:44.986623Z","shell.execute_reply":"2022-01-06T19:23:44.989567Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"1. Tổng quan\n","metadata":{}},{"cell_type":"markdown","source":"Mô hình dùng để làm gì:\n>Mô hình được tạo ra với mục đích xét khả năng toxic của dữ liệu text mà mình mong muốn, to lớn hơn là giúp cho không gian mạng sạch sẽ hơn rất nhiều","metadata":{}},{"cell_type":"code","source":"#Khởi chạy thử data cần train và nhận xét\n# import pandas as pd\ntrain_df = pd.read_csv('../input/jigsaw-toxic-comment-classification-challenge/train.csv')# Loading data\n# if train_df is None:\n#     print(1)\ntrain_df.head() #show data","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:23:45.025987Z","iopub.execute_input":"2022-01-06T19:23:45.026855Z","iopub.status.idle":"2022-01-06T19:23:47.094345Z","shell.execute_reply.started":"2022-01-06T19:23:45.026813Z","shell.execute_reply":"2022-01-06T19:23:47.093051Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Nhận xét tổng quan về dữ liệu: Dữ liệu sử dụng để training chưa được tối ưu hóa trong trường comment_text, cần có bước làm sạch dữ liệu trước khi training\n","metadata":{}},{"cell_type":"code","source":"#Khởi chạy thử data cần test\n\n# test_df = pd.read_csv('/kaggle/working/test.csv')# Loading data\ntest_df = pd.read_csv('../input/jigsaw-toxic-comment-classification-challenge/test.csv')# Loading data\ntest_df.head() #show data","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:23:47.095819Z","iopub.execute_input":"2022-01-06T19:23:47.096116Z","iopub.status.idle":"2022-01-06T19:23:48.696039Z","shell.execute_reply.started":"2022-01-06T19:23:47.096075Z","shell.execute_reply":"2022-01-06T19:23:48.695247Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import nltk\nfrom nltk.corpus import stopwords  # Xóa ký tự không cần thiết\nfrom nltk.stem.lancaster import LancasterStemmer # convert words to base form","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:23:48.697078Z","iopub.execute_input":"2022-01-06T19:23:48.697311Z","iopub.status.idle":"2022-01-06T19:23:50.388345Z","shell.execute_reply.started":"2022-01-06T19:23:48.697283Z","shell.execute_reply":"2022-01-06T19:23:50.38747Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Tải stopword trong thư viện nltk\nnltk.download('stopwords')","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:23:50.390624Z","iopub.execute_input":"2022-01-06T19:23:50.39094Z","iopub.status.idle":"2022-01-06T19:23:50.545934Z","shell.execute_reply.started":"2022-01-06T19:23:50.390892Z","shell.execute_reply":"2022-01-06T19:23:50.544878Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"set(stopwords.words('english')) # Chon English","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:23:50.547272Z","iopub.execute_input":"2022-01-06T19:23:50.547496Z","iopub.status.idle":"2022-01-06T19:23:50.560361Z","shell.execute_reply.started":"2022-01-06T19:23:50.547469Z","shell.execute_reply":"2022-01-06T19:23:50.559157Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Chúng ta sử dụng stopword để làm gì?\n* Vì trong một câu có những từ có tần số xuất hiện nhiều như the, to... các từ này thường mang ít giá trị ý nghĩa và không khác nhau nhiều trong các văn bản khác nhau.\n* Ví dụ từ \"the\" hay \"to\" thì ở văn bản nào nó cũng không bị thay đổi về ý nghĩa.\n* Vì thế chúng ta có thể xóa bỏ những từ này","metadata":{}},{"cell_type":"code","source":"# Hiển thị thử info của data\ntrain_df.info()","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:23:50.561987Z","iopub.execute_input":"2022-01-06T19:23:50.562849Z","iopub.status.idle":"2022-01-06T19:23:50.618936Z","shell.execute_reply.started":"2022-01-06T19:23:50.562801Z","shell.execute_reply":"2022-01-06T19:23:50.618129Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['comment_text'][0]","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:23:50.619888Z","iopub.execute_input":"2022-01-06T19:23:50.620588Z","iopub.status.idle":"2022-01-06T19:23:50.626132Z","shell.execute_reply.started":"2022-01-06T19:23:50.620554Z","shell.execute_reply":"2022-01-06T19:23:50.625346Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['comment_text'][1]","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:23:50.627444Z","iopub.execute_input":"2022-01-06T19:23:50.627656Z","iopub.status.idle":"2022-01-06T19:23:50.641449Z","shell.execute_reply.started":"2022-01-06T19:23:50.62763Z","shell.execute_reply":"2022-01-06T19:23:50.640679Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Như đã nói ở trên, comment_text chưa được xử lý trước khi training","metadata":{}},{"cell_type":"code","source":"#Trả về một Chuỗi chứa số lượng các giá trị\ntrain_df.toxic.value_counts(normalize=True)\n# normalize=True thì đối tượng được trả về sẽ chứa các tần số tương đối của các giá trị","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:23:50.642546Z","iopub.execute_input":"2022-01-06T19:23:50.642959Z","iopub.status.idle":"2022-01-06T19:23:50.659178Z","shell.execute_reply.started":"2022-01-06T19:23:50.642926Z","shell.execute_reply":"2022-01-06T19:23:50.658299Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Trả về một Chuỗi chứa số lượng các giá trị \ntrain_df.severe_toxic.value_counts(normalize=True)\n# normalize=True thì đối tượng được trả về sẽ chứa các tần số tương đối của các giá tr","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:23:50.662807Z","iopub.execute_input":"2022-01-06T19:23:50.66318Z","iopub.status.idle":"2022-01-06T19:23:50.672266Z","shell.execute_reply.started":"2022-01-06T19:23:50.663152Z","shell.execute_reply":"2022-01-06T19:23:50.6714Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Trả về một Chuỗi chứa số lượng các giá trị\ntrain_df.obscene.value_counts(normalize=True)\n# normalize=True thì đối tượng được trả về sẽ chứa các tần số tương đối của các giá trị","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:23:50.673498Z","iopub.execute_input":"2022-01-06T19:23:50.673716Z","iopub.status.idle":"2022-01-06T19:23:50.686926Z","shell.execute_reply.started":"2022-01-06T19:23:50.673691Z","shell.execute_reply":"2022-01-06T19:23:50.686059Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Trả về một Chuỗi chứa số lượng các giá trị\ntrain_df.insult.value_counts(normalize=True)\n# normalize=True thì đối tượng được trả về sẽ chứa các tần số tương đối của các giá trị","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:23:50.68815Z","iopub.execute_input":"2022-01-06T19:23:50.688528Z","iopub.status.idle":"2022-01-06T19:23:50.703996Z","shell.execute_reply.started":"2022-01-06T19:23:50.68849Z","shell.execute_reply":"2022-01-06T19:23:50.703166Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Trả về một Chuỗi chứa số lượng các giá trị\ntrain_df.threat.value_counts(normalize=True)\n# normalize=True thì đối tượng được trả về sẽ chứa các tần số tương đối của các giá trị","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:23:50.705515Z","iopub.execute_input":"2022-01-06T19:23:50.705775Z","iopub.status.idle":"2022-01-06T19:23:50.722372Z","shell.execute_reply.started":"2022-01-06T19:23:50.705746Z","shell.execute_reply":"2022-01-06T19:23:50.721392Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Trả về một Chuỗi chứa số lượng các giá trị\ntrain_df.identity_hate.value_counts(normalize=True)\n# normalize=True thì đối tượng được trả về sẽ chứa các tần số tương đối của các giá tr","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:23:50.723497Z","iopub.execute_input":"2022-01-06T19:23:50.723703Z","iopub.status.idle":"2022-01-06T19:23:50.737Z","shell.execute_reply.started":"2022-01-06T19:23:50.723679Z","shell.execute_reply":"2022-01-06T19:23:50.736333Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:23:50.738023Z","iopub.execute_input":"2022-01-06T19:23:50.738589Z","iopub.status.idle":"2022-01-06T19:23:50.756206Z","shell.execute_reply.started":"2022-01-06T19:23:50.738554Z","shell.execute_reply":"2022-01-06T19:23:50.755437Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#tính tổng các giá trị lấy từ cột thứ 3\ndata_count=train_df.iloc[:,2:].sum() \ndata_count","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:23:50.757428Z","iopub.execute_input":"2022-01-06T19:23:50.757796Z","iopub.status.idle":"2022-01-06T19:23:50.777952Z","shell.execute_reply.started":"2022-01-06T19:23:50.757758Z","shell.execute_reply":"2022-01-06T19:23:50.777276Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Thứ tự tăng dần là:\n* threat\n* identity_hate\n* severe_toxic\n* insult\n* obscene\n* toxic","metadata":{}},{"cell_type":"markdown","source":"Hiển thị thử các trường hợp thành một biểu đồ cột","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport nltk\nimport re\nimport string\nimport seaborn as sns\n\nfrom sklearn.feature_extraction.text import CountVectorizer, TfidfVectorizer","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:23:50.778994Z","iopub.execute_input":"2022-01-06T19:23:50.779306Z","iopub.status.idle":"2022-01-06T19:23:50.922273Z","shell.execute_reply.started":"2022-01-06T19:23:50.779281Z","shell.execute_reply":"2022-01-06T19:23:50.921289Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#sử dụng plot biểu đồ\nplt.figure(figsize=(8,4))\n\n#Sử dụng phương thức barplot trong Seaborn\n#Hiển thị ước tính điểm và khoảng tin cậy dưới dạng thanh hình chữ nhật.\nax = sns.barplot(data_count.index, data_count.values, alpha=0.8)\n\nplt.title(\"Bảng giá trị\") # đặt tên biểu đồ\nplt.ylabel(\"\", fontsize=12)\nplt.xlabel(\"Loại\", fontsize=12) # đặt tên cho trục hoành , set font là 12\n\n#Thêm text cho mỗi cột\nrects = ax.patches\nlabels = data_count.values\nfor rect, label in zip(rects, labels):\n  height = rect.get_height()\n  ax.text(rect.get_x() + rect.get_width()/2, height + 5, label, ha='center', va='bottom')\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:23:50.923437Z","iopub.execute_input":"2022-01-06T19:23:50.923669Z","iopub.status.idle":"2022-01-06T19:23:51.176261Z","shell.execute_reply.started":"2022-01-06T19:23:50.923642Z","shell.execute_reply":"2022-01-06T19:23:51.17567Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Thứ tự tăng dần là:\n* threat\n* identity_hate\n* severe_toxic\n* insult\n* obscene\n* toxic # Nhìn vào biểu đồ có thể thấy toxic là trường có nhiều nhất","metadata":{}},{"cell_type":"code","source":"#lấy độ dài của data\nnum_rows=len(train_df)\nprint(num_rows)","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:23:51.177119Z","iopub.execute_input":"2022-01-06T19:23:51.177469Z","iopub.status.idle":"2022-01-06T19:23:51.182281Z","shell.execute_reply.started":"2022-01-06T19:23:51.177437Z","shell.execute_reply":"2022-01-06T19:23:51.181671Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Xét thử phần trăm của các trường so với tổng độ dài Data","metadata":{}},{"cell_type":"code","source":"#Create bar graph\nsum_tox = train_df['toxic'].sum() / num_rows * 100\nsum_sev = train_df['severe_toxic'].sum() / num_rows * 100\nsum_obs = train_df['obscene'].sum() / num_rows * 100\nsum_thr = train_df['threat'].sum() / num_rows * 100\nsum_ins = train_df['insult'].sum() / num_rows * 100\nsum_ide = train_df['identity_hate'].sum() / num_rows * 100\n\nind = np.arange(6)\n\nax = plt.barh(ind, [sum_tox, sum_sev, sum_obs, sum_thr, sum_ins, sum_ide])\nplt.xlabel('Percentage (%)', size=20)\nplt.xticks(np.arange(0, 30, 5), size=20)\nplt.yticks(ind, ('Toxic', 'Severe Toxic', 'Obscene', 'Threat', 'Insult', 'Identity Hate'), size=15)\n\nplt.gca().invert_yaxis()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:23:51.183482Z","iopub.execute_input":"2022-01-06T19:23:51.183843Z","iopub.status.idle":"2022-01-06T19:23:51.30483Z","shell.execute_reply.started":"2022-01-06T19:23:51.183806Z","shell.execute_reply":"2022-01-06T19:23:51.304215Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Ta điều chỉnh lại 1 chút cho dễ nhìn nhận hơn","metadata":{}},{"cell_type":"code","source":"#Create bar graph\nsum_tox = train_df['toxic'].sum() / num_rows * 100\nsum_sev = train_df['severe_toxic'].sum() / num_rows * 100\nsum_obs = train_df['obscene'].sum() / num_rows * 100\nsum_thr = train_df['threat'].sum() / num_rows * 100\nsum_ins = train_df['insult'].sum() / num_rows * 100\nsum_ide = train_df['identity_hate'].sum() / num_rows * 100\n\nind = np.arange(6)\n\nax = plt.barh(ind, [sum_tox, sum_obs, sum_ins,  sum_sev, sum_ide , sum_thr])\nplt.xlabel('Percentage (%)', size=20)\nplt.xticks(np.arange(0, 30, 5), size=20)\nplt.yticks(ind, ('Toxic', 'Obscene', 'Insult', 'Severe Toxic', 'Identity Hate', 'Threat'), size=15)\n\nplt.gca().invert_yaxis()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:23:51.305964Z","iopub.execute_input":"2022-01-06T19:23:51.306389Z","iopub.status.idle":"2022-01-06T19:23:51.416189Z","shell.execute_reply.started":"2022-01-06T19:23:51.306352Z","shell.execute_reply":"2022-01-06T19:23:51.415585Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* Có tận 3 trường là dưới 5%\n* 2 trường Obscene và Insult khoảng 5%\n* Và trường Toxic vượt quá 10%","metadata":{}},{"cell_type":"markdown","source":"2. Thực thi","metadata":{}},{"cell_type":"markdown","source":"Giai đoạn tiền xử lý dữ liệu","metadata":{}},{"cell_type":"code","source":"#Tiền xử lý\n# xóa tất cả các số có chữ cái gắn liền với chúng\nalphanumeric = lambda x: re.sub('\\w*\\d\\w*', ' ', x)\n\n# thay thế dấu câu bằng khoảng trắng\n#convert tất cả chuỗi thành chữ thường\npunc_lower = lambda x: re.sub('[%s]' % re.escape(string.punctuation), ' ', x.lower())\n\n#xóa tất cả '\\n'\nremove_n = lambda x:re.sub(\"\\n\", \" \", x)\n\n#xóa các ký tự không phải ascii\nremove_non_ascii = lambda x: re.sub(r'[^\\x00-\\x7f]',r' ',x)\n\n#Apply map\ntrain_df['comment_text'] = train_df['comment_text'].map(alphanumeric).map(punc_lower).map(remove_n).map(remove_non_ascii)\n\n#Show comment_text 0\ntrain_df['comment_text'][0]","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:23:51.417362Z","iopub.execute_input":"2022-01-06T19:23:51.417727Z","iopub.status.idle":"2022-01-06T19:24:02.717928Z","shell.execute_reply.started":"2022-01-06T19:23:51.417686Z","shell.execute_reply":"2022-01-06T19:24:02.717117Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- So sánh với comment_text[0] trước đó:\n  \n  * \"Explanation\\nWhy the edits made under my username Hardcore Metallica Fan were reverted? They weren't vandalisms, just closure on some GAs after I voted at New York Dolls FAC. And please don't remove the template from the talk page since I'm retired now.89.205.38.27\"","metadata":{}},{"cell_type":"markdown","source":"- Có thể thấy, các dấu đã bị bỏ đi, các số , ký tự \"/n\" đã bị loại bỏ và chuyển về dạng viết thường hết\n","metadata":{}},{"cell_type":"code","source":"# Chia thanh 6 section\ndata_tox = train_df.loc[:,['id', 'comment_text', 'toxic']]\ndata_sev = train_df.loc[:,['id', 'comment_text', 'severe_toxic']]\ndata_obs = train_df.loc[:,['id', 'comment_text', 'obscene']]\ndata_ins = train_df.loc[:,['id', 'comment_text', 'insult']]\ndata_thr = train_df.loc[:,['id', 'comment_text', 'threat']]\ndata_ide = train_df.loc[:,['id', 'comment_text', 'identity_hate']]","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:24:02.719186Z","iopub.execute_input":"2022-01-06T19:24:02.719932Z","iopub.status.idle":"2022-01-06T19:24:02.773158Z","shell.execute_reply.started":"2022-01-06T19:24:02.719896Z","shell.execute_reply":"2022-01-06T19:24:02.77229Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pip install wordcloud","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:24:02.774375Z","iopub.execute_input":"2022-01-06T19:24:02.774608Z","iopub.status.idle":"2022-01-06T19:24:11.999778Z","shell.execute_reply.started":"2022-01-06T19:24:02.77458Z","shell.execute_reply":"2022-01-06T19:24:11.998688Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Tạo các WordCloud\nNhững đám mây chứa rất nhiều từ ở các kích thước khác nhau, chúng thể hiện tần suất hoặc tầm quan trọng của mỗi từ.","metadata":{}},{"cell_type":"code","source":"import wordcloud\nfrom PIL import Image\nfrom wordcloud import WordCloud, STOPWORDS, ImageColorGenerator\nfrom nltk.corpus import stopwords","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:24:12.001929Z","iopub.execute_input":"2022-01-06T19:24:12.002339Z","iopub.status.idle":"2022-01-06T19:24:12.215838Z","shell.execute_reply.started":"2022-01-06T19:24:12.002292Z","shell.execute_reply":"2022-01-06T19:24:12.214939Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def wordcloud(df, label):\n    subset=df[df[label]==1]\n    text=subset.comment_text.values\n    wc=WordCloud(background_color='black',max_words=4000)\n    \n    wc.generate(\" \".join(text))\n    \n    plt.figure(figsize=(20,20))\n    plt.subplot(221)\n    plt.axis('off')\n    plt.title(\"Words frequented in {}\".format(label), fontsize=20)\n    plt.imshow(wc.recolor(colormap='gist_earth', random_state=244), alpha=0.98)","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:24:12.21718Z","iopub.execute_input":"2022-01-06T19:24:12.217457Z","iopub.status.idle":"2022-01-06T19:24:12.224067Z","shell.execute_reply.started":"2022-01-06T19:24:12.217428Z","shell.execute_reply":"2022-01-06T19:24:12.222979Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"wordcloud(data_tox,'toxic')","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:24:12.225188Z","iopub.execute_input":"2022-01-06T19:24:12.225438Z","iopub.status.idle":"2022-01-06T19:24:15.383894Z","shell.execute_reply.started":"2022-01-06T19:24:12.225411Z","shell.execute_reply":"2022-01-06T19:24:15.383267Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"wordcloud(data_ide,'identity_hate')","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:24:15.388598Z","iopub.execute_input":"2022-01-06T19:24:15.389054Z","iopub.status.idle":"2022-01-06T19:24:16.410374Z","shell.execute_reply.started":"2022-01-06T19:24:15.389021Z","shell.execute_reply":"2022-01-06T19:24:16.409627Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"wordcloud(data_sev,'severe_toxic')\n","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:24:16.411825Z","iopub.execute_input":"2022-01-06T19:24:16.412304Z","iopub.status.idle":"2022-01-06T19:24:17.166195Z","shell.execute_reply.started":"2022-01-06T19:24:16.412263Z","shell.execute_reply":"2022-01-06T19:24:17.165477Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"wordcloud(data_obs,'obscene')","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:24:17.167523Z","iopub.execute_input":"2022-01-06T19:24:17.16796Z","iopub.status.idle":"2022-01-06T19:24:19.013178Z","shell.execute_reply.started":"2022-01-06T19:24:17.167928Z","shell.execute_reply":"2022-01-06T19:24:19.012452Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"wordcloud(data_ins,'insult')\n","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:24:19.014491Z","iopub.execute_input":"2022-01-06T19:24:19.01491Z","iopub.status.idle":"2022-01-06T19:24:20.779761Z","shell.execute_reply.started":"2022-01-06T19:24:19.014877Z","shell.execute_reply":"2022-01-06T19:24:20.778921Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"wordcloud(data_thr,'threat')","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:24:20.780856Z","iopub.execute_input":"2022-01-06T19:24:20.781085Z","iopub.status.idle":"2022-01-06T19:24:21.675902Z","shell.execute_reply.started":"2022-01-06T19:24:20.781058Z","shell.execute_reply":"2022-01-06T19:24:21.675012Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_tox.head()","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:24:21.677217Z","iopub.execute_input":"2022-01-06T19:24:21.677472Z","iopub.status.idle":"2022-01-06T19:24:21.690776Z","shell.execute_reply.started":"2022-01-06T19:24:21.677443Z","shell.execute_reply":"2022-01-06T19:24:21.689609Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Lập chỉ mục hoàn toàn dựa trên vị trí số nguyên để lựa chọn theo vị trí.","metadata":{}},{"cell_type":"code","source":"data_tox_1 = data_tox[data_tox['toxic'] == 1].iloc[0:5000,:]\ndata_tox_1.shape","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:24:21.692246Z","iopub.execute_input":"2022-01-06T19:24:21.692551Z","iopub.status.idle":"2022-01-06T19:24:21.704546Z","shell.execute_reply.started":"2022-01-06T19:24:21.69252Z","shell.execute_reply":"2022-01-06T19:24:21.703187Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_tox_0 = data_tox[data_tox['toxic'] == 0].iloc[0:5000,:]","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:24:21.706141Z","iopub.execute_input":"2022-01-06T19:24:21.706996Z","iopub.status.idle":"2022-01-06T19:24:21.738602Z","shell.execute_reply.started":"2022-01-06T19:24:21.706951Z","shell.execute_reply":"2022-01-06T19:24:21.737815Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Nối lại bằng concat","metadata":{}},{"cell_type":"markdown","source":"Axis: Trục nối dọc","metadata":{}},{"cell_type":"code","source":"data_tox_done = pd.concat([data_tox_1, data_tox_0], axis = 0)\ndata_tox_done.shape","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:24:21.739996Z","iopub.execute_input":"2022-01-06T19:24:21.740218Z","iopub.status.idle":"2022-01-06T19:24:21.747937Z","shell.execute_reply.started":"2022-01-06T19:24:21.740194Z","shell.execute_reply":"2022-01-06T19:24:21.747251Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Đếm Severe Toxic","metadata":{}},{"cell_type":"code","source":"data_sev[data_sev['severe_toxic'] == 1].count()","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:24:21.749034Z","iopub.execute_input":"2022-01-06T19:24:21.749383Z","iopub.status.idle":"2022-01-06T19:24:21.764875Z","shell.execute_reply.started":"2022-01-06T19:24:21.74935Z","shell.execute_reply":"2022-01-06T19:24:21.763551Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_sev_1 = data_sev[data_sev['severe_toxic'] == 1].iloc[0:1595,:]\ndata_sev_0 = data_sev[data_sev['severe_toxic'] == 0].iloc[0:1595,:]\ndata_sev_done = pd.concat([data_sev_1, data_sev_0], axis = 0)\ndata_sev_done.shape","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:24:21.766542Z","iopub.execute_input":"2022-01-06T19:24:21.767204Z","iopub.status.idle":"2022-01-06T19:24:21.789758Z","shell.execute_reply.started":"2022-01-06T19:24:21.76716Z","shell.execute_reply":"2022-01-06T19:24:21.788752Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_obs[data_obs['obscene'] == 1].count()","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:24:21.791439Z","iopub.execute_input":"2022-01-06T19:24:21.791753Z","iopub.status.idle":"2022-01-06T19:24:21.809155Z","shell.execute_reply.started":"2022-01-06T19:24:21.791713Z","shell.execute_reply":"2022-01-06T19:24:21.808553Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_obs_1 = data_obs[data_obs['obscene'] == 1].iloc[0:5000,:]\ndata_obs_0 = data_obs[data_obs['obscene'] == 0].iloc[0:5000,:]\ndata_obs_done = pd.concat([data_obs_1, data_obs_0], axis = 0)\ndata_obs_done.shape","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:24:21.810907Z","iopub.execute_input":"2022-01-06T19:24:21.811399Z","iopub.status.idle":"2022-01-06T19:24:21.834796Z","shell.execute_reply.started":"2022-01-06T19:24:21.811354Z","shell.execute_reply":"2022-01-06T19:24:21.8342Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_thr[data_thr['threat'] == 1].count()","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:24:21.835819Z","iopub.execute_input":"2022-01-06T19:24:21.83612Z","iopub.status.idle":"2022-01-06T19:24:21.845967Z","shell.execute_reply.started":"2022-01-06T19:24:21.836095Z","shell.execute_reply":"2022-01-06T19:24:21.845372Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_thr_1 = data_thr[data_thr['threat'] == 1].iloc[0:478,:] #20%\ndata_thr_0 = data_thr[data_thr['threat'] == 0].iloc[0:1912,:]#80%\ndata_thr_done = pd.concat([data_thr_1, data_thr_0], axis = 0)\ndata_thr_done.shape","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:24:21.847254Z","iopub.execute_input":"2022-01-06T19:24:21.847634Z","iopub.status.idle":"2022-01-06T19:24:21.869783Z","shell.execute_reply.started":"2022-01-06T19:24:21.847573Z","shell.execute_reply":"2022-01-06T19:24:21.869195Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_ins[data_ins['insult'] == 1].count()","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:24:21.870817Z","iopub.execute_input":"2022-01-06T19:24:21.871195Z","iopub.status.idle":"2022-01-06T19:24:21.885738Z","shell.execute_reply.started":"2022-01-06T19:24:21.871165Z","shell.execute_reply":"2022-01-06T19:24:21.884706Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_ins_1 = data_ins[data_ins['insult'] == 1].iloc[0:5000,:]\ndata_ins_0 = data_ins[data_ins['insult'] == 0].iloc[0:5000,:]\ndata_ins_done = pd.concat([data_ins_1, data_ins_0], axis = 0)\ndata_ins_done.shape","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:24:21.886795Z","iopub.execute_input":"2022-01-06T19:24:21.887005Z","iopub.status.idle":"2022-01-06T19:24:21.911289Z","shell.execute_reply.started":"2022-01-06T19:24:21.886979Z","shell.execute_reply":"2022-01-06T19:24:21.91043Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_ide[data_ide['identity_hate'] == 1].count()","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:24:21.912685Z","iopub.execute_input":"2022-01-06T19:24:21.913141Z","iopub.status.idle":"2022-01-06T19:24:21.925173Z","shell.execute_reply.started":"2022-01-06T19:24:21.9131Z","shell.execute_reply":"2022-01-06T19:24:21.924488Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_ide_1 = data_ide[data_ide['identity_hate'] == 1].iloc[0:1405,:] #20%\ndata_ide_0 = data_ide[data_ide['identity_hate'] == 0].iloc[0:5620,:] #80%\ndata_ide_done = pd.concat([data_ide_1, data_ide_0], axis = 0)\ndata_ide_done.shape","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:24:21.926284Z","iopub.execute_input":"2022-01-06T19:24:21.92687Z","iopub.status.idle":"2022-01-06T19:24:21.951324Z","shell.execute_reply.started":"2022-01-06T19:24:21.926833Z","shell.execute_reply":"2022-01-06T19:24:21.95024Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Nhắc nhở: Số lượng nhận xét thuộc các danh mục sau:\n* Toxic ( 14000+)\n* Severe Toxic (1595)\n* Obscene (8449)\n* Threat (478)\n* Insult (7877)\n* Identity Hate (1405)\n\n\n\n//-------------------------------------------------------//","metadata":{}},{"cell_type":"markdown","source":"Part 3 : Chạy Mô hình ML trên dữ liệu","metadata":{}},{"cell_type":"code","source":"#nhập các gói để xử lý trước\nfrom sklearn import preprocessing\nfrom sklearn.feature_selection import SelectFromModel\n\n#nhập các công cụ để chia nhỏ dữ liệu và đánh giá hiệu suất mô hình\nfrom sklearn.model_selection import train_test_split, KFold, cross_val_score\nfrom sklearn.metrics import f1_score, precision_score, recall_score, precision_recall_curve, fbeta_score, confusion_matrix\nfrom sklearn.metrics import roc_auc_score, roc_curve\n\n#import ML algos\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.neighbors import KNeighborsClassifier\nfrom sklearn.naive_bayes import MultinomialNB, BernoulliNB\nfrom sklearn.svm import LinearSVC\nfrom sklearn.ensemble import RandomForestClassifier","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:24:21.952728Z","iopub.execute_input":"2022-01-06T19:24:21.952961Z","iopub.status.idle":"2022-01-06T19:24:22.092908Z","shell.execute_reply.started":"2022-01-06T19:24:21.952933Z","shell.execute_reply":"2022-01-06T19:24:22.092282Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Tạo chức năng đơn giản có trong tập dữ liệu và cho phép người dùng chọn tập dữ liệu, nhãn độc tính, vectơ và số lượng ngam","metadata":{}},{"cell_type":"code","source":"def cv_tf_train_test(df_done,label,vectorizer,ngram):\n    #Chia dữ liệu thành các tập dữ liệu X và y\n    X = df_done.comment_text\n    y = df_done[label]\n    \n    #Chia ngày của chúng tôi thành dữ liệu đào tạo và kiểm tra\n    X_train, X_test, y_train, y_test = train_test_split(X,y, test_size=0.3, random_state=42)\n    \n    #Tạo một đối tượng Vectorizer và xóa các stopword dừng khỏi bảng\n    cv1 = vectorizer(ngram_range=(ngram), stop_words='english')\n    \n    X_train_cv1 = cv1.fit_transform(X_train) # Học từ điển từ vựng và trả về ma trận tài liệu thuật ngữ\n    X_test_sv1 = cv1.transform(X_test) #Học từ điển từ vựng của tất cả các mã thông báo trong tài liệu thô\n    \n    # Dùng các model để train\n    lr = LogisticRegression()\n    lr.fit(X_train_cv1, y_train)\n    \n    knn = KNeighborsClassifier(n_neighbors=5)\n    knn.fit(X_train_cv1, y_train)\n    \n    \n    bnb = BernoulliNB()\n    bnb.fit(X_train_cv1, y_train)\n    \n    mnb = MultinomialNB()\n    mnb.fit(X_train_cv1, y_train)\n    \n    svm_model = LinearSVC()\n    svm_model.fit(X_train_cv1, y_train)\n    \n    randomforest = RandomForestClassifier(n_estimators=100, random_state = 42)\n    randomforest.fit(X_train_cv1, y_train)\n    \n    f1_score_data = {'F1 Score':[f1_score(lr.predict(X_test_sv1), y_test), \n                                 f1_score(knn.predict(X_test_sv1), y_test),\n                                 f1_score(bnb.predict(X_test_sv1), y_test),\n                                 f1_score(mnb.predict(X_test_sv1), y_test), \n                                 f1_score(svm_model.predict(X_test_sv1), y_test),\n                                f1_score(randomforest.predict(X_test_sv1), y_test),]}\n    \n    \n    df_f1 = pd.DataFrame(f1_score_data, index=['Log Regression', 'KNN', 'BernoulliNB', 'MultinomialNB', 'SVM', 'Random Forest'])\n    return df_f1","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:24:22.09432Z","iopub.execute_input":"2022-01-06T19:24:22.09502Z","iopub.status.idle":"2022-01-06T19:24:22.10562Z","shell.execute_reply.started":"2022-01-06T19:24:22.094984Z","shell.execute_reply":"2022-01-06T19:24:22.104685Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Tạo một Frame hiện thị tất cả những mô hình mà mình vừa sử dụng\n(Nhận xét sau khi tạo biểu đồ đường)","metadata":{}},{"cell_type":"code","source":"df_tox_cv = cv_tf_train_test(data_tox_done, 'toxic', TfidfVectorizer, (1,1))\ndf_tox_cv.rename(columns={'F1 Score': 'F1 Score(toxic)'}, inplace=True)\ndf_tox_cv","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:24:22.107097Z","iopub.execute_input":"2022-01-06T19:24:22.107684Z","iopub.status.idle":"2022-01-06T19:24:31.904139Z","shell.execute_reply.started":"2022-01-06T19:24:22.10754Z","shell.execute_reply":"2022-01-06T19:24:31.903179Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_sev_cv = cv_tf_train_test(data_sev_done, 'severe_toxic', TfidfVectorizer, (1,1))\ndf_sev_cv.rename(columns={'F1 Score': 'F1 Score(Severe toxic)'}, inplace=True)\ndf_sev_cv","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:24:31.905565Z","iopub.execute_input":"2022-01-06T19:24:31.905814Z","iopub.status.idle":"2022-01-06T19:24:33.743526Z","shell.execute_reply.started":"2022-01-06T19:24:31.905786Z","shell.execute_reply":"2022-01-06T19:24:33.742668Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_obs_cv = cv_tf_train_test(data_obs_done, 'obscene', TfidfVectorizer, (1,1))\ndf_obs_cv.rename(columns={'F1 Score': 'F1 Score(Obscene)'}, inplace=True)\ndf_obs_cv","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:24:33.745048Z","iopub.execute_input":"2022-01-06T19:24:33.745315Z","iopub.status.idle":"2022-01-06T19:24:42.732977Z","shell.execute_reply.started":"2022-01-06T19:24:33.74528Z","shell.execute_reply":"2022-01-06T19:24:42.732075Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_ins_cv = cv_tf_train_test(data_ins_done, 'insult', TfidfVectorizer, (1,1))\ndf_ins_cv.rename(columns={'F1 Score': 'F1 Score(Insult)'}, inplace=True)\ndf_ins_cv","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:24:42.734172Z","iopub.execute_input":"2022-01-06T19:24:42.734439Z","iopub.status.idle":"2022-01-06T19:24:52.126714Z","shell.execute_reply.started":"2022-01-06T19:24:42.734412Z","shell.execute_reply":"2022-01-06T19:24:52.126087Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_thr_cv = cv_tf_train_test(data_thr_done, 'threat', TfidfVectorizer, (1,1))\ndf_thr_cv.rename(columns={'F1 Score': 'F1 Score(Threat)'}, inplace=True)\ndf_thr_cv","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:24:52.127775Z","iopub.execute_input":"2022-01-06T19:24:52.128105Z","iopub.status.idle":"2022-01-06T19:24:53.359974Z","shell.execute_reply.started":"2022-01-06T19:24:52.128077Z","shell.execute_reply":"2022-01-06T19:24:53.359073Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_ide_cv = cv_tf_train_test(data_ide_done, 'identity_hate', TfidfVectorizer, (1,1))\ndf_ide_cv.rename(columns={'F1 Score': 'F1 Score(Identity Hate)'}, inplace=True)\ndf_ide_cv","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:24:53.361175Z","iopub.execute_input":"2022-01-06T19:24:53.361421Z","iopub.status.idle":"2022-01-06T19:24:58.540107Z","shell.execute_reply.started":"2022-01-06T19:24:53.36139Z","shell.execute_reply":"2022-01-06T19:24:58.539285Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Cùng xem xét lại thành 1 bảng hoàn chỉnh","metadata":{}},{"cell_type":"code","source":"f1_all = pd.concat([df_tox_cv, df_sev_cv, df_obs_cv, df_ins_cv, df_thr_cv, df_ide_cv], axis=1)\nf1_all","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:24:58.541451Z","iopub.execute_input":"2022-01-06T19:24:58.54167Z","iopub.status.idle":"2022-01-06T19:24:58.560752Z","shell.execute_reply.started":"2022-01-06T19:24:58.541645Z","shell.execute_reply":"2022-01-06T19:24:58.55999Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Thay đổi góc nhìn 1 chút ta được bảng sau:","metadata":{}},{"cell_type":"code","source":"f1_all_trp = f1_all.transpose() #Trả về chế độ xem của mảng có các trục được hoán vị.\nf1_all_trp","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:24:58.563132Z","iopub.execute_input":"2022-01-06T19:24:58.563609Z","iopub.status.idle":"2022-01-06T19:24:58.576942Z","shell.execute_reply.started":"2022-01-06T19:24:58.563573Z","shell.execute_reply":"2022-01-06T19:24:58.576307Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Tạo một đồ thị đường để so sánh các giá trị mới tìm được:","metadata":{}},{"cell_type":"code","source":"sns.lineplot(data=f1_all_trp, markers=True)\n#sns.relplot(data=flights, x=\"year\", y=\"passengers\", hue=\"month\", kind=\"line\")\nplt.xticks(rotation='90', fontsize=14)\nplt.yticks(fontsize=14)\nplt.legend(loc='best')\nplt.title('F1 Score', fontsize=20)","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:24:58.578013Z","iopub.execute_input":"2022-01-06T19:24:58.5787Z","iopub.status.idle":"2022-01-06T19:24:58.973184Z","shell.execute_reply.started":"2022-01-06T19:24:58.578663Z","shell.execute_reply":"2022-01-06T19:24:58.972257Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Từ các bảng và biểu đồ trên có thể thấy được\n* Random Forest và SVM có độ ổn định và đánh giá cao hơn hẳn\n* KNN và BernoulliNB không cao lắm # ==>Sẽ dùng RandomForest để dự đoán","metadata":{}},{"cell_type":"code","source":"#Xây dựng lại 1 hàm để huấn luyện data\ndef train_test(df_done, label, vectorizer, ngram):\n    X = df_done.comment_text\n    y = df_done[label]\n    \n    #chia thành các phần tử con để thực hiện huấn luyện\n    X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42)\n    \n    #Tạo Vectorizer object và xóa stopwords khỏi bảng\n    cv1 = vectorizer(ngram_range=(ngram), stop_words='english')\n    \n    X_train_cv1 = cv1.fit_transform(X_train) \n    print(X_train_cv1)\n    X_test_cv1 = cv1.transform(X_test)\n    \n    #Chọn mô hình Random Forest\n    randomforest = RandomForestClassifier(n_estimators=100, random_state=42)\n    randomforest.fit(X_train_cv1, y_train)\n    return randomforest.predict(X_test_cv1)","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:24:58.974557Z","iopub.execute_input":"2022-01-06T19:24:58.974802Z","iopub.status.idle":"2022-01-06T19:24:58.981804Z","shell.execute_reply.started":"2022-01-06T19:24:58.974772Z","shell.execute_reply":"2022-01-06T19:24:58.980709Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Chuyển đổi một bộ sưu tập các tài liệu thô thành một ma trận các tính năng TF-IDF.","metadata":{}},{"cell_type":"code","source":"df_tox = train_test(data_tox_done,'toxic', TfidfVectorizer, (1,1))\ndf_tox","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:24:58.983129Z","iopub.execute_input":"2022-01-06T19:24:58.983475Z","iopub.status.idle":"2022-01-06T19:25:07.365087Z","shell.execute_reply.started":"2022-01-06T19:24:58.983434Z","shell.execute_reply":"2022-01-06T19:25:07.36414Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_sev = train_test(data_sev_done,'severe_toxic', TfidfVectorizer, (1,1))\ndf_sev","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:25:07.366128Z","iopub.execute_input":"2022-01-06T19:25:07.366347Z","iopub.status.idle":"2022-01-06T19:25:08.8852Z","shell.execute_reply.started":"2022-01-06T19:25:07.366323Z","shell.execute_reply":"2022-01-06T19:25:08.884286Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_obs = train_test(data_obs_done,'obscene', TfidfVectorizer, (1,1))\ndf_obs","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:25:08.886643Z","iopub.execute_input":"2022-01-06T19:25:08.886957Z","iopub.status.idle":"2022-01-06T19:25:16.674352Z","shell.execute_reply.started":"2022-01-06T19:25:08.886914Z","shell.execute_reply":"2022-01-06T19:25:16.673511Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_ins = train_test(data_ins_done,'insult', TfidfVectorizer, (1,1))\ndf_ins","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:25:16.675456Z","iopub.execute_input":"2022-01-06T19:25:16.675673Z","iopub.status.idle":"2022-01-06T19:25:25.039637Z","shell.execute_reply.started":"2022-01-06T19:25:16.67565Z","shell.execute_reply":"2022-01-06T19:25:25.038362Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_thr = train_test(data_thr_done,'threat', TfidfVectorizer, (1,1))\ndf_thr","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:25:25.042388Z","iopub.execute_input":"2022-01-06T19:25:25.042993Z","iopub.status.idle":"2022-01-06T19:25:26.045759Z","shell.execute_reply.started":"2022-01-06T19:25:25.042891Z","shell.execute_reply":"2022-01-06T19:25:26.044912Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_ide = train_test(data_ide_done,'identity_hate', TfidfVectorizer, (1,1))\ndf_ide","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:25:26.04684Z","iopub.execute_input":"2022-01-06T19:25:26.047055Z","iopub.status.idle":"2022-01-06T19:25:30.410975Z","shell.execute_reply.started":"2022-01-06T19:25:26.047029Z","shell.execute_reply":"2022-01-06T19:25:30.410303Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"//---------------------------------------------//","metadata":{}},{"cell_type":"markdown","source":"Bây giờ sẽ dự đoán thử khả năng Toxic của một câu","metadata":{}},{"cell_type":"code","source":" X = data_tox_done.comment_text\n    y = data_tox_done['toxic']\n    \n    #chia thành các phần tử con để thực hiện huấn luyện\n    X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42)\n    \n    #Tạo Vectorizer object và xóa stopwords khỏi bảng\n    tfv = TfidfVectorizer(ngram_range=(1,1), stop_words='english')\n    \n    X_train_fit = tfv.fit_transform(X_train) \n    X_test_fit = tfv.transform(X_test)\n    \n    #Chọn mô hình Random Forest\n    randomforest = RandomForestClassifier(n_estimators=100, random_state=42)\n    randomforest.fit(X_train_fit, y_train)\n    randomforest.predict(X_test_fit)","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:25:30.412067Z","iopub.execute_input":"2022-01-06T19:25:30.412414Z","iopub.status.idle":"2022-01-06T19:25:30.418858Z","shell.execute_reply.started":"2022-01-06T19:25:30.412378Z","shell.execute_reply":"2022-01-06T19:25:30.417654Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Một function dự đoán khả năng toxic","metadata":{}},{"cell_type":"code","source":"#Predict Toxic Function \ndef predictToxic(sample):\n    vect = tfv.transform(sample)\n    return randomforest.predict_proba(vect)[:,1:]","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:25:30.419852Z","iopub.status.idle":"2022-01-06T19:25:30.420312Z","shell.execute_reply.started":"2022-01-06T19:25:30.420061Z","shell.execute_reply":"2022-01-06T19:25:30.420086Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Dự đoán khả năng củ từ \"Fuck you Nigga\"","metadata":{}},{"cell_type":"code","source":"print('Du doan cua Toxic: ', predictToxic(['Fuck you nigga']))","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:25:30.42211Z","iopub.status.idle":"2022-01-06T19:25:30.422703Z","shell.execute_reply.started":"2022-01-06T19:25:30.422357Z","shell.execute_reply":"2022-01-06T19:25:30.422423Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":" >Khá là chính xác","metadata":{}},{"cell_type":"markdown","source":"Dự đoán từ \"I Love You\"","metadata":{}},{"cell_type":"code","source":"print('Du doan cua Toxic: ', predictToxic(['I Love You']))","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:25:30.4247Z","iopub.status.idle":"2022-01-06T19:25:30.425468Z","shell.execute_reply.started":"2022-01-06T19:25:30.42516Z","shell.execute_reply":"2022-01-06T19:25:30.425187Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":">Không có vấn đề gì cả, dưới 0.5","metadata":{}},{"cell_type":"markdown","source":"Dự đoán từ\"How are you today\"","metadata":{}},{"cell_type":"code","source":"print('Du doan cua Toxic: ', predictToxic(['How are you today']))","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:25:30.42685Z","iopub.status.idle":"2022-01-06T19:25:30.427696Z","shell.execute_reply.started":"2022-01-06T19:25:30.427414Z","shell.execute_reply":"2022-01-06T19:25:30.427453Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Bây giờ sẽ dự đoán thử khả năng Severe Toxic của một câu","metadata":{}},{"cell_type":"code","source":"X = data_sev_done.comment_text\n    y = data_sev_done['severe_toxic']\n    \n    #chia thành các phần tử con để thực hiện huấn luyện\n    X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42)\n    \n    #Tạo Vectorizer object và xóa stopwords khỏi bảng\n    tfv = TfidfVectorizer(ngram_range=(1,1), stop_words='english')\n    \n    X_train_fit = tfv.fit_transform(X_train) \n    X_test_fit = tfv.transform(X_test)\n    \n    #Chọn mô hình Random Forest\n    randomforest = RandomForestClassifier(n_estimators=100, random_state=42)\n    randomforest.fit(X_train_fit, y_train)\n    randomforest.predict(X_test_fit)","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:25:30.42916Z","iopub.status.idle":"2022-01-06T19:25:30.429901Z","shell.execute_reply.started":"2022-01-06T19:25:30.429643Z","shell.execute_reply":"2022-01-06T19:25:30.42967Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Severe Toxic Function","metadata":{}},{"cell_type":"code","source":"#Predict Severe Toxic Function \ndef predict(sample):\n    vect = tfv.transform(sample)\n    return randomforest.predict_proba(vect)[:,1:]","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:25:30.431339Z","iopub.status.idle":"2022-01-06T19:25:30.432044Z","shell.execute_reply.started":"2022-01-06T19:25:30.431788Z","shell.execute_reply":"2022-01-06T19:25:30.431816Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Du doan cua Severe Toxic: ', predict(['Fuck you nigga']))","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:25:30.43317Z","iopub.status.idle":"2022-01-06T19:25:30.433497Z","shell.execute_reply.started":"2022-01-06T19:25:30.433346Z","shell.execute_reply":"2022-01-06T19:25:30.433362Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Du doan cua Severe Toxic: ', predict(['I love you']))","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:25:30.434731Z","iopub.status.idle":"2022-01-06T19:25:30.435022Z","shell.execute_reply.started":"2022-01-06T19:25:30.434873Z","shell.execute_reply":"2022-01-06T19:25:30.434889Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Severe Toxic khá là chính xác**","metadata":{}},{"cell_type":"markdown","source":"Dự đoán obscene","metadata":{}},{"cell_type":"code","source":"X = data_obs_done.comment_text\n    y = data_obs_done['obscene']\n    \n    #chia thành các phần tử con để thực hiện huấn luyện\n    X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42)\n    \n    #Tạo Vectorizer object và xóa stopwords khỏi bảng\n    tfv = TfidfVectorizer(ngram_range=(1,1), stop_words='english')\n    \n    X_train_fit = tfv.fit_transform(X_train) \n    X_test_fit = tfv.transform(X_test)\n    \n    #Chọn mô hình Random Forest\n    randomforest = RandomForestClassifier(n_estimators=100, random_state=42)\n    randomforest.fit(X_train_fit, y_train)\n    randomforest.predict(X_test_fit)\n    \n    #Predict Obscene Function \n    def predict(sample):\n        vect = tfv.transform(sample)\n        return randomforest.predict_proba(vect)[:,1:]\n\n    print('Du doan cua Obscene: ', predict(['Fuck you nigga']))\n    print('Du doan cua Obscene: ', predict(['Ilove you']))","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:25:30.435898Z","iopub.status.idle":"2022-01-06T19:25:30.436174Z","shell.execute_reply.started":"2022-01-06T19:25:30.43603Z","shell.execute_reply":"2022-01-06T19:25:30.436045Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Chính xác tuyệt đối**","metadata":{}},{"cell_type":"markdown","source":"Dự đoán Insult","metadata":{}},{"cell_type":"code","source":" X = data_ins_done.comment_text\n    y = data_ins_done['insult']\n    \n    #chia thành các phần tử con để thực hiện huấn luyện\n    X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42)\n    \n    #Tạo Vectorizer object và xóa stopwords khỏi bảng\n    tfv = TfidfVectorizer(ngram_range=(1,1), stop_words='english')\n    \n    X_train_fit = tfv.fit_transform(X_train) \n    X_test_fit = tfv.transform(X_test)\n    \n    #Chọn mô hình Random Forest\n    randomforest = RandomForestClassifier(n_estimators=100, random_state=42)\n    randomforest.fit(X_train_fit, y_train)\n    randomforest.predict(X_test_fit)\n    \n    #Predict Obscene Function \n    def predict(sample):\n        vect = tfv.transform(sample)\n        return randomforest.predict_proba(vect)[:,1:]\n\n    print('Du doan cua Insult: ', predict(['Fuck you nigga']))\n    print('Du doan cua Insult: ', predict(['Ilove you']))","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:25:30.437728Z","iopub.status.idle":"2022-01-06T19:25:30.438178Z","shell.execute_reply.started":"2022-01-06T19:25:30.438023Z","shell.execute_reply":"2022-01-06T19:25:30.43804Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Gần như là hoàn hảo**","metadata":{}},{"cell_type":"markdown","source":"Dự đoán Threat","metadata":{}},{"cell_type":"code","source":" X = data_thr_done.comment_text\n    y = data_thr_done['threat']\n    \n    #chia thành các phần tử con để thực hiện huấn luyện\n    X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42)\n    \n    #Tạo Vectorizer object và xóa stopwords khỏi bảng\n    tfv = TfidfVectorizer(ngram_range=(1,1), stop_words='english')\n    \n    X_train_fit = tfv.fit_transform(X_train) \n    X_test_fit = tfv.transform(X_test)\n    \n    #Chọn mô hình Random Forest\n    randomforest = RandomForestClassifier(n_estimators=100, random_state=42)\n    randomforest.fit(X_train_fit, y_train)\n    randomforest.predict(X_test_fit)\n    \n    #Predict Obscene Function \n    def predict(sample):\n        vect = tfv.transform(sample)\n        return randomforest.predict_proba(vect)[:,1:]\n\n    print('Du doan cua Threat: ', predict(['Fuck you nigga']))\n    print('Du doan cua Threat: ', predict(['Ilove you']))","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:25:30.439261Z","iopub.status.idle":"2022-01-06T19:25:30.439792Z","shell.execute_reply.started":"2022-01-06T19:25:30.439616Z","shell.execute_reply":"2022-01-06T19:25:30.439635Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Lỗi**","metadata":{}},{"cell_type":"markdown","source":"Dự đoán identity Hate","metadata":{}},{"cell_type":"code","source":"X = data_ide_done.comment_text\n    y = data_ide_done['identity_hate']\n    \n    #chia thành các phần tử con để thực hiện huấn luyện\n    X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42)\n    \n    #Tạo Vectorizer object và xóa stopwords khỏi bảng\n    tfv = TfidfVectorizer(ngram_range=(1,1), stop_words='english')\n    \n    X_train_fit = tfv.fit_transform(X_train) \n    X_test_fit = tfv.transform(X_test)\n    \n    #Chọn mô hình Random Forest\n    randomforest = RandomForestClassifier(n_estimators=100, random_state=42)\n    randomforest.fit(X_train_fit, y_train)\n    randomforest.predict(X_test_fit)\n    \n    #Predict Obscene Function \n    def predict(sample):\n        vect = tfv.transform(sample)\n        return randomforest.predict_proba(vect)[:,1:]\n\n    print('Du doan cua Identity Hate: ', predict(['Fuck you nigga']))\n    print('Du doan cua Identity Hate: ', predict(['Ilove you']))","metadata":{"execution":{"iopub.status.busy":"2022-01-06T19:25:30.440751Z","iopub.status.idle":"2022-01-06T19:25:30.441024Z","shell.execute_reply.started":"2022-01-06T19:25:30.440884Z","shell.execute_reply":"2022-01-06T19:25:30.440898Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Như vậy là chúng ta đã test thử với 2 câu\n* \"Fuck you nigga\"\n* I Love You # Có thể thấy, kết quả khá là cao nhưng vẫn còn trường threat là hơi sai 1 chút nhưng nhìn chung là mô hình của chúng ta khá hoàn chỉnh với RandomForest\n\n\n\n-------------------------------------------","metadata":{}},{"cell_type":"markdown","source":"Bây giờ chúng ta cùng nói về thuật toán mình chọn trong bài toán","metadata":{}},{"cell_type":"markdown","source":"Random Forest là gì và tại sao nó lại tốt\n* Random = Tính ngẫu nhiên\n* Forest = Rừng = Nhiều cây quyết định","metadata":{}},{"cell_type":"markdown","source":"**Random forest là một trong những thuật toán học máy mạnh mẽ và phổ biến nhất, nó là thuật toán supervised learning, có thể giải quyết cả bài toán regression và classification. Random Forest là sự cải tiến của bagging. Nó sử dụng các cây (tree) để làm nền tảng, là một tập hợp của hàng trăm cây quyết định, trong đó mỗi cây được tạo nên ngẫu nhiên từ việc tái chọn mẫu (chọn random 1 phần của dữ liệu để xây dựng) và random các đặc trưng (feature) từ toàn bộ dữ liệu.******","metadata":{}},{"cell_type":"markdown","source":"**Ví dụ cho dễ hiểu: Bạn muốn đi mua một thứ gì đó nhưng bạn muốn có được những thứ tốt nhất, bạn phải cân nhắc địa điểm mua hàng cho nên, bạn phải tham khảo ý kiến của nhiều nơi khác nhau. Mỗi một ý kiến ở đây sẽ đóng vai trò như một CÂY QUYẾT ĐỊNH trả lời cho những câu hỏi của bạn. Rồi sau đó bạn sẽ có một loạt câu trả lời cho những câu hỏi của bạn, từ đó chọn được phương án tốt nhất.**","metadata":{}},{"cell_type":"markdown","source":"**Random Forest hoạt động cũng như thế, mỗi cây quyết định được xây dựng dùng thuật toán Decision Tree trên tập dữ liệu khác nhau và dùng tập thuộc tính khác nhau. Sau đó bằng cách đánh giá các cây quyết định sử dụng cách thức voting để đưa ra kết quả cuối cùng cho bài toán**","metadata":{}},{"cell_type":"markdown","source":"**Nếu thuật toán Random Forest có 6 cây quyết đinh, 5 cây dự đoán 1, 1 cây dự đoán 0, do đó mình sẽ lấy và cho ra dự đoán cuối cùng là 1**","metadata":{}},{"cell_type":"markdown","source":"Ưu và nhược điểm của thuật toán Random Forest","metadata":{}},{"cell_type":"markdown","source":"1. Ưu điểm\n* Thuật toán dễ sử dụng và mạnh mẽ\n* Giảm phương án sai, tránh bị Overfitting\n* Có thể sử dụng cho cả 2 loại bài toán là\n* Nó có thể làm được với những bài toán bị thiếu dữ liệu\n2. Nhược điểm\n* Mất khả năng diễn giải của mô hình\n* Bagging rất mạnh, cho chúng ta độ chính xác cao hơn, nhưng lại nặng về mặt tính toán và có thể có trường hợp không có được kết quả mong muốn giống như trường threat của chúng ta ở trên.","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}}]}