{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Quora Incinsere Questions Classification sử dụng Logistic Regression\n","metadata":{}},{"cell_type":"code","source":"# import các thư viện cần thiết\nimport os\nimport time\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt \nfrom sklearn.model_selection import train_test_split\nfrom sklearn.feature_extraction.text import CountVectorizer\nfrom wordcloud import WordCloud, STOPWORDS, ImageColorGenerator\nimport warnings\nwarnings.filterwarnings(\"ignore\")\n# import spacy\nimport re\nfrom tqdm import tqdm\nimport nltk\nfrom tqdm import tqdm_notebook\ntqdm_notebook().pandas()","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2021-06-10T16:05:17.410444Z","iopub.execute_input":"2021-06-10T16:05:17.410896Z","iopub.status.idle":"2021-06-10T16:05:19.121571Z","shell.execute_reply.started":"2021-06-10T16:05:17.410809Z","shell.execute_reply":"2021-06-10T16:05:19.120627Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 1. Phân tích dữ liệu\n#### Trước hết ta load dữ liệu vào dataframe và in ra để quan sát. Mục đích của bước này là để quan sát được dữ liệu câu hỏi gồm những trường gì, các trường thuộc kiểu dữ liệu gì để có thể xử lý và quyết định hướng tiếp cận giải quyết vấn đề.","metadata":{}},{"cell_type":"code","source":"train_raw = pd.read_csv(\"../input/quora-insincere-questions-classification/train.csv\")\nvalidation_data = pd.read_csv(\"../input/quora-insincere-questions-classification/test.csv\")\ntrain_raw","metadata":{"execution":{"iopub.status.busy":"2021-06-10T16:05:19.122989Z","iopub.execute_input":"2021-06-10T16:05:19.123288Z","iopub.status.idle":"2021-06-10T16:05:25.200860Z","shell.execute_reply.started":"2021-06-10T16:05:19.123258Z","shell.execute_reply":"2021-06-10T16:05:25.199885Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Dữ liệu\n- 1306122 hàng × 3 cột\n- Chưa nhìn thấy được phân bố phân lớp của câu hỏi\n#### Các feature gồm id câu hỏi, câu hỏi, phân lớp của câu hỏi:\n- qid: id độc nhất của câu hỏi, qid chắc sẽ không tham gia vào bước phân lớp câu hỏi nên có thể bỏ được\n- question_text: dữ liệu câu hỏi, do trường này là duy nhất tác động trực tiếp vào phân lớp của câu hỏi nên cần phải thực hiện tiền xử lý\n- target: phân lớp của câu hỏi, target = 0 với câu hỏi sincere và target = 1 với câu hỏi incinsere\n\n","metadata":{}},{"cell_type":"code","source":"y = train_raw['target']\ny.value_counts().plot(kind='bar', rot=0)","metadata":{"execution":{"iopub.status.busy":"2021-06-10T16:05:25.203158Z","iopub.execute_input":"2021-06-10T16:05:25.203508Z","iopub.status.idle":"2021-06-10T16:05:25.507514Z","shell.execute_reply.started":"2021-06-10T16:05:25.203475Z","shell.execute_reply":"2021-06-10T16:05:25.506463Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- Số lượng câu hỏi sincere nhiều hơn rất nhiều so với insincere","metadata":{}},{"cell_type":"markdown","source":"### Tạo thêm một số feature để dễ quan sát, có thể sử dụng sau từ dữ liệu câu hỏi:\n#### Dữ liệu của bài toán chỉ có trường văn bản của câu hỏi nên cần phải tạo thêm một số các feature để có thể quan sát kỹ hơn. Trong đó các feature mới được tạo ra là:\n- số từ\n- số từ độc nhất\n- số ký tự đặc biệt\n- số từ in hoa\n- số từ không in hoa\n- in hoa đầu từ","metadata":{}},{"cell_type":"code","source":"def create_features(df_):\n    \n    df_[\"nb_words\"] = df_[\"question_text\"].apply(lambda x: len(x.split())) # số từ\n    df_[\"nb_unique_words\"] = df_[\"question_text\"].apply(lambda x: len(set(str(x).split()))) # độ dài set\n    df_[\"nb_chars\"] = df_[\"question_text\"].apply(lambda x: len(str(x))) # số ký tự\n    df_['spe_chars'] = df_['question_text'].str.findall(r'[^a-zA-Z0-9 ]').str.len()\n    df_[\"nb_uppercase\"] = df_[\"question_text\"].apply(lambda x : len([nu for nu in str(x).split() if nu.isupper()]))\n    df_[\"nb_lowercase\"] = df_[\"question_text\"].apply(lambda x : len([nl for nl in str(x).split() if nl.islower()]))\n    df_[\"nb_title\"] = df_[\"question_text\"].apply(lambda x : len([nl for nl in str(x).split() if nl.istitle()]))\n\n    return df_\n\ntrain_features = create_features(train_raw)","metadata":{"execution":{"iopub.status.busy":"2021-06-10T16:05:25.509068Z","iopub.execute_input":"2021-06-10T16:05:25.509379Z","iopub.status.idle":"2021-06-10T16:05:48.868292Z","shell.execute_reply.started":"2021-06-10T16:05:25.509348Z","shell.execute_reply":"2021-06-10T16:05:48.867267Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Dữ liệu câu hỏi sincere","metadata":{}},{"cell_type":"code","source":"train_features[train_features['target'] == 0].describe().round(1)","metadata":{"execution":{"iopub.status.busy":"2021-06-10T16:05:48.869655Z","iopub.execute_input":"2021-06-10T16:05:48.869979Z","iopub.status.idle":"2021-06-10T16:05:49.421110Z","shell.execute_reply.started":"2021-06-10T16:05:48.869935Z","shell.execute_reply":"2021-06-10T16:05:49.419892Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Dữ liệu câu hỏi insincere","metadata":{}},{"cell_type":"code","source":"train_features[train_features['target'] == 1].describe().round(1)","metadata":{"execution":{"iopub.status.busy":"2021-06-10T16:05:49.422470Z","iopub.execute_input":"2021-06-10T16:05:49.422836Z","iopub.status.idle":"2021-06-10T16:05:49.514645Z","shell.execute_reply.started":"2021-06-10T16:05:49.422753Z","shell.execute_reply":"2021-06-10T16:05:49.513467Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Nhận xét:\n- Câu hỏi toxic có trung bình(mean) của số từ, số lượng ký tự nhiều hơn câu hỏi non toxic \n- Có thể sử dụng các feature này vào model (đã thử và không có hiệu quả với các mô hình tuyến tính)\n- Có thể sử dụng vector đếm từ, về căn bản vector đếm từ có khả năng lưu lại các feature như số lượng từ, số lượng từ độc nhất\n- Các câu hỏi có số lượng ký tự đặc biệt hoặc số lượng từ đạt max thường có giá trị lớn hơn nhiều so với trung bình nên cần phải xem xét thêm","metadata":{}},{"cell_type":"code","source":"## Câu hỏi insincere\nprint(train_raw['question_text'][(train_raw['target']==1)].sample(10).values)","metadata":{"execution":{"iopub.status.busy":"2021-06-10T16:05:49.516482Z","iopub.execute_input":"2021-06-10T16:05:49.516923Z","iopub.status.idle":"2021-06-10T16:05:49.540204Z","shell.execute_reply.started":"2021-06-10T16:05:49.516872Z","shell.execute_reply":"2021-06-10T16:05:49.538996Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Nhận xét:\n- Các câu hỏi insincere thường có các từ, cụm từ mang nghĩa xấu, ít khi phụ thuộc vào ngữ pháp\n- Có thể sử dụng vector đếm từ vì không cần giữ lại ngữ pháp câu\n- Các ký tự đặc biệt, chữ số, đường dẫn, in hoa hay in thường không ảnh hưởng nhiều đến phân lớp của câu hỏi nên có thể bỏ","metadata":{}},{"cell_type":"markdown","source":"#### Xem xét thêm về dữ liệu câu hỏi insincere\n- Số lượng ký tự, ký tự đặc biệt trong câu hỏi\n- Tần suất các từ trong câu hỏi","metadata":{}},{"cell_type":"code","source":"plt.hist(train_features['nb_chars'][train_features['target'] == 1], bins=[0,50,100,150,200,250,300,400,500,1000])\nplt.title('number of characters')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2021-06-10T16:05:49.544400Z","iopub.execute_input":"2021-06-10T16:05:49.544711Z","iopub.status.idle":"2021-06-10T16:05:49.709322Z","shell.execute_reply.started":"2021-06-10T16:05:49.544682Z","shell.execute_reply":"2021-06-10T16:05:49.708182Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.hist(train_features['spe_chars'], bins=[1,2,5,10,20,50])\nplt.title('special characters')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2021-06-10T16:05:49.710923Z","iopub.execute_input":"2021-06-10T16:05:49.711226Z","iopub.status.idle":"2021-06-10T16:05:49.894456Z","shell.execute_reply.started":"2021-06-10T16:05:49.711198Z","shell.execute_reply":"2021-06-10T16:05:49.893403Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- Hầu hết các câu hỏi insincere có số lượng ký tự nằm trong khoảng [0,300] và số lượng ký tự đặc biệt từ [0,10]. \n- Như vậy có rất nhiều dữ liệu nhiễu do từ bảng thống kê trên, câu hỏi có số lượng ký tự nhiều nhất là 1017 chữ và 411 ký tự.","metadata":{}},{"cell_type":"markdown","source":"### Sử dụng wordcloud để xem tần suất của các từ trong câu hỏi insincere","metadata":{}},{"cell_type":"code","source":"wordcloud = WordCloud(width=800, height=600, collocations=False).generate(\" \".join(train_raw['question_text'][train_raw['target']==1]))\nplt.figure(figsize=(8,8))\nplt.axis(\"off\")\nplt.imshow(wordcloud,interpolation='bilinear')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2021-06-10T16:05:49.895890Z","iopub.execute_input":"2021-06-10T16:05:49.896301Z","iopub.status.idle":"2021-06-10T16:05:53.254467Z","shell.execute_reply.started":"2021-06-10T16:05:49.896256Z","shell.execute_reply":"2021-06-10T16:05:53.253376Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- Các câu hỏi insincere thường có nhiều các từ mang nghĩa xấu\n- Tuy nhiên một số các từ không mang nghĩa xấu có tần suất cao như people, will, many, much. Các từ này thuộc stopwords, tức là các từ cần thiết trong ngữ pháp nhưng không mang lại nhiều ý nghĩa khi xét từng từ riêng lẻ.","metadata":{}},{"cell_type":"markdown","source":"### Xem xét dữ liệu nhiễu, biên\n- Từ các bảng trên, ta nhận thấy một số câu hỏi có số lượng từ, ký tự, ký tự đặc biệt nhiều hơn rất nhiều so với mean của toàn tập nên cần xem xét để bỏ","metadata":{}},{"cell_type":"code","source":"print(train_features['question_text'][(train_raw['target']==1) & (train_features['nb_chars']>600.0)].values)\nprint(train_features['question_text'][(train_raw['target']==1) & (train_features['spe_chars']>30.0)].values)","metadata":{"execution":{"iopub.status.busy":"2021-06-10T16:05:53.255897Z","iopub.execute_input":"2021-06-10T16:05:53.256224Z","iopub.status.idle":"2021-06-10T16:05:53.274484Z","shell.execute_reply.started":"2021-06-10T16:05:53.256192Z","shell.execute_reply":"2021-06-10T16:05:53.273395Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Nhận xét:\n- Hầu hết các câu hỏi có nhiều ký tự đặc biệt là câu hỏi chứa công thức toán hoặc chữ tượng hình\n- Một số câu hỏi liên quan đến toán bị xếp vào insincere, có thể là nhiễu nên bỏ được","metadata":{}},{"cell_type":"markdown","source":"## 2. Xử lý dữ liệu câu hỏi\n### Sau bước phân tích dữ liệu, ta nhận thấy dữ liệu cần phải được xử lý vì các vấn đề:\n- Tập dataset có tỉ lệ câu hỏi sincere:insincere là 15:1, không cân bằng\n- Phải xử lý dữ liệu thô của văn bản do có chứa các ký tự đặc biệt, các từ stopwords không hữu ích,...\n- Loại bỏ các câu hỏi có số lượng ký tự đặc biệt, số lượng từ lớn hơn nhiều so với trung bình.","metadata":{}},{"cell_type":"markdown","source":"### Bỏ các câu hỏi có số kí tự, số kí tự đặc biệt vượt ngưỡng\n- Dữ liệu câu hỏi sẽ được lọc để cắt bớt phần biên","metadata":{}},{"cell_type":"code","source":"# train_features_filtered = train_features.drop(train_features[(train_features['nb_chars'] >= 600) & (train_features['nb_unique_words']>35.0)].index)\ntrain_features_filtered = train_features[(train_features['nb_chars']<600.0) & (train_features['nb_words']<70.0) & (train_features['spe_chars']<12.0)]\n\ntrain_features_filtered.describe().round(1)","metadata":{"execution":{"iopub.status.busy":"2021-06-10T16:13:06.991117Z","iopub.execute_input":"2021-06-10T16:13:06.991520Z","iopub.status.idle":"2021-06-10T16:13:07.551202Z","shell.execute_reply.started":"2021-06-10T16:13:06.991477Z","shell.execute_reply":"2021-06-10T16:13:07.550204Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Nhận xét:\n- Đã bỏ 3494 hàng\n- Số ký tự nhiều nhất là 300, ký tự đặc biệt là 11 sau khi bỏ 3494 câu hỏi, như vậy vẫn giữ được phần lớn số lượng câu hỏi nhưng bỏ được các câu hỏi nhiễu.","metadata":{}},{"cell_type":"markdown","source":"## Resample\n- Số lượng câu hỏi sincere nhiều hơn rất nhiều so với câu hỏi insincere (gấp hơn 15 lần), dataset không cân bằng, gây ra overfit trên câu hỏi sincere, làm cho accuracy khi dự đoán câu hỏi insincere thấp\n- Cần resample lại dataset cho cân bằng \n- Sau khi chia theo nhiều tỉ lệ, tỉ lệ 4:1 cho kết quả tốt nhất","metadata":{}},{"cell_type":"code","source":"# Resampling\nfrom sklearn.utils import resample\n\n# sincere = train_raw[train_raw.target == 0]\n# insincere = train_raw[train_raw.target == 1]\n\nsincere = train_features_filtered[train_features_filtered.target == 0]\ninsincere = train_features_filtered[train_features_filtered.target == 1]\n\n# Tỉ lệ 1:1\n# x = pd.concat([resample(sincere,\n#                      replace = False,\n#                      n_samples = len(insincere)), insincere])\n\n# Tỉ lệ 2:1\n# x = pd.concat([resample(sincere,\n#                      replace = True,\n#                      n_samples = len(insincere)*2), insincere])\n\n# Tỉ lệ 3:1\n# x = pd.concat([resample(sincere,\n#                      replace = True,\n#        n_samples = len(insincere)*3), insincere])\n\n# 4:1\nx = pd.concat([resample(sincere,\n                     replace = True,\n                     n_samples = len(insincere)*4), insincere])\n","metadata":{"execution":{"iopub.status.busy":"2021-06-10T16:05:53.757021Z","iopub.execute_input":"2021-06-10T16:05:53.757314Z","iopub.status.idle":"2021-06-10T16:05:54.136404Z","shell.execute_reply.started":"2021-06-10T16:05:53.757286Z","shell.execute_reply":"2021-06-10T16:05:54.135132Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y = x['target']\ny.value_counts().plot(kind='bar', rot=0)","metadata":{"execution":{"iopub.status.busy":"2021-06-10T16:05:54.137919Z","iopub.execute_input":"2021-06-10T16:05:54.138397Z","iopub.status.idle":"2021-06-10T16:05:54.269653Z","shell.execute_reply.started":"2021-06-10T16:05:54.138340Z","shell.execute_reply":"2021-06-10T16:05:54.268751Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Nhận xét: Tập train giờ đã cân bằng hơn","metadata":{}},{"cell_type":"markdown","source":"## Xử lý dữ liệu văn bản của câu hỏi\n#### Từ kết quả của bước phân tích dữ liệu, ta có thể loại bỏ khỏi câu hỏi các dữ liệu không cần thiết và chuyển một số dữ liệu về dạng nguyên gốc:\n- Bỏ đường link\n- Loại bỏ ký tự đặc biệt\n- Chuyển các từ cùng biến thể của một từ về một từ duy nhất\n- Chuyển dạng rút gọn từ thành nguyên bản\n- Bỏ chữ số\n- Bỏ công thức toán trong tag latex\n- Bỏ stopword","metadata":{}},{"cell_type":"code","source":"# Bỏ đường link\ndef clean_tag(x):\n    if 'http' in x or 'www' in x:\n        x = re.sub('(?:(?:https?|ftp):\\/\\/)?[\\w/\\-?=%.]+\\.[\\w/\\-?=%.]+', '[url]', x) #replacing with [url]\n    return x","metadata":{"execution":{"iopub.status.busy":"2021-06-10T16:05:54.270892Z","iopub.execute_input":"2021-06-10T16:05:54.271195Z","iopub.status.idle":"2021-06-10T16:05:54.276266Z","shell.execute_reply.started":"2021-06-10T16:05:54.271166Z","shell.execute_reply":"2021-06-10T16:05:54.275205Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Chuyển dạng rút gọn từ thành nguyên bản\n\ncontraction_mapping = {\"We'd\": \"We had\", \"That'd\": \"That had\", \"AREN'T\": \"Are not\", \"HADN'T\": \"Had not\", \"Could've\": \"Could have\", \"LeT's\": \"Let us\", \"How'll\": \"How will\", \"They'll\": \"They will\", \"DOESN'T\": \"Does not\", \"HE'S\": \"He has\", \"O'Clock\": \"Of the clock\", \"Who'll\": \"Who will\", \"What'S\": \"What is\", \"Ain't\": \"Am not\", \"WEREN'T\": \"Were not\", \"Y'all\": \"You all\", \"Y'ALL\": \"You all\", \"Here's\": \"Here is\", \"It'd\": \"It had\", \"Should've\": \"Should have\", \"I'M\": \"I am\", \"ISN'T\": \"Is not\", \"Would've\": \"Would have\", \"He'll\": \"He will\", \"DON'T\": \"Do not\", \"She'd\": \"She had\", \"WOULDN'T\": \"Would not\", \"She'll\": \"She will\", \"IT's\": \"It is\", \"There'd\": \"There had\", \"It'll\": \"It will\", \"You'll\": \"You will\", \"He'd\": \"He had\", \"What'll\": \"What will\", \"Ma'am\": \"Madam\", \"CAN'T\": \"Can not\", \"THAT'S\": \"That is\", \"You've\": \"You have\", \"She's\": \"She is\", \"Weren't\": \"Were not\", \"They've\": \"They have\", \"Couldn't\": \"Could not\", \"When's\": \"When is\", \"Haven't\": \"Have not\", \"We'll\": \"We will\", \"That's\": \"That is\", \"We're\": \"We are\", \"They're\": \"They' are\", \"You'd\": \"You would\", \"How'd\": \"How did\", \"What're\": \"What are\", \"Hasn't\": \"Has not\", \"Wasn't\": \"Was not\", \"Won't\": \"Will not\", \"There's\": \"There is\", \"Didn't\": \"Did not\", \"Doesn't\": \"Does not\", \"You're\": \"You are\", \"He's\": \"He is\", \"SO's\": \"So is\", \"We've\": \"We have\", \"Who's\": \"Who is\", \"Wouldn't\": \"Would not\", \"Why's\": \"Why is\", \"WHO's\": \"Who is\", \"Let's\": \"Let us\", \"How's\": \"How is\", \"Can't\": \"Can not\", \"Where's\": \"Where is\", \"They'd\": \"They had\", \"Don't\": \"Do not\", \"Shouldn't\":\"Should not\", \"Aren't\":\"Are not\", \"ain't\": \"is not\", \"What's\": \"What is\", \"It's\": \"It is\", \"Isn't\":\"Is not\", \"aren't\": \"are not\",\"can't\": \"cannot\", \"'cause\": \"because\", \"could've\": \"could have\", \"couldn't\": \"could not\", \"didn't\": \"did not\",  \"doesn't\": \"does not\", \"don't\": \"do not\", \"hadn't\": \"had not\", \"hasn't\": \"has not\", \"haven't\": \"have not\", \"he'd\": \"he would\",\"he'll\": \"he will\", \"he's\": \"he is\", \"how'd\": \"how did\", \"how'd'y\": \"how do you\", \"how'll\": \"how will\", \"how's\": \"how is\",  \"I'd\": \"I would\", \"I'd've\": \"I would have\", \"I'll\": \"I will\", \"I'll've\": \"I will have\",\"I'm\": \"I am\", \"I've\": \"I have\", \"i'd\": \"i would\", \"i'd've\": \"i would have\", \"i'll\": \"i will\",  \"i'll've\": \"i will have\",\"i'm\": \"i am\", \"i've\": \"i have\", \"isn't\": \"is not\", \"it'd\": \"it would\", \"it'd've\": \"it would have\", \"it'll\": \"it will\", \"it'll've\": \"it will have\",\"it's\": \"it is\", \"let's\": \"let us\", \"ma'am\": \"madam\", \"mayn't\": \"may not\", \"might've\": \"might have\",\"mightn't\": \"might not\",\"mightn't've\": \"might not have\", \"must've\": \"must have\", \"mustn't\": \"must not\", \"mustn't've\": \"must not have\", \"needn't\": \"need not\", \"needn't've\": \"need not have\",\"o'clock\": \"of the clock\", \"oughtn't\": \"ought not\", \"oughtn't've\": \"ought not have\", \"shan't\": \"shall not\", \"sha'n't\": \"shall not\", \"shan't've\": \"shall not have\", \"she'd\": \"she would\", \"she'd've\": \"she would have\", \"she'll\": \"she will\", \"she'll've\": \"she will have\", \"she's\": \"she is\", \"should've\": \"should have\", \"shouldn't\": \"should not\", \"shouldn't've\": \"should not have\", \"so've\": \"so have\",\"so's\": \"so as\", \"this's\": \"this is\",\"that'd\": \"that would\", \"that'd've\": \"that would have\", \"that's\": \"that is\", \"there'd\": \"there would\", \"there'd've\": \"there would have\", \"there's\": \"there is\", \"here's\": \"here is\",\"they'd\": \"they would\", \"they'd've\": \"they would have\", \"they'll\": \"they will\", \"they'll've\": \"they will have\", \"they're\": \"they are\", \"they've\": \"they have\", \"to've\": \"to have\", \"wasn't\": \"was not\", \"we'd\": \"we would\", \"we'd've\": \"we would have\", \"we'll\": \"we will\", \"we'll've\": \"we will have\", \"we're\": \"we are\", \"we've\": \"we have\", \"weren't\": \"were not\", \"what'll\": \"what will\", \"what'll've\": \"what will have\", \"what're\": \"what are\",  \"what's\": \"what is\", \"what've\": \"what have\", \"when's\": \"when is\", \"when've\": \"when have\", \"where'd\": \"where did\", \"where's\": \"where is\", \"where've\": \"where have\", \"who'll\": \"who will\", \"who'll've\": \"who will have\", \"who's\": \"who is\", \"who've\": \"who have\", \"why's\": \"why is\", \"why've\": \"why have\", \"will've\": \"will have\", \"won't\": \"will not\", \"won't've\": \"will not have\", \"would've\": \"would have\", \"wouldn't\": \"would not\", \"wouldn't've\": \"would not have\", \"y'all\": \"you all\", \"y'all'd\": \"you all would\",\"y'all'd've\": \"you all would have\",\"y'all're\": \"you all are\",\"y'all've\": \"you all have\",\"you'd\": \"you would\", \"you'd've\": \"you would have\", \"you'll\": \"you will\", \"you'll've\": \"you will have\", \"you're\": \"you are\", \"you've\": \"you have\" }\n\ndef clean_contractions(x):\n    specials = [\"’\", \"‘\", \"´\", \"`\"]\n    for s in specials:\n        x = x.replace(s, \"'\")\n    \n    x = ' '.join([contraction_mapping[t] if t in contraction_mapping else t for t in x.split(\" \")])\n    return x","metadata":{"execution":{"iopub.status.busy":"2021-06-10T16:05:54.277870Z","iopub.execute_input":"2021-06-10T16:05:54.278201Z","iopub.status.idle":"2021-06-10T16:05:54.302296Z","shell.execute_reply.started":"2021-06-10T16:05:54.278170Z","shell.execute_reply":"2021-06-10T16:05:54.301058Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# bỏ stopword\ndef remove_stopwords(x):\n  x = [word for word in x.split() if word not in STOPWORDS]\n  x = ' '.join(x)\n  return x","metadata":{"execution":{"iopub.status.busy":"2021-06-10T16:05:54.304011Z","iopub.execute_input":"2021-06-10T16:05:54.304449Z","iopub.status.idle":"2021-06-10T16:05:54.318435Z","shell.execute_reply.started":"2021-06-10T16:05:54.304402Z","shell.execute_reply":"2021-06-10T16:05:54.317422Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Chuyển các từ cùng biến thể của một từ về một từ duy nhất\nfrom nltk.tokenize import word_tokenize\nfrom nltk.stem import WordNetLemmatizer\nl = WordNetLemmatizer()\n\ndef lemmatize_text(x):\n    x = ' '.join([l.lemmatize(word) for word in word_tokenize(x)])\n    return x","metadata":{"execution":{"iopub.status.busy":"2021-06-10T16:05:54.319804Z","iopub.execute_input":"2021-06-10T16:05:54.320148Z","iopub.status.idle":"2021-06-10T16:05:54.334344Z","shell.execute_reply.started":"2021-06-10T16:05:54.320119Z","shell.execute_reply":"2021-06-10T16:05:54.332114Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Loại bỏ ký tự đặc biệt\npuncts = [',', '.', '\"', ':', ')', '(', '-', '!', '?', '|', ';', \"'\", '$', '&', '/', '[', ']', '>', '%', '=', '#', '*', '+', '\\\\', \n        '•', '~', '@', '£', '·', '_', '{', '}', '©', '^', '®', '`', '<', '→', '°', '€', '™', '›', '♥', '←', '×', '§', '″', '′', \n        '█', '…', '“', '★', '”', '–', '●', '►', '−', '¢', '¬', '░', '¡', '¶', '↑', '±', '¿', '▾', '═', '¦', '║', '―', '¥', '▓', \n        '—', '‹', '─', '▒', '：', '⊕', '▼', '▪', '†', '■', '’', '▀', '¨', '▄', '♫', '☆', '¯', '♦', '¤', '▲', '¸', '⋅', '‘', '∞', \n        '∙', '）', '↓', '、', '│', '（', '»', '，', '♪', '╩', '╚', '・', '╦', '╣', '╔', '╗', '▬', '❤', '≤', '‡', '√', '◄', '━', \n        '⇒', '▶', '≥', '╝', '♡', '◊', '。', '✈', '≡', '☺', '✔', '↵', '≈', '✓', '♣', '☎', '℃', '◦', '└', '‟', '～', '！', '○', \n        '◆', '№', '♠', '▌', '✿', '▸', '⁄', '□', '❖', '✦', '．', '÷', '｜', '┃', '／', '￥', '╠', '↩', '✭', '▐', '☼', '☻', '┐', \n        '├', '«', '∼', '┌', '℉', '☮', '฿', '≦', '♬', '✧', '〉', '－', '⌂', '✖', '･', '◕', '※', '‖', '◀', '‰', '\\x97', '↺', \n        '∆', '┘', '┬', '╬', '،', '⌘', '⊂', '＞', '〈', '⎙', '？', '☠', '⇐', '▫', '∗', '∈', '≠', '♀', '♔', '˚', '℗', '┗', '＊', \n        '┼', '❀', '＆', '∩', '♂', '‿', '∑', '‣', '➜', '┛', '⇓', '☯', '⊖', '☀', '┳', '；', '∇', '⇑', '✰', '◇', '♯', '☞', '´', \n        '↔', '┏', '｡', '◘', '∂', '✌', '♭', '┣', '┴', '┓', '✨', '\\xa0', '˜', '❥', '┫', '℠', '✒', '［', '∫', '\\x93', '≧', '］', \n        '\\x94', '∀', '♛', '\\x96', '∨', '◎', '↻', '⇩', '＜', '≫', '✩', '✪', '♕', '؟', '₤', '☛', '╮', '␊', '＋', '┈', '％', \n        '╋', '▽', '⇨', '┻', '⊗', '￡', '।', '▂', '✯', '▇', '＿', '➤', '✞', '＝', '▷', '△', '◙', '▅', '✝', '∧', '␉', '☭', \n        '┊', '╯', '☾', '➔', '∴', '\\x92', '▃', '↳', '＾', '׳', '➢', '╭', '➡', '＠', '⊙', '☢', '˝', '∏', '„', '∥', '❝', '☐', \n        '▆', '╱', '⋙', '๏', '☁', '⇔', '▔', '\\x91', '➚', '◡', '╰', '\\x85', '♢', '˙', '۞', '✘', '✮', '☑', '⋆', 'ⓘ', '❒', \n        '☣', '✉', '⌊', '➠', '∣', '❑', '◢', 'ⓒ', '\\x80', '〒', '∕', '▮', '⦿', '✫', '✚', '⋯', '♩', '☂', '❞', '‗', '܂', '☜', \n        '‾', '✜', '╲', '∘', '⟩', '＼', '⟨', '·', '✗', '♚', '∅', 'ⓔ', '◣', '͡', '‛', '❦', '◠', '✄', '❄', '∃', '␣', '≪', '｢', \n        '≅', '◯', '☽', '∎', '｣', '❧', '̅', 'ⓐ', '↘', '⚓', '▣', '˘', '∪', '⇢', '✍', '⊥', '＃', '⎯', '↠', '۩', '☰', '◥', \n        '⊆', '✽', '⚡', '↪', '❁', '☹', '◼', '☃', '◤', '❏', 'ⓢ', '⊱', '➝', '̣', '✡', '∠', '｀', '▴', '┤', '∝', '♏', 'ⓐ', \n        '✎', ';', '␤', '＇', '❣', '✂', '✤', 'ⓞ', '☪', '✴', '⌒', '˛', '♒', '＄', '✶', '▻', 'ⓔ', '◌', '◈', '❚', '❂', '￦', \n        '◉', '╜', '̃', '✱', '╖', '❉', 'ⓡ', '↗', 'ⓣ', '♻', '➽', '׀', '✲', '✬', '☉', '▉', '≒', '☥', '⌐', '♨', '✕', 'ⓝ', \n        '⊰', '❘', '＂', '⇧', '̵', '➪', '▁', '▏', '⊃', 'ⓛ', '‚', '♰', '́', '✏', '⏑', '̶', 'ⓢ', '⩾', '￠', '❍', '≃', '⋰', '♋', \n        '､', '̂', '❋', '✳', 'ⓤ', '╤', '▕', '⌣', '✸', '℮', '⁺', '▨', '╨', 'ⓥ', '♈', '❃', '☝', '✻', '⊇', '≻', '♘', '♞', \n        '◂', '✟', '⌠', '✠', '☚', '✥', '❊', 'ⓒ', '⌈', '❅', 'ⓡ', '♧', 'ⓞ', '▭', '❱', 'ⓣ', '∟', '☕', '♺', '∵', '⍝', 'ⓑ', \n        '✵', '✣', '٭', '♆', 'ⓘ', '∶', '⚜', '◞', '்', '✹', '➥', '↕', '̳', '∷', '✋', '➧', '∋', '̿', 'ͧ', '┅', '⥤', '⬆', '⋱', \n        '☄', '↖', '⋮', '۔', '♌', 'ⓛ', '╕', '♓', '❯', '♍', '▋', '✺', '⭐', '✾', '♊', '➣', '▿', 'ⓑ', '♉', '⏠', '◾', '▹', \n        '⩽', '↦', '╥', '⍵', '⌋', '։', '➨', '∮', '⇥', 'ⓗ', 'ⓓ', '⁻', '⎝', '⌥', '⌉', '◔', '◑', '✼', '♎', '♐', '╪', '⊚', \n        '☒', '⇤', 'ⓜ', '⎠', '◐', '⚠', '╞', '◗', '⎕', 'ⓨ', '☟', 'ⓟ', '♟', '❈', '↬', 'ⓓ', '◻', '♮', '❙', '♤', '∉', '؛', \n        '⁂', 'ⓝ', '־', '♑', '╫', '╓', '╳', '⬅', '☔', '☸', '┄', '╧', '׃', '⎢', '❆', '⋄', '⚫', '̏', '☏', '➞', '͂', '␙', \n        'ⓤ', '◟', '̊', '⚐', '✙', '↙', '̾', '℘', '✷', '⍺', '❌', '⊢', '▵', '✅', 'ⓖ', '☨', '▰', '╡', 'ⓜ', '☤', '∽', '╘', \n        '˹', '↨', '♙', '⬇', '♱', '⌡', '⠀', '╛', '❕', '┉', 'ⓟ', '̀', '♖', 'ⓚ', '┆', '⎜', '◜', '⚾', '⤴', '✇', '╟', '⎛', \n        '☩', '➲', '➟', 'ⓥ', 'ⓗ', '⏝', '◃', '╢', '↯', '✆', '˃', '⍴', '❇', '⚽', '╒', '̸', '♜', '☓', '➳', '⇄', '☬', '⚑', \n        '✐', '⌃', '◅', '▢', '❐', '∊', '☈', '॥', '⎮', '▩', 'ு', '⊹', '‵', '␔', '☊', '➸', '̌', '☿', '⇉', '⊳', '╙', 'ⓦ', \n        '⇣', '｛', '̄', '↝', '⎟', '▍', '❗', '״', '΄', '▞', '◁', '⛄', '⇝', '⎪', '♁', '⇠', '☇', '✊', 'ி', '｝', '⭕', '➘', \n        '⁀', '☙', '❛', '❓', '⟲', '⇀', '≲', 'ⓕ', '⎥', '\\u06dd', 'ͤ', '₋', '̱', '̎', '♝', '≳', '▙', '➭', '܀', 'ⓖ', '⇛', '▊', \n        '⇗', '̷', '⇱', '℅', 'ⓧ', '⚛', '̐', '̕', '⇌', '␀', '≌', 'ⓦ', '⊤', '̓', '☦', 'ⓕ', '▜', '➙', 'ⓨ', '⌨', '◮', '☷', \n        '◍', 'ⓚ', '≔', '⏩', '⍳', '℞', '┋', '˻', '▚', '≺', 'ْ', '▟', '➻', '̪', '⏪', '̉', '⎞', '┇', '⍟', '⇪', '▎', '⇦', '␝', \n        '⤷', '≖', '⟶', '♗', '̴', '♄', 'ͨ', '̈', '❜', '̡', '▛', '✁', '➩', 'ா', '˂', '↥', '⏎', '⎷', '̲', '➖', '↲', '⩵', '̗', '❢', \n        '≎', '⚔', '⇇', '̑', '⊿', '̖', '☍', '➹', '⥊', '⁁', '✢']\n\ndef clean_punct(x):\n  for punct in puncts:\n    if punct in x:\n      x = x.replace(punct, f' {punct} ')\n  return x","metadata":{"execution":{"iopub.status.busy":"2021-06-10T16:05:54.336992Z","iopub.execute_input":"2021-06-10T16:05:54.337415Z","iopub.status.idle":"2021-06-10T16:05:54.386544Z","shell.execute_reply.started":"2021-06-10T16:05:54.337371Z","shell.execute_reply":"2021-06-10T16:05:54.385122Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# bỏ chữ số\ndef clean_numbers(x):\n    x = re.sub(r'(\\d+)([a-zA-Z])', '\\g<1> \\g<2>', x)\n    x = re.sub(r'(\\d+) (th|st|nd|rd) ', '\\g<1>\\g<2> ', x)\n    x = re.sub(r'(\\d+),(\\d+)', '\\g<1>\\g<2>', x)\n    return x","metadata":{"execution":{"iopub.status.busy":"2021-06-10T16:05:54.389146Z","iopub.execute_input":"2021-06-10T16:05:54.389837Z","iopub.status.idle":"2021-06-10T16:05:54.402979Z","shell.execute_reply.started":"2021-06-10T16:05:54.389784Z","shell.execute_reply":"2021-06-10T16:05:54.401808Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# thay các tag latex\ndef clean_latex_tag(x):\n    corr_t = []\n    for t in x.split(\" \"):\n        t = t.strip()\n        if t != '':\n            corr_t.append(t)\n    x = ' '.join(corr_t)\n    x = re.sub('(\\[ math \\]).+(\\[ / math \\])', 'math formula', x)\n    return x","metadata":{"execution":{"iopub.status.busy":"2021-06-10T16:05:54.404450Z","iopub.execute_input":"2021-06-10T16:05:54.404872Z","iopub.status.idle":"2021-06-10T16:05:54.417833Z","shell.execute_reply.started":"2021-06-10T16:05:54.404828Z","shell.execute_reply":"2021-06-10T16:05:54.415756Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#gộp các hàm xử lý lại\ndef data_cleaning(x):\n    x = clean_tag(x)\n    x = clean_contractions(x)\n    x = clean_punct(x)\n    x = lemmatize_text(x)\n    x = clean_latex_tag(x)\n    x = clean_numbers(x)\n    x = remove_stopwords(x)\n    return x","metadata":{"execution":{"iopub.status.busy":"2021-06-10T16:05:54.420785Z","iopub.execute_input":"2021-06-10T16:05:54.423328Z","iopub.status.idle":"2021-06-10T16:05:54.435241Z","shell.execute_reply.started":"2021-06-10T16:05:54.423267Z","shell.execute_reply":"2021-06-10T16:05:54.434023Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#xử lý dữ liệu câu hỏi trên tập train và validation\nx['question_text'] = x['question_text'].progress_map(lambda x: data_cleaning(x))\nvalidation_data['question_text']=validation_data['question_text'].progress_map(lambda x: data_cleaning(x))","metadata":{"execution":{"iopub.status.busy":"2021-06-10T16:05:54.438855Z","iopub.execute_input":"2021-06-10T16:05:54.439292Z","iopub.status.idle":"2021-06-10T16:10:55.690030Z","shell.execute_reply.started":"2021-06-10T16:05:54.439254Z","shell.execute_reply":"2021-06-10T16:10:55.689060Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 3. Đưa dữ liệu vào mô hình phân lớp\n- Từ bước phân tích dữ liệu, ta nhận thấy các câu hỏi insincere thường chứa các từ ngữ mang nghĩa xấu, không phụ thuộc vào ngữ pháp nên chọn hướng tiếp cận sử dụng CountVectorizer là vector đếm số lượng từ.\n- Bước xử lý dữ liệu cũng đã thu nhỏ được tập từ vựng khi sử dụng CountVectorizer qua việc loại bỏ dữ liệu biên và xử lý dữ liệu văn bản.","metadata":{}},{"cell_type":"markdown","source":"## Chia tập test, train","metadata":{}},{"cell_type":"code","source":"x_train, x_test, y_train, y_test = train_test_split(\n    x['question_text'], x['target'], test_size=0.2, random_state=0)\nprint('x_train: ', x_train.shape, y_train.shape)\nprint('x_test: ',x_test.shape, y_test.shape)","metadata":{"scrolled":true,"_uuid":"3dbc763f4a123450ddb0c1c467be33f04625dfce","execution":{"iopub.status.busy":"2021-06-10T16:10:55.691768Z","iopub.execute_input":"2021-06-10T16:10:55.692088Z","iopub.status.idle":"2021-06-10T16:10:55.779870Z","shell.execute_reply.started":"2021-06-10T16:10:55.692059Z","shell.execute_reply":"2021-06-10T16:10:55.778691Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Tạo ra các vector đếm từ tập train, test\n- Vector đếm từ: chuyển câu thành vector chứa các từ và số lần xuất hiện của từ đó trong câu\n- Phân lớp câu hỏi không phụ thuộc vào ngữ pháp\n- Học trên tập từ vựng của toàn bộ tập train và test do ở bước test, vector đếm có thể phải mã hoá những từ có ở tập test mà không nằm trong tập train","metadata":{}},{"cell_type":"code","source":"vectorizer = CountVectorizer()\n# Học trên tập từ vựng của toàn bộ tập train và validation\nvectorizer.fit(list(x['question_text'].values)+ list(validation_data['question_text'].values))\n# Tạo vector đếm cho tập train, test, validation dựa trên tập từ vựng đã học\nx_tr = vectorizer.transform(x_train) \nx_te = vectorizer.transform(x_test)\nx_val = vectorizer.transform(validation_data['question_text'])\nprint(x_tr.shape)\nprint(x_te.shape)\nprint(x_val.shape)","metadata":{"_uuid":"bc1feb442a5d439db3b20f2f7eda3ab4cb620d27","execution":{"iopub.status.busy":"2021-06-10T16:10:55.781147Z","iopub.execute_input":"2021-06-10T16:10:55.781436Z","iopub.status.idle":"2021-06-10T16:11:18.297373Z","shell.execute_reply.started":"2021-06-10T16:10:55.781407Z","shell.execute_reply":"2021-06-10T16:11:18.296602Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Sử dụng Logistic Regression để phân lớp câu hỏi từ các vector đếm\n- Sau khi có được vector đếm, chỉ cần đưa vector vào model để train\n- Sử dụng Logistic Regression vì model đơn giản, hiệu quả hơn so với các model khác: Random Forest Classifier, Naive Bayes, SVM","metadata":{}},{"cell_type":"code","source":"# import\nfrom sklearn.metrics import accuracy_score, f1_score, classification_report\nfrom sklearn.linear_model import LogisticRegression","metadata":{"execution":{"iopub.status.busy":"2021-06-10T16:11:18.298492Z","iopub.execute_input":"2021-06-10T16:11:18.298934Z","iopub.status.idle":"2021-06-10T16:11:18.302502Z","shell.execute_reply.started":"2021-06-10T16:11:18.298889Z","shell.execute_reply":"2021-06-10T16:11:18.301666Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Tối ưu tham số bằng GridSearchCV\n- Đầu vào của model có tham số C, sử dụng GridSearchCV nhằm tìm ra tham số C tốt nhất cho model","metadata":{}},{"cell_type":"code","source":"# Tìm kiếm tham số \nfrom sklearn.model_selection import  GridSearchCV\nparams = {'C': [0.1, 1, 2, 3, 4, 5, 10]}\n\ngridsearch = GridSearchCV(LogisticRegression(), params, scoring='f1', n_jobs=-1, verbose=1)\ngridsearch.fit(x_tr, y_train)\nprint(gridsearch.best_params_)","metadata":{"execution":{"iopub.status.busy":"2021-06-10T16:11:18.303608Z","iopub.execute_input":"2021-06-10T16:11:18.304076Z","iopub.status.idle":"2021-06-10T16:12:53.250343Z","shell.execute_reply.started":"2021-06-10T16:11:18.304045Z","shell.execute_reply":"2021-06-10T16:12:53.247922Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Chạy mô hình trên tham số đã tìm được","metadata":{}},{"cell_type":"code","source":"%%time\nmodel = LogisticRegression(C=gridsearch.best_params_['C'])\n## Chạy mô hình\nprint(f\"Running Logistic Regression\")\nmodel.fit(x_tr, y_train)\n\ntrain_predictions = model.predict(x_tr)\ntrain_acc = accuracy_score(y_train, train_predictions)\ntrain_f1 = f1_score(y_train, train_predictions)\nprint(f\"Train accuracy: {train_acc:.2%}, F1: {train_f1:.4f}\") \ntest_predictions = model.predict(x_te)\ntest_acc = accuracy_score(y_test, test_predictions) \ntest_f1 = f1_score(y_test, test_predictions) \nprint(f\"Test accuracy:  {test_acc:.2%}, F1: {test_f1:.4f}\")","metadata":{"execution":{"iopub.status.busy":"2021-06-10T16:28:20.870135Z","iopub.execute_input":"2021-06-10T16:28:20.871514Z","iopub.status.idle":"2021-06-10T16:28:34.103324Z","shell.execute_reply.started":"2021-06-10T16:28:20.871453Z","shell.execute_reply":"2021-06-10T16:28:34.102228Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# In ra confusion matrix, classification_report\nimport seaborn as sns\nfrom sklearn.metrics import confusion_matrix\nsns.set(font_scale=1.4)\nsns.heatmap(pd.DataFrame(confusion_matrix(y_test, test_predictions), range(2),range(2)), annot=True, fmt='g')\nprint(classification_report(y_test, test_predictions))","metadata":{"execution":{"iopub.status.busy":"2021-06-10T16:13:05.394492Z","iopub.execute_input":"2021-06-10T16:13:05.394740Z","iopub.status.idle":"2021-06-10T16:13:06.086515Z","shell.execute_reply.started":"2021-06-10T16:13:05.394715Z","shell.execute_reply":"2021-06-10T16:13:06.085438Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Nhận xét:\n- F1 và accuracy đã khá gần nhau sau khi cân bằng dataset\n- F1 của model khi dự đoán câu hỏi insincere tăng lên đáng kể sau khi cân bằng dataset\n- Accuracy của model khi dự đoán câu hỏi sincere giảm nhưng không đáng kể sau khi cân bằng dataset\n- Không bị overfit trên câu hỏi sincere","metadata":{}},{"cell_type":"markdown","source":"## 4. Submission\n- Sau khi train xong, chạy mô hình trên tập validation và lưu lại để submit","metadata":{}},{"cell_type":"code","source":"# Submission\n\nvalidation_predictions = model.predict(x_val)\nsubmission = pd.DataFrame({'qid':validation_data['qid'], 'prediction':validation_predictions })\nsubmission.to_csv('submission.csv', index=False)\nsubmission","metadata":{"execution":{"iopub.status.busy":"2021-06-10T16:13:06.088215Z","iopub.execute_input":"2021-06-10T16:13:06.088643Z","iopub.status.idle":"2021-06-10T16:13:06.988984Z","shell.execute_reply.started":"2021-06-10T16:13:06.088600Z","shell.execute_reply":"2021-06-10T16:13:06.987991Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Nhận xét:\n- Kết quả F1 trên tập validation: ~0.6\n- Điểm F1 tập validation lệch nhiều so với F1 tập test và tập train nhưng tương đối cao với mô hình tuyến tính","metadata":{}}]}