{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Tổng quan vấn đề\n\n* Một vấn đề tồn tại với bất kì trang web lớn nào hiện nay là làm thế nào xử lý nội dung độc hại. Quora muốn giải quyết vấn đề này trực tiếp để giữ cho nền tảng của họ trở thành một nơi mà người dùng có thể cảm thấy an toàn khi chia sẻ kiến thức của họ với thế giới.\n* Quora là một nền tảng cho phép mọi người học hỏi lẫn nhau. Trên Quora, mọi người có thể đặt câu hỏi và kết nối với những người khác, những người đóng góp thông tin chi tiết độc đáo và câu trả lời chất lượng. Một thách thức quan trọng là loại bỏ những câu hỏi insincere - những câu hỏi được đặt ra dựa trên những tiền đề sai lầm hoặc có ý định đưa ra một tuyên bố hơn là tìm kiếm câu trả lời hữu ích.\n* Mục tiêu: loại bỏ những câu hỏi insincere\n","metadata":{}},{"cell_type":"code","source":"import numpy as np \nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport re\nimport os\nimport math","metadata":{"execution":{"iopub.status.busy":"2021-05-27T02:45:25.821920Z","iopub.execute_input":"2021-05-27T02:45:25.822360Z","iopub.status.idle":"2021-05-27T02:45:26.758052Z","shell.execute_reply.started":"2021-05-27T02:45:25.822248Z","shell.execute_reply":"2021-05-27T02:45:26.757047Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train = pd.read_csv('../input/quora-insincere-questions-classification/train.csv')\ntest = pd.read_csv('../input/quora-insincere-questions-classification/test.csv')","metadata":{"execution":{"iopub.status.busy":"2021-05-27T02:45:26.761159Z","iopub.execute_input":"2021-05-27T02:45:26.761604Z","iopub.status.idle":"2021-05-27T02:45:32.749698Z","shell.execute_reply.started":"2021-05-27T02:45:26.761558Z","shell.execute_reply":"2021-05-27T02:45:32.748717Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.info()","metadata":{"execution":{"iopub.status.busy":"2021-05-27T02:45:32.751689Z","iopub.execute_input":"2021-05-27T02:45:32.752009Z","iopub.status.idle":"2021-05-27T02:45:33.022294Z","shell.execute_reply.started":"2021-05-27T02:45:32.751977Z","shell.execute_reply":"2021-05-27T02:45:33.021269Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"+ Tập dữ liệu train gôm 13 triệu dòng và 3 cột\n+ Các trường dữ liệu:\n * qid: mã định danh\n * question_text: các câu hỏi trên quora\n * target: câu hỏi sincere khi có giá trị 0 và insincere khi có giá trị 1","metadata":{}},{"cell_type":"code","source":"train.head(10)","metadata":{"execution":{"iopub.status.busy":"2021-05-27T02:45:33.024115Z","iopub.execute_input":"2021-05-27T02:45:33.024709Z","iopub.status.idle":"2021-05-27T02:45:33.047305Z","shell.execute_reply.started":"2021-05-27T02:45:33.024664Z","shell.execute_reply":"2021-05-27T02:45:33.046194Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train[\"target\"].value_counts()","metadata":{"execution":{"iopub.status.busy":"2021-05-27T02:45:33.048846Z","iopub.execute_input":"2021-05-27T02:45:33.049483Z","iopub.status.idle":"2021-05-27T02:45:33.072973Z","shell.execute_reply.started":"2021-05-27T02:45:33.049423Z","shell.execute_reply":"2021-05-27T02:45:33.071923Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Tập dữ liệu train gồm 80810 dòng đã được xác nhận là câu hỏi insincere và 1225312 dòng là sincere","metadata":{}},{"cell_type":"code","source":"test.info()","metadata":{"execution":{"iopub.status.busy":"2021-05-27T02:45:33.074486Z","iopub.execute_input":"2021-05-27T02:45:33.075080Z","iopub.status.idle":"2021-05-27T02:45:33.159330Z","shell.execute_reply.started":"2021-05-27T02:45:33.075037Z","shell.execute_reply":"2021-05-27T02:45:33.158154Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Tập dữ liệu test gồm hơn 300 nghìn dòng","metadata":{}},{"cell_type":"code","source":"test.head(10)","metadata":{"execution":{"iopub.status.busy":"2021-05-27T02:45:33.161020Z","iopub.execute_input":"2021-05-27T02:45:33.161703Z","iopub.status.idle":"2021-05-27T02:45:33.173337Z","shell.execute_reply.started":"2021-05-27T02:45:33.161657Z","shell.execute_reply":"2021-05-27T02:45:33.172431Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = plt.figure(figsize=(5,3))\nax = fig.add_axes([0,0,2,1])\ntype_check = ['sincere','insincere']\ncount = [1225312,80810]\nax.bar(type_check,count)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2021-05-27T02:45:33.176398Z","iopub.execute_input":"2021-05-27T02:45:33.176757Z","iopub.status.idle":"2021-05-27T02:45:33.331000Z","shell.execute_reply.started":"2021-05-27T02:45:33.176725Z","shell.execute_reply":"2021-05-27T02:45:33.329912Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('- Phần trăm câu hỏi sincere (target = 0): {}%'.format(100 - round(train['target'].mean() * 100, 2))) \nprint('- Phần trăm câu hỏi insincere (target = 1): {}%'.format(round(train['target'].mean() * 100, 2)))","metadata":{"execution":{"iopub.status.busy":"2021-05-27T02:45:33.333442Z","iopub.execute_input":"2021-05-27T02:45:33.333895Z","iopub.status.idle":"2021-05-27T02:45:33.344711Z","shell.execute_reply.started":"2021-05-27T02:45:33.333842Z","shell.execute_reply":"2021-05-27T02:45:33.343770Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* Câu hỏi sincere chiếm 93,81% trong khi câu hỏi insincere chỉ chiếm 6,19% trên tập dữ liệu train\n* Như vậy, tập dữ liệu của chúng ta rất không cân bằng\n* Giải pháp: làm cho mô hình cân bằng hơn sao cho không ảnh hưởng đáng kể đến khả năng dự báo của mô hình","metadata":{}},{"cell_type":"code","source":"# chia lại tập dữ liệu theo tỉ lệ 4:1\nfrom sklearn.utils import resample\nsincere = train[train.target == 0]\ninsincere = train[train.target == 1]\ntrain_sample = pd.concat([resample(sincere,replace = True,n_samples = len(insincere)*4), \ninsincere])\ntrain_sample","metadata":{"execution":{"iopub.status.busy":"2021-05-27T02:45:33.345945Z","iopub.execute_input":"2021-05-27T02:45:33.346230Z","iopub.status.idle":"2021-05-27T02:45:33.825959Z","shell.execute_reply.started":"2021-05-27T02:45:33.346202Z","shell.execute_reply":"2021-05-27T02:45:33.825292Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Xử lý dữ liệu","metadata":{}},{"cell_type":"code","source":"puncts = [',', '.', '\"', ':', ')', '(', '-', '!', '?', '|', ';', \"'\", '$', '&', '/', '[',\n          ']', '>', '%', '=', '#', '*', '+', '\\\\', '•',  '~', '@', '£', '·', '_', '{', '}',\n          '©', '^', '®', '`',  '<', '→', '°', '€', '™', '›',  '♥', '←', '×', '§', '″', '′',\n          'Â', '█', '½', 'à', '…', '“', '★', '”', '–', '●', 'â', '►', '−', '¢', '²', '¬',\n          '░', '¶', '↑', '±', '¿', '▾', '═', '¦', '║', '―', '¥', '▓', '—', '‹', '─', '▒',\n          '：', '¼', '⊕', '▼', '▪', '†', '■', '’', '▀', '¨', '▄', '♫', '☆', 'é', '¯', '♦',\n          '¤', '▲', 'è', '¸', '¾', 'Ã', '⋅', '‘', '∞', '∙', '）', '↓', '、', '│', '（', '»',\n          '，', '♪', '╩', '╚', '³', '・', '╦', '╣', '╔', '╗', '▬', '❤', 'ï', 'Ø', '¹', '≤',\n          '‡', '√' ]","metadata":{"execution":{"iopub.status.busy":"2021-05-27T02:45:33.827080Z","iopub.execute_input":"2021-05-27T02:45:33.827543Z","iopub.status.idle":"2021-05-27T02:45:33.836749Z","shell.execute_reply.started":"2021-05-27T02:45:33.827496Z","shell.execute_reply":"2021-05-27T02:45:33.835779Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Loại bỏ các kí tự đặc biệt\ndef clean_text(x):\n    x = str(x)\n    for punct in puncts:\n        x = x.replace(punct, f' {punct} ')\n    return x","metadata":{"execution":{"iopub.status.busy":"2021-05-27T02:45:33.838032Z","iopub.execute_input":"2021-05-27T02:45:33.838323Z","iopub.status.idle":"2021-05-27T02:45:33.854579Z","shell.execute_reply.started":"2021-05-27T02:45:33.838289Z","shell.execute_reply":"2021-05-27T02:45:33.853573Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Loại bỏ số\ndef clean_numbers(x):\n    x = re.sub('[0-9]{5,}', '#####', x)\n    x = re.sub('[0-9]{4}', '####', x)\n    x = re.sub('[0-9]{3}', '###', x)\n    x = re.sub('[0-9]{2}', '##', x)\n    return x","metadata":{"execution":{"iopub.status.busy":"2021-05-27T02:45:33.856180Z","iopub.execute_input":"2021-05-27T02:45:33.856521Z","iopub.status.idle":"2021-05-27T02:45:33.874762Z","shell.execute_reply.started":"2021-05-27T02:45:33.856477Z","shell.execute_reply":"2021-05-27T02:45:33.873931Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mispell_dict = {\"aren't\" : \"are not\",\n\"can't\" : \"cannot\",\n\"couldn't\" : \"could not\",\n\"didn't\" : \"did not\",\n\"doesn't\" : \"does not\",\n\"don't\" : \"do not\",\n\"hadn't\" : \"had not\",\n\"hasn't\" : \"has not\",\n\"haven't\" : \"have not\",\n\"he'd\" : \"he would\",\n\"he'll\" : \"he will\",\n\"he's\" : \"he is\",\n\"i'd\" : \"I would\",\n\"i'd\" : \"I had\",\n\"i'll\" : \"I will\",\n\"i'm\" : \"I am\",\n\"isn't\" : \"is not\",\n\"it's\" : \"it is\",\n\"it'll\":\"it will\",\n\"i've\" : \"I have\",\n\"let's\" : \"let us\",\n\"mightn't\" : \"might not\",\n\"mustn't\" : \"must not\",\n\"shan't\" : \"shall not\",\n\"she'd\" : \"she would\",\n\"she'll\" : \"she will\",\n\"she's\" : \"she is\",\n\"shouldn't\" : \"should not\",\n\"that's\" : \"that is\",\n\"there's\" : \"there is\",\n\"they'd\" : \"they would\",\n\"they'll\" : \"they will\",\n\"they're\" : \"they are\",\n\"they've\" : \"they have\",\n\"we'd\" : \"we would\",\n\"we're\" : \"we are\",\n\"weren't\" : \"were not\",\n\"we've\" : \"we have\",\n\"what'll\" : \"what will\",\n\"what're\" : \"what are\",\n\"what's\" : \"what is\",\n\"what've\" : \"what have\",\n\"where's\" : \"where is\",\n\"who'd\" : \"who would\",\n\"who'll\" : \"who will\",\n\"who're\" : \"who are\",\n\"who's\" : \"who is\",\n\"who've\" : \"who have\",\n\"won't\" : \"will not\",\n\"wouldn't\" : \"would not\",\n\"you'd\" : \"you would\",\n\"you'll\" : \"you will\",\n\"you're\" : \"you are\",\n\"you've\" : \"you have\",\n\"'re\": \" are\",\n\"wasn't\": \"was not\",\n\"we'll\":\" will\",\n\"didn't\": \"did not\",\n\"tryin'\":\"trying\"}","metadata":{"execution":{"iopub.status.busy":"2021-05-27T02:45:33.875841Z","iopub.execute_input":"2021-05-27T02:45:33.876234Z","iopub.status.idle":"2021-05-27T02:45:33.889140Z","shell.execute_reply.started":"2021-05-27T02:45:33.876204Z","shell.execute_reply":"2021-05-27T02:45:33.888016Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Thay thế các từ viết tắt\ndef _get_mispell(mispell_dict):\n    mispell_re = re.compile('(%s)' % '|'.join(mispell_dict.keys()))\n    return mispell_dict, mispell_re\n\nmispellings, mispellings_re = _get_mispell(mispell_dict)\ndef replace_typical_misspell(text):\n    def replace(match):\n        return mispellings[match.group(0)]\n    return mispellings_re.sub(replace, text)","metadata":{"execution":{"iopub.status.busy":"2021-05-27T02:45:33.890345Z","iopub.execute_input":"2021-05-27T02:45:33.890663Z","iopub.status.idle":"2021-05-27T02:45:33.906423Z","shell.execute_reply.started":"2021-05-27T02:45:33.890635Z","shell.execute_reply":"2021-05-27T02:45:33.905231Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Loại bỏ các từ trong stopword\nimport nltk\nfrom nltk.corpus import stopwords\nstopword_list = nltk.corpus.stopwords.words('english')\ndef remove_stopwords(text, is_lower_case=True):\n    tokenizer = ToktokTokenizer()\n    tokens = tokenizer.tokenize(text)\n    tokens = [token.strip() for token in tokens]\n    if is_lower_case:\n        filtered_tokens = [token for token in tokens if token not in stopword_list]\n    else:\n        filtered_tokens = [token for token in tokens if token.lower() not in stopword_list]\n    filtered_text = ' '.join(filtered_tokens)\n    return filtered_text","metadata":{"execution":{"iopub.status.busy":"2021-05-27T02:45:33.907800Z","iopub.execute_input":"2021-05-27T02:45:33.908115Z","iopub.status.idle":"2021-05-27T02:45:34.800917Z","shell.execute_reply.started":"2021-05-27T02:45:33.908086Z","shell.execute_reply":"2021-05-27T02:45:34.800094Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Stemming\nfrom nltk.stem import SnowballStemmer\nfrom nltk.tokenize.toktok import ToktokTokenizer\ndef stem_text(text):\n    tokenizer = ToktokTokenizer()\n    stemmer = SnowballStemmer('english')\n    tokens = tokenizer.tokenize(text)\n    tokens = [token.strip() for token in tokens]\n    tokens = [stemmer.stem(token) for token in tokens]\n    return ' '.join(tokens)","metadata":{"execution":{"iopub.status.busy":"2021-05-27T02:45:34.802422Z","iopub.execute_input":"2021-05-27T02:45:34.802887Z","iopub.status.idle":"2021-05-27T02:45:34.809967Z","shell.execute_reply.started":"2021-05-27T02:45:34.802840Z","shell.execute_reply":"2021-05-27T02:45:34.808821Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Stemming là kỹ thuật dùng để biến đổi 1 từ về dạng gốc (được gọi là stem hoặc root form) bằng cách cực kỳ đơn giản là loại bỏ 1 số ký tự nằm ở cuối từ mà nó nghĩ rằng là biến thể của từ","metadata":{}},{"cell_type":"code","source":"# Lemmatization\nfrom nltk.stem import WordNetLemmatizer\nfrom nltk.tokenize.toktok import ToktokTokenizer\nwordnet_lemmatizer = WordNetLemmatizer()\ndef lemma_text(text):\n    tokenizer = ToktokTokenizer()\n    tokens = tokenizer.tokenize(text)\n    tokens = [token.strip() for token in tokens]\n    tokens = [wordnet_lemmatizer.lemmatize(token) for token in tokens]\n    return ' '.join(tokens)","metadata":{"execution":{"iopub.status.busy":"2021-05-27T02:45:34.811528Z","iopub.execute_input":"2021-05-27T02:45:34.811954Z","iopub.status.idle":"2021-05-27T02:45:34.826110Z","shell.execute_reply.started":"2021-05-27T02:45:34.811914Z","shell.execute_reply":"2021-05-27T02:45:34.825169Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Lemmatization khác với Stemming là xử lý bằng cách loại bỏ các ký tự cuối từ một cách rất heuristic, Lemmatization sẽ xử lý thông minh hơn bằng một bộ từ điển hoặc một bộ ontology nào đó","metadata":{}},{"cell_type":"code","source":"def clean_sentence(x):\n    x = x.lower()\n    x = clean_text(x)\n    x = clean_numbers(x)\n    x = replace_typical_misspell(x)\n    x = remove_stopwords(x)\n    x = stem_text(x)\n    x = lemma_text(x)\n    x = x.replace(\"'\",\"\")\n    return x","metadata":{"execution":{"iopub.status.busy":"2021-05-27T02:45:34.827567Z","iopub.execute_input":"2021-05-27T02:45:34.827994Z","iopub.status.idle":"2021-05-27T02:45:34.840389Z","shell.execute_reply.started":"2021-05-27T02:45:34.827950Z","shell.execute_reply":"2021-05-27T02:45:34.839311Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# word cloud cho câu hỏi sincere\nfrom wordcloud import WordCloud, STOPWORDS\nstop_words = set(STOPWORDS)\nsincere_wordcloud = WordCloud().generate(str(train[train[\"target\"] == 0][\"question_text\"]))\nplt.figure(figsize=(8,7))\nplt.imshow(sincere_wordcloud)\nplt.axis(\"off\")\nplt.tight_layout(pad=0)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2021-05-27T02:45:34.842276Z","iopub.execute_input":"2021-05-27T02:45:34.842642Z","iopub.status.idle":"2021-05-27T02:45:35.310739Z","shell.execute_reply.started":"2021-05-27T02:45:34.842611Z","shell.execute_reply":"2021-05-27T02:45:35.309604Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# word cloud cho câu hỏi insincere\ninsincere_wordcloud = WordCloud().generate(str(train[train[\"target\"] == 1][\"question_text\"]))\nplt.figure(figsize=(8,7))\nplt.imshow(insincere_wordcloud)\nplt.axis(\"off\")\nplt.tight_layout(pad=0)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2021-05-27T02:45:35.312051Z","iopub.execute_input":"2021-05-27T02:45:35.312487Z","iopub.status.idle":"2021-05-27T02:45:35.611300Z","shell.execute_reply.started":"2021-05-27T02:45:35.312426Z","shell.execute_reply":"2021-05-27T02:45:35.610308Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Xử lí dữ liệu trên cả tập train và tập test \ntrain_sample['question_text'] = train_sample['question_text'].apply(lambda x: clean_sentence(x))\ntest['question_text'] = test['question_text'].apply(lambda x: clean_sentence(x))","metadata":{"execution":{"iopub.status.busy":"2021-05-27T02:45:35.612959Z","iopub.execute_input":"2021-05-27T02:45:35.613661Z","iopub.status.idle":"2021-05-27T02:51:44.885759Z","shell.execute_reply.started":"2021-05-27T02:45:35.613618Z","shell.execute_reply":"2021-05-27T02:51:44.884725Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.metrics import confusion_matrix\nfrom sklearn.metrics import accuracy_score, log_loss\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.feature_extraction.text import CountVectorizer\nfrom sklearn.feature_extraction.text import TfidfVectorizer\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.metrics import normalized_mutual_info_score\nfrom sklearn.model_selection import cross_val_score\nfrom sklearn import model_selection\nfrom sklearn.metrics import accuracy_score, f1_score, confusion_matrix, classification_report\n\nX_train, X_test, y_train, y_test = train_test_split(train_sample['question_text'], \ntrain_sample['target'], test_size=0.3)","metadata":{"execution":{"iopub.status.busy":"2021-05-27T02:51:44.889114Z","iopub.execute_input":"2021-05-27T02:51:44.889474Z","iopub.status.idle":"2021-05-27T02:51:44.969649Z","shell.execute_reply.started":"2021-05-27T02:51:44.889414Z","shell.execute_reply":"2021-05-27T02:51:44.968674Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Bag of Words (BoW)\n* Với một văn bản thì feature vector sẽ có dạng như thế nào? Làm sao đưa các từ, các câu, đoạn văn ở dạng text trong các văn bản về một vector mà mỗi phần tử là một số?\n* Có một phương pháp rất phổ biến giúp bạn trả lời những câu hỏi này. Phương pháp đó có tên là Bag of Words (BoW) (Túi đựng Từ)\n* Nhược điểm: không mang thông tin về thứ tự của các từ; cũng như sự liên kết giữa các câu, các đoạn văn trong văn bản\n* CountVectorizer để chuyển đổi văn bản thành một vector","metadata":{}},{"cell_type":"code","source":"vectorizer = CountVectorizer()\nvectorizer.fit(list(X_train) + list(X_test))\nX_train = vectorizer.transform(X_train) \nX_test = vectorizer.transform(X_test)","metadata":{"execution":{"iopub.status.busy":"2021-05-27T02:51:44.971236Z","iopub.execute_input":"2021-05-27T02:51:44.971659Z","iopub.status.idle":"2021-05-27T02:51:55.689349Z","shell.execute_reply.started":"2021-05-27T02:51:44.971617Z","shell.execute_reply":"2021-05-27T02:51:55.688560Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Logistic Regression","metadata":{}},{"cell_type":"markdown","source":"Phương pháp logistic regression là một mô hình hồi quy nhằm dự đoán giá trị đầu ra rời rạc (discrete target variable) y ứng với một véc-tơ đầu vào x. Việc này tương đương với chuyện phân loại các đầu vào x vào các nhóm y tương ứng\n","metadata":{}},{"cell_type":"code","source":"logistic = LogisticRegression() \nlogistic.fit(X_train,y_train)","metadata":{"execution":{"iopub.status.busy":"2021-05-27T02:51:55.690370Z","iopub.execute_input":"2021-05-27T02:51:55.690820Z","iopub.status.idle":"2021-05-27T02:52:06.342274Z","shell.execute_reply.started":"2021-05-27T02:51:55.690788Z","shell.execute_reply":"2021-05-27T02:52:06.341167Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_predictions = logistic.predict(X_train)\ntrain_acc = accuracy_score(y_train, train_predictions)  \ntrain_f1 = f1_score(y_train, train_predictions) \nprint(f\"Training accuracy: {train_acc:.2%}, F1: {train_f1:.4f}\") \ntest_predictions = logistic.predict(X_test)\ntest_acc = accuracy_score(y_test, test_predictions) \ntest_f1 = f1_score(y_test, test_predictions) \nprint(f\"Testing accuracy:  {test_acc:.2%}, F1: {test_f1:.4f}\")","metadata":{"execution":{"iopub.status.busy":"2021-05-27T02:52:06.343900Z","iopub.execute_input":"2021-05-27T02:52:06.344339Z","iopub.status.idle":"2021-05-27T02:52:06.571248Z","shell.execute_reply.started":"2021-05-27T02:52:06.344295Z","shell.execute_reply":"2021-05-27T02:52:06.570156Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"F1_score xấp xỉ 0,73","metadata":{}},{"cell_type":"markdown","source":"# Result","metadata":{}},{"cell_type":"code","source":"x_val = vectorizer.transform(test['question_text'])\nval_predictions = logistic.predict(x_val)\ntest['prediction'] = val_predictions\nsubmission = test[['qid', 'prediction']]\nsubmission.to_csv('submission.csv', index=False)\nsubmission","metadata":{"execution":{"iopub.status.busy":"2021-05-27T02:52:06.572388Z","iopub.execute_input":"2021-05-27T02:52:06.572675Z","iopub.status.idle":"2021-05-27T02:52:12.074686Z","shell.execute_reply.started":"2021-05-27T02:52:06.572648Z","shell.execute_reply":"2021-05-27T02:52:12.073706Z"},"trusted":true},"execution_count":null,"outputs":[]}]}