{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# 🇬🇧 BBC NEWS CLASSIFICATION PROJECT \n## *(Unsupervised Learning Project)*\n#### 🏫 University of Colorado, Boulder - Unsupervised Algorithms in Machine Learning \n*By Mattison Hineline*\n\n\n**Overview**\n\nThis project looks at an unsupervised learning model technique called matrix factorization. All the required libraries to run the notebook are in the first coding cell. This notebook explores the training data while using common natural language processing techniques. Our goal is to classify different news articles into five different groups: business, tech, sport, entertainment or politics. We will train two models: (1) an unsupervised model using matrix factorization and then compare it with (2) a supervised model using KMeans clustering. Both models are submitted to have their testing accuracy. \n\n\n**1. Exploratory Data Analysis (EDA)🥸** \n\nIn this section we will explore the data given for the competition. It is important to look at what you have to work with before beginning any model building or model training. Normally, it is suggested to first split the data into training and testing sets before any EDA because you, as a researcher, do not want to be biased for the results which can lead to overfitting. For this reason, we will not visualize or explore the test data given, and use that data only for testing the models. \n\n**2. Model Building and Training 🦾**\n\nNext, we will build multiple unsupervised models and train them on the test data. Once we have at least two strong models built, we will fine-tune those models and test on the test dataset. We will be using the train-test split on this dataset that is already given in the CSV files.\n\n**3. Model Comparisons 🧐**\n\nOnce we have good, working models we will compare them for their performance. We will discuss which model is the best out of the models tested and why. \n\n**4. Conclusions 🤓**\n\nFinally, we will summary the project and results, including future suggestions for further analysis. ","metadata":{}},{"cell_type":"code","source":"#import important libraries\nimport numpy as np \nimport pandas as pd \nimport seaborn as sns\nimport matplotlib.pyplot as plt\nimport os\n\n#EDA and preprocessing\nimport re\nimport nltk.corpus\nfrom nltk.corpus import stopwords\nfrom nltk.tokenize import word_tokenize\nfrom nltk.stem import WordNetLemmatizer\nfrom string import digits\n\n#modeling\nfrom sklearn.feature_extraction.text import TfidfVectorizer, CountVectorizer\nfrom sklearn.decomposition import NMF\nfrom sklearn.metrics import accuracy_score\nimport sklearn.metrics as metrics\nimport itertools\nfrom sklearn.cluster import KMeans\nfrom sklearn.model_selection import train_test_split\n\n#find the files names\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))","metadata":{"execution":{"iopub.status.busy":"2022-07-04T18:16:56.412511Z","iopub.execute_input":"2022-07-04T18:16:56.413556Z","iopub.status.idle":"2022-07-04T18:16:58.906176Z","shell.execute_reply.started":"2022-07-04T18:16:56.413446Z","shell.execute_reply":"2022-07-04T18:16:58.904072Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#label file paths\npath_dir = '/kaggle/input/learn-ai-bbc/'\ntrain_path = path_dir + 'BBC News Train.csv'\nsample_solution_path = path_dir + 'BBC News Sample Solution.csv'\ntest_path = path_dir + 'BBC News Test.csv'","metadata":{"execution":{"iopub.status.busy":"2022-07-04T18:16:58.908558Z","iopub.execute_input":"2022-07-04T18:16:58.909709Z","iopub.status.idle":"2022-07-04T18:16:58.915986Z","shell.execute_reply.started":"2022-07-04T18:16:58.909662Z","shell.execute_reply":"2022-07-04T18:16:58.914410Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#import data\ntrain = pd.read_csv(train_path)\nsample_solution = pd.read_csv(sample_solution_path)\ntest = pd.read_csv(test_path)","metadata":{"execution":{"iopub.status.busy":"2022-07-04T18:16:58.917670Z","iopub.execute_input":"2022-07-04T18:16:58.918691Z","iopub.status.idle":"2022-07-04T18:16:59.121472Z","shell.execute_reply.started":"2022-07-04T18:16:58.918644Z","shell.execute_reply":"2022-07-04T18:16:59.119589Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"________\n# 1. Exploratory Analysis (EDA)\n_____","metadata":{}},{"cell_type":"code","source":"#look at what we are trying to submit \nsample_solution","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-04T18:16:59.124395Z","iopub.execute_input":"2022-07-04T18:16:59.124827Z","iopub.status.idle":"2022-07-04T18:16:59.154274Z","shell.execute_reply.started":"2022-07-04T18:16:59.124779Z","shell.execute_reply":"2022-07-04T18:16:59.152796Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After importing the data, we first take a look at what our ultimate goal for the project is by looking at the submission dataframe (above). We need to create a dataframe comprising of classified articles (ArticleId) and which category they belong to (category). We also see that each ArticleID is unique, while the categories repeat themselves. This is good information to know before continuing. Now that we know our end goal structure, let's move on to the training data. \n\nWe can see in the dataframe below that we have three columns: \n1. ArticleID: which is the identifying number for the article\n2. Test: the article header and text\n3. Category: category given to the article","metadata":{}},{"cell_type":"code","source":"#look at the training data\ntrain","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-04T18:16:59.159698Z","iopub.execute_input":"2022-07-04T18:16:59.160411Z","iopub.status.idle":"2022-07-04T18:16:59.187925Z","shell.execute_reply.started":"2022-07-04T18:16:59.160361Z","shell.execute_reply":"2022-07-04T18:16:59.186806Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# gain more information from the dataframe\ntrain.describe()","metadata":{"execution":{"iopub.status.busy":"2022-07-04T18:16:59.189424Z","iopub.execute_input":"2022-07-04T18:16:59.192862Z","iopub.status.idle":"2022-07-04T18:16:59.228439Z","shell.execute_reply.started":"2022-07-04T18:16:59.192814Z","shell.execute_reply":"2022-07-04T18:16:59.226924Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#check the type of data, null value counts and number of entries\ntrain.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-04T18:16:59.234975Z","iopub.execute_input":"2022-07-04T18:16:59.238079Z","iopub.status.idle":"2022-07-04T18:16:59.271349Z","shell.execute_reply.started":"2022-07-04T18:16:59.238024Z","shell.execute_reply":"2022-07-04T18:16:59.269306Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Great. All this information look as we expected. We have two object columns and one integer column (for the ID's). It looks like we are not missing any rows. Since the data we are working with is text, we don't need to worry about numbers that are missing such as 9999, 0, etc. Before continuing, I want to make sure that we have no repeated articles in the data. From the code below, we can see that we have 1490 unique IDs and we know the dataframe has 1490 rows, we can assume that each article is unique and continue. I also wanted to see how many categories there are. We can see below that there are five categories in total: business, tech, politics, sport, entertainment.","metadata":{}},{"cell_type":"code","source":"# check for repeated articles\ntrain['ArticleId'].nunique()","metadata":{"execution":{"iopub.status.busy":"2022-07-04T18:16:59.277959Z","iopub.execute_input":"2022-07-04T18:16:59.278704Z","iopub.status.idle":"2022-07-04T18:16:59.298852Z","shell.execute_reply.started":"2022-07-04T18:16:59.278648Z","shell.execute_reply":"2022-07-04T18:16:59.296460Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train['Category'].unique()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-04T18:16:59.300942Z","iopub.execute_input":"2022-07-04T18:16:59.302106Z","iopub.status.idle":"2022-07-04T18:16:59.311889Z","shell.execute_reply.started":"2022-07-04T18:16:59.302059Z","shell.execute_reply":"2022-07-04T18:16:59.310579Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The first thing I want to look at is a couple of texts to get an idea of how they are saved in the dataframe. We will look at the first row.  We can see that the text starts with the header, then has a fairly good amount of text afterwards which would be the article text. We can see that the data has already been preprocessed a bit because there are no uppercase characters. Since these are news articles, we can also assume that there are no spelling mistakes. Capitalization and spelling are two important factors when it comes to natural language processing. For this reason, we will be thankful that this has already been processed as such. ","metadata":{}},{"cell_type":"code","source":"# first row \ntrain['Text'][0]","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-04T18:16:59.319049Z","iopub.execute_input":"2022-07-04T18:16:59.321416Z","iopub.status.idle":"2022-07-04T18:16:59.334805Z","shell.execute_reply.started":"2022-07-04T18:16:59.321360Z","shell.execute_reply":"2022-07-04T18:16:59.332443Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now let's visualize the data as much as we can before running some models. We can see that overall we have about even number of entries for each category. This is good because if one or two categories was severely underrepresentated or, in contrast, overrepresentative in the data, then it may cause our model to be biased and/or perform poorly on some or all of the test data. ","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(8, 5))\nsns.histplot(\n    data = train,\n    x = 'Category',\n    hue = 'Category',\n    palette = 'colorblind',\n    legend = False,\n    ).set(\n        title = 'Category Counts');","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-04T18:16:59.338241Z","iopub.execute_input":"2022-07-04T18:16:59.342127Z","iopub.status.idle":"2022-07-04T18:16:59.912451Z","shell.execute_reply.started":"2022-07-04T18:16:59.342003Z","shell.execute_reply":"2022-07-04T18:16:59.910275Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"While working with text data, it is important to make the text \"readable\" for the computer. To do this, we will take three steps:\n1. remove punctuation\n2. remove stop words (common English words such as 'to', 'the', 'of', etc","metadata":{}},{"cell_type":"code","source":"def clean_text(dataframe, text_col):\n    '''\n    A helper function which takes a dataframe \n    and removes punction and stopwords.\n    '''\n    #remove all punctuation\n    dataframe['no_punct'] = dataframe[text_col].apply(lambda row: re.sub(r'[^\\w\\s]+', '', row))\n    \n    #remove numbers \n    dataframe['no_punct_num'] = dataframe['no_punct'].apply(lambda row: re.sub(r'[0-9]+', '', row))\n    \n    #remove stopwords\n    stop_words = stopwords.words('english')\n    dataframe['no_stopwords'] = dataframe['no_punct_num'].apply(lambda x: ' '.join([word for word in x.split() if word not in (stop_words)]))\n    \n    #remove extra spaces\n    dataframe['clean_text'] = dataframe['no_stopwords'].apply(lambda x: re.sub(' +', ' ', x))\n    return ","metadata":{"execution":{"iopub.status.busy":"2022-07-04T18:16:59.918002Z","iopub.execute_input":"2022-07-04T18:16:59.918484Z","iopub.status.idle":"2022-07-04T18:16:59.947530Z","shell.execute_reply.started":"2022-07-04T18:16:59.918434Z","shell.execute_reply":"2022-07-04T18:16:59.945027Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#clean dataframe text column\nclean_text(train, 'Text')","metadata":{"execution":{"iopub.status.busy":"2022-07-04T18:16:59.957566Z","iopub.execute_input":"2022-07-04T18:16:59.964065Z","iopub.status.idle":"2022-07-04T18:17:03.079920Z","shell.execute_reply.started":"2022-07-04T18:16:59.963970Z","shell.execute_reply":"2022-07-04T18:17:03.078809Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train['clean_text'][1]","metadata":{"execution":{"iopub.status.busy":"2022-07-04T18:17:03.085242Z","iopub.execute_input":"2022-07-04T18:17:03.087929Z","iopub.status.idle":"2022-07-04T18:17:03.100565Z","shell.execute_reply.started":"2022-07-04T18:17:03.087883Z","shell.execute_reply":"2022-07-04T18:17:03.099458Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After cleaning the dataframe text cells, we will also tokenize and lemmatize the text. Tokenize entails splitting a string of words into a list of words. For example \"cat sat on dog\" would be converted to ['cat', 'sat', 'on', 'dog']. Tokenizing splits up each word which we can then use later on to train models easier. Next, we will lemmatize the text. We can choose to lemmatize or stem the words. For this project, I chose to lemmatize the words because it keeps a bit more informationn than would stemming. An example of lemmatizing would be to take the words 'running', 'horses', and 'adjustable', and we lemmatize them to be 'run', 'horse', 'adjust'. This keeps the words general meaning but allows the model to learn better. Additionally, we will ensure that all words are in lowercase form. These cleaning steps mentioned before and here are important because, for example, the computer could look at two sentences such as \"Running big reddish dogs.\" versus \"Run big red dog!\" and these would be considered different even though they are quite similar. After cleaning, both sentences would be converted to ['run', 'big', 'red', 'dog'] and therefore these two \"articles\" probably would be classified together, which is the goal of our model.","metadata":{}},{"cell_type":"code","source":"# tokenize text function\nwordnet_lemmatizer = WordNetLemmatizer()\ndef lemmatizer(text):\n    ''' \n    A helper function to lemmatize an entire sentence/string\n    '''\n    lem = [wordnet_lemmatizer.lemmatize(word.lower()) for word in text] \n    return lem\n\ndef tokenize_lemmatize(dataframe, text_col):\n    '''\n    A helper function to tokenize then lemmatize the string.\n    Also, add column which counts the number of words in that string.\n    '''\n    dataframe['tokenized'] = dataframe.apply(lambda row: nltk.word_tokenize(row[text_col]), axis=1)\n    dataframe['lemmatized'] = dataframe['tokenized'].apply(lambda string: lemmatizer(string))\n    dataframe['num_words'] = dataframe['lemmatized'].apply(lambda lst: len(lst))\n    return","metadata":{"execution":{"iopub.status.busy":"2022-07-04T18:17:03.106513Z","iopub.execute_input":"2022-07-04T18:17:03.109170Z","iopub.status.idle":"2022-07-04T18:17:03.121584Z","shell.execute_reply.started":"2022-07-04T18:17:03.109125Z","shell.execute_reply":"2022-07-04T18:17:03.120233Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tokenize_lemmatize(train, 'clean_text')","metadata":{"execution":{"iopub.status.busy":"2022-07-04T18:17:03.127562Z","iopub.execute_input":"2022-07-04T18:17:03.131205Z","iopub.status.idle":"2022-07-04T18:17:10.465402Z","shell.execute_reply.started":"2022-07-04T18:17:03.131150Z","shell.execute_reply":"2022-07-04T18:17:10.464248Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After cleaning, we can see (below) the number of words per article. We see most articles are around 200 words. However, we also see we have some severe outliers that reach up to more than 750 words! We actually will remove these outliers as they might actually impact our model later on, in addition to creating more features (words) to have to calculate within the model. ","metadata":{}},{"cell_type":"code","source":"# number of tokens (words) per article\nfig, ax = plt.subplots(figsize=(15, 5))\nsns.histplot(\n    data = train, \n    x = 'num_words',\n    palette = 'colorblind',\n    ).set(\n        title = 'Number of Words per Article');","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-04T18:17:10.467266Z","iopub.execute_input":"2022-07-04T18:17:10.467685Z","iopub.status.idle":"2022-07-04T18:17:10.854400Z","shell.execute_reply.started":"2022-07-04T18:17:10.467641Z","shell.execute_reply":"2022-07-04T18:17:10.853264Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#remove outlier articles (longer than 750 words)\ntrain = train[train['num_words'] < 750]\nlen(train)","metadata":{"execution":{"iopub.status.busy":"2022-07-04T18:17:10.855958Z","iopub.execute_input":"2022-07-04T18:17:10.857099Z","iopub.status.idle":"2022-07-04T18:17:10.868977Z","shell.execute_reply.started":"2022-07-04T18:17:10.857053Z","shell.execute_reply":"2022-07-04T18:17:10.867669Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's also look at the number of words per category (boxplot below). Per category, we also see quite a few outliers. We will leave these this time. We also see that the mean of each category is similar, with tech and politics having more words, and variance, than the rest of the topics.","metadata":{}},{"cell_type":"code","source":"# words per category\nfig, ax = plt.subplots(figsize=(15, 5))\nsns.boxplot(\n    data = train, \n    x = 'num_words', \n    y = 'Category',\n    palette = 'colorblind'\n    ).set(\n        title = 'Number of Words Per Category');","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-04T18:17:10.871593Z","iopub.execute_input":"2022-07-04T18:17:10.872068Z","iopub.status.idle":"2022-07-04T18:17:11.257870Z","shell.execute_reply.started":"2022-07-04T18:17:10.872026Z","shell.execute_reply":"2022-07-04T18:17:11.256802Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"-------\n# 2. Model Building and Training \n-----\n\n","metadata":{}},{"cell_type":"code","source":"train_df = train.copy()","metadata":{"execution":{"iopub.status.busy":"2022-07-04T18:17:11.262894Z","iopub.execute_input":"2022-07-04T18:17:11.265978Z","iopub.status.idle":"2022-07-04T18:17:11.273269Z","shell.execute_reply.started":"2022-07-04T18:17:11.265929Z","shell.execute_reply":"2022-07-04T18:17:11.272176Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def predict(w_matrix):\n    sortedW = np.argsort(w_matrix)\n    n_predictions, maxValue = sortedW.shape\n    predictions = [[sortedW[i][maxValue - 1]] for i in range(n_predictions)]\n    topics = np.empty(n_predictions, dtype = np.int64)\n    for i in range(n_predictions):\n        topics[i] = predictions[i][0]\n    return topics","metadata":{"execution":{"iopub.status.busy":"2022-07-04T18:17:11.278289Z","iopub.execute_input":"2022-07-04T18:17:11.279281Z","iopub.status.idle":"2022-07-04T18:17:11.291739Z","shell.execute_reply.started":"2022-07-04T18:17:11.279236Z","shell.execute_reply":"2022-07-04T18:17:11.290368Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def label_permute(ytdf,yp,n=5):\n    \"\"\"\n    ytdf: labels dataframe object\n    yp: clustering label prediction output\n    Returns permuted label order and accuracy. \n    Example output: (3, 4, 1, 2, 0), 0.74 \n    \"\"\"\n    perms = list(itertools.permutations([0, 1, 2, 3, 4]))    #create permutation list\n    best_labels = []\n    best_acc = 0 \n    current = {}\n    labels = ['business', 'tech', 'politics', 'sport', 'entertainment']\n    for perm in perms:\n        for i in range(n):\n            current[labels[i]] = perm[i]\n            if len(current) == 5:\n                conditions = [\n                    (ytdf['Category'] == current['business']),\n                    (ytdf['Category'] == current['tech']),\n                    (ytdf['Category'] == current['politics']),\n                    (ytdf['Category'] == current['sport']),\n                    (ytdf['Category'] == current['entertainment'])]\n                ytdf['test'] = ytdf['Category'].map(current)\n                current_accuracy = accuracy_score(ytdf['test'], yp)\n                if current_accuracy > best_acc: \n                    best_acc = current_accuracy\n                    best_labels = perm\n                    ytdf['best'] = ytdf['test']\n    return best_labels, best_acc","metadata":{"execution":{"iopub.status.busy":"2022-07-04T18:17:11.297659Z","iopub.execute_input":"2022-07-04T18:17:11.298748Z","iopub.status.idle":"2022-07-04T18:17:11.317835Z","shell.execute_reply.started":"2022-07-04T18:17:11.298681Z","shell.execute_reply":"2022-07-04T18:17:11.316405Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#create vectorizer\ntfidvec = TfidfVectorizer(min_df = 2,\n                          max_df = 0.95,\n                          norm = 'l2',\n                          stop_words = 'english')\ntfidvec_train = tfidvec.fit_transform(train_df['clean_text'])\n\n#create model\nnmf_model = NMF(n_components=5, \n                init='nndsvda', \n                solver = 'mu',\n                beta_loss = 'kullback-leibler',\n                l1_ratio = 0.5,\n                random_state = 101)\nnmf_model.fit(tfidvec_train)\n\n#view results\nyhat_train = predict(nmf_model.transform(tfidvec_train))\nlabel_order, accuracy = label_permute(train_df, yhat_train )\nprint('accuracy=', accuracy)","metadata":{"execution":{"iopub.status.busy":"2022-07-04T18:17:11.320169Z","iopub.execute_input":"2022-07-04T18:17:11.320991Z","iopub.status.idle":"2022-07-04T18:17:18.502630Z","shell.execute_reply.started":"2022-07-04T18:17:11.320950Z","shell.execute_reply":"2022-07-04T18:17:18.501470Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"For this model, using matrix factorization, I found the best combination of parameters as above which resulted in the highest accuracy. I also tried other combinations, changing the TfidVectorizor and/or NMF model parameters. Particularly, I played around with min_df and man\\x_df in the TfidVectorizor\n- using max_df values of 0.85, 0.90, 0.95\n- using min_df values of 0, 1, and 2\n- using beta_loss of 'frobenius' and 'kullback-leibler'\n- using solver of 'mu' and 'cd'\n\nAlthough all models performed with higher than 85%, the combination above (in the code) was the best model.","metadata":{}},{"cell_type":"code","source":"#show best labels for the trained model \nlabel_dict = {4:'business', 2:'tech', 1:'politics', 0:'sport', 3:'entertainment'}\nfor i in range(5):\n    print(f'{label_order[i]}:  {label_dict[label_order[i]]}')","metadata":{"execution":{"iopub.status.busy":"2022-07-04T18:17:18.504424Z","iopub.execute_input":"2022-07-04T18:17:18.507040Z","iopub.status.idle":"2022-07-04T18:17:18.515468Z","shell.execute_reply.started":"2022-07-04T18:17:18.506991Z","shell.execute_reply":"2022-07-04T18:17:18.514191Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#first clean testing data as we did with the training data\nclean_text(test, 'Text')\ntfidvec_test = tfidvec.transform(test['clean_text'])\nyhat_test = predict(nmf_model.transform(tfidvec_test))","metadata":{"execution":{"iopub.status.busy":"2022-07-04T18:17:18.516995Z","iopub.execute_input":"2022-07-04T18:17:18.518112Z","iopub.status.idle":"2022-07-04T18:17:19.810337Z","shell.execute_reply.started":"2022-07-04T18:17:18.518038Z","shell.execute_reply":"2022-07-04T18:17:19.808796Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#create a submission dataframe\ntest_predictions = pd.DataFrame(columns=['ArticleId', 'Category', 'yhat'])\ntest_predictions['ArticleId'] = test['ArticleId']\ntest_predictions['yhat'] = yhat_test\ntest_predictions['Category'] = test_predictions['yhat'].apply(lambda i: label_dict[i])\n\n#delete columns unneeded for submission\ntest_predictions = test_predictions.drop('yhat', 1)\nprint(test_predictions.head(15))","metadata":{"execution":{"iopub.status.busy":"2022-07-04T18:17:19.812178Z","iopub.execute_input":"2022-07-04T18:17:19.812650Z","iopub.status.idle":"2022-07-04T18:17:19.838043Z","shell.execute_reply.started":"2022-07-04T18:17:19.812604Z","shell.execute_reply":"2022-07-04T18:17:19.836720Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# #save and submit test dataframe\n# try: \n#     test_predictions.to_csv('submission.csv', index=False)\n# except: \n#     pass\n\n# #public and private score was 0.96326","metadata":{"execution":{"iopub.status.busy":"2022-07-04T18:17:19.840005Z","iopub.execute_input":"2022-07-04T18:17:19.840793Z","iopub.status.idle":"2022-07-04T18:17:19.846107Z","shell.execute_reply.started":"2022-07-04T18:17:19.840741Z","shell.execute_reply":"2022-07-04T18:17:19.844837Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"To get the accuracy on the test set, I went ahead and submitted the results from the unsupervised model. Our model got a test accuracy score of 0.96326, or 96.3%, which is pretty good! Let's see if we can do even better using a supervised model. ","metadata":{}},{"cell_type":"markdown","source":"_______\n# 3. Model Comparisons \n\nFor this project, we are asked to use unsupervised learning to classify text articles uing a matrix factorization model. Traditionally, supervised models would perform better with this type of data if we have pre-labeled text, which we do. Therefore, we will first compare the unsupervised learning model above to another unsupervised model below using KMeans Clustering. After we have our results we will also train a supervised model, such as K-Nearest Neighbors.  \nFor good measure, we will re-import the data again since we'll be working with a new model. ","metadata":{}},{"cell_type":"code","source":"#import data\ntrain = pd.read_csv(train_path)\ntest = pd.read_csv(test_path)","metadata":{"execution":{"iopub.status.busy":"2022-07-04T18:17:19.859789Z","iopub.execute_input":"2022-07-04T18:17:19.860753Z","iopub.status.idle":"2022-07-04T18:17:19.973133Z","shell.execute_reply.started":"2022-07-04T18:17:19.860677Z","shell.execute_reply":"2022-07-04T18:17:19.971960Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#clean data\nclean_text(train, 'Text')\n\n#split data into X and y\ny_train = train['Category'].values\nX_train = train['clean_text'].values\n\n#create new vectorizer for supervised learning model\ntfidfvec_supervised = TfidfVectorizer(min_df = 2,\n                          max_df = 0.95,\n                          norm = 'l2',\n                          stop_words = 'english')\ntfSuper_train = tfidfvec_supervised.fit_transform(X_train) \n\n#create KMeans Model and train\nkmeans = KMeans(n_clusters = 5, \n                init = 'k-means++', \n                algorithm = 'full', \n                random_state = 101)\nyhat_train_super = kmeans.fit_predict(tfSuper_train)\n\n#get accuracy\ny_train_df = pd.DataFrame(y_train, columns=['Category'])\nlabel_order, accuracy = label_permute(y_train_df, yhat_train_super)\nprint('accuracy=', accuracy)\nprint(label_order, '\\n')\n\n#show label order\nlabel_dict = {3:'business', 1:'tech', 4:'politics', 2:'sport', 0:'entertainment'}\nfor i in range(5):\n    print(f'{label_order[i]}:  {label_dict[label_order[i]]}')","metadata":{"execution":{"iopub.status.busy":"2022-07-04T18:17:19.974893Z","iopub.execute_input":"2022-07-04T18:17:19.975315Z","iopub.status.idle":"2022-07-04T18:17:24.772976Z","shell.execute_reply.started":"2022-07-04T18:17:19.975270Z","shell.execute_reply":"2022-07-04T18:17:24.771708Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we will test the model on the test set. Again, we need to repeat the same steps to clean the data like we did in the training set. ","metadata":{}},{"cell_type":"code","source":"#clean data\nclean_text(test, 'Text')\n\n#split data\nX_test = test['clean_text'].values\n\n#create vectorizer (do not fit it!)\ntfSuper_test = tfidfvec_supervised.transform(X_test)\nyhat_test = kmeans.predict(tfSuper_test)\n\n#create a submission dataframe\ntest_predictions = pd.DataFrame(columns=['ArticleId', 'Category', 'yhat'])\ntest_predictions['ArticleId'] = test['ArticleId']\ntest_predictions['yhat'] = yhat_test\ntest_predictions['Category'] = test_predictions['yhat'].apply(lambda i: label_dict[i])\n\n#delete columns unneeded for submission\ntest_predictions = test_predictions.drop('yhat', 1)\nprint(test_predictions.head(2))","metadata":{"execution":{"iopub.status.busy":"2022-07-04T18:17:24.775018Z","iopub.execute_input":"2022-07-04T18:17:24.775459Z","iopub.status.idle":"2022-07-04T18:17:25.677217Z","shell.execute_reply.started":"2022-07-04T18:17:24.775410Z","shell.execute_reply":"2022-07-04T18:17:25.676194Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#save and submit test dataframe\n# try: \n#     test_predictions.to_csv('submission.csv', index=False)\n# except: \n#     pass\n\n# #public and private score was 0.62993","metadata":{"execution":{"iopub.status.busy":"2022-07-04T18:17:25.679001Z","iopub.execute_input":"2022-07-04T18:17:25.679824Z","iopub.status.idle":"2022-07-04T18:17:25.686549Z","shell.execute_reply.started":"2022-07-04T18:17:25.679775Z","shell.execute_reply":"2022-07-04T18:17:25.685570Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It looks like we got a training accuracy of 93.69% but only a testing accuracy of 62.99%. Wow! This suggests that our model is probably overfitting to the training data and does a poor job predicting on new data. If we wanted to train a more powerful model, we could use techniques such as ensemble methods and/or cross-validation (Kfold), or different models such as decision tree, random forest, SVM, etc. Let's now compare our two unsupervised learning models to a supervised model, such as KNN (K Nearest Neighbors). ","metadata":{}},{"cell_type":"code","source":"from sklearn.ensemble import RandomForestClassifier\nfrom sklearn.feature_extraction.text import CountVectorizer\n\n#import data\ntrain = pd.read_csv(train_path)\ntest = pd.read_csv(test_path)\n\n#clean data and split into x and y training\nclean_text(train, 'Text')\n\nvectorizer = CountVectorizer() \ny_train = train['Category'].values\nX_train = train['clean_text']\nX_train_vector = vectorizer.fit_transform(X_train)\n\n\n#build model\nclf = RandomForestClassifier(random_state=101)\nclf.fit(X_train_vector, y_train)\n\n#train accuracy \nprint(\"Train Accuracy\",accuracy_score(y_train, clf.predict(X_train_vector)))\n\n#test set  accuracy\nclean_text(test, 'Text')\nX_test = test['clean_text']\nX_test_vector = vectorizer.transform(X_test)\ntest_predictions = pd.DataFrame(columns=['ArticleId', 'Category'])\ntest_predictions['ArticleId'] = test['ArticleId']\ntest_predictions['Category'] = clf.predict(X_test_vector)\n\n#save and submit test dataframe to get accuracy \n# try: \n#     test_predictions.to_csv('submission.csv', index=False)\n# except: \n#     pass\n\n#TEST ACCURACY FROM SUBMISSION = 0.96190, or 96.19%","metadata":{"execution":{"iopub.status.busy":"2022-07-04T18:45:55.126751Z","iopub.execute_input":"2022-07-04T18:45:55.127408Z","iopub.status.idle":"2022-07-04T18:45:59.006594Z","shell.execute_reply.started":"2022-07-04T18:45:55.127372Z","shell.execute_reply":"2022-07-04T18:45:59.005673Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This is very interesting now because we can see that the supervised learning model using Random Forest Classifeir got 96.19% which is really good! Our unsupervised model using matrix factorization happened to perform better on this dataset but that is always to be expected. We also see that the training accuracy for the supervised learning model is 100%, which makes sense because we trained the model with labels on the data then tested the accuracy on the data. In other words, it already knew the answers before guessing! What a fun experiment for unsupervised and supervised models. ","metadata":{}},{"cell_type":"markdown","source":"______ \n# 4. Conclusions\n______\n\nTo summarize this project, we first cleaned the training data in common NLP preprocessing ways and explored the data. Then we created a matrix factorization model and got a testing accuracy of 96.3%. We got this score by find-tuning some parameters and using the training accuracy as a guide. The unsupervised model did quite well coompared to the supervised learning model. We made sure to preprocess the data in the same way for all training and testing runs. We also notice that the KMeans clustering (unsupervised model) performed the worst. \n\n\n**Future Project Enhancement Options**: \nHere are a few things that could be performed in future NLP projects which may impact the results in this study by improving or decreasing the performance. It is worth trying out. This project removed captialized letters, all numbers, and punctuation. In addition, we did not deal with mispelled or 'non-real' words. These things can impact the outcomes of a model and are worth modifying and trying out in different ways to see if better results can be obtained.","metadata":{}},{"cell_type":"markdown","source":"\nEnd 🥳\n______","metadata":{}}]}