{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<img src = \"https://mytechdecisions.com/wp-content/uploads/2021/02/AdobeStock_382844018-1000x500.jpeg\"> <br><br>\n<h1 style = \"font-family: Times New Roman; font-weight: 1000; color:black\">🤖 Getting Started: Natural Language Processing with Disaster Tweets</h1>","metadata":{"execution":{"iopub.status.busy":"2022-08-02T12:41:09.049999Z","iopub.execute_input":"2022-08-02T12:41:09.050383Z","iopub.status.idle":"2022-08-02T12:41:09.058064Z","shell.execute_reply.started":"2022-08-02T12:41:09.050350Z","shell.execute_reply":"2022-08-02T12:41:09.056546Z"}}},{"cell_type":"markdown","source":"<div style=\"color:white;\n           display:fill;\n           border-radius:10px;\n           background-color:#00008B;\n           font-size:100%;\n           font-family:Times New Roman;\n           letter-spacing:0.5px\">\n\n<h1 id=\"basics\" style=\"color:white; font-family:Times New Roman; padding:20px\"> \n    Introduction\n</h1>\n</div>","metadata":{}},{"cell_type":"markdown","source":"<span style=\"font-size:18px; font-family:Times New Roman\">Welcome to my first attempt at a Natural Language Processing competition here on Kaggle!<br><br>\n    Thanks to smartphones and easy access to Internet everywhere, people are able to announce an emergency they're observing in real-time and that can be very useful to help people to avoid certain areas, to alert authorities, disaster relief organizations and news agencies. However, although a human could quite easily distinguish between a tweet that is about a real emergency and one that uses certain emergency-related words in a metaphorical way, it isn't quite the same when it comes to machines.<br><br>\n    The goal of this competition is to train a machine learning model that will learn how to differentiate between a tweet about a real disaster and one that isn't.<br><br>\n    To help me with that goal, I'm going to use PyCaret which is an AutoML library for Python and the reason I'm doing so is that whenever I'm learning something new about machine learning, I always like to use AutoML tools to help me observe how these models can actually work in practice!</span>","metadata":{}},{"cell_type":"markdown","source":"<div style=\"color:white;\n           display:fill;\n           border-radius:5px;\n           background-color:#00008B;\n           font-size:100%;\n           font-family:Times New Roman;\n           letter-spacing:0.5px\">\n\n<h1 id=\"basics\" style=\"color:white; font-family:Times New Roman; padding:10px\"> \n    What is Natural Language Processing?\n</h1>\n</div>","metadata":{}},{"cell_type":"markdown","source":"<span id=\"Introduction\" style=\"font-size:18px; font-family:Times New Roman\"> <b>Natural Language Processing (NLP)</b> is a branch of artificial intelligence that's focused on analyzing and understanding the languages that humans use naturally in order to interface with computers in written and spoken contexts.<br><br>\n    Nowadays, NLP in machine learning is used for sentiment analysis, which helps the machine to identify the mood or opinions within large amounts of text, document summarization, transforming voice commands into written text, and also automatic translation.</span>","metadata":{}},{"cell_type":"markdown","source":"<div style=\"color:white;\n           display:fill;\n           border-radius:5px;\n           background-color:#00008B;\n           font-size:100%;\n           font-family:Times New Roman;\n           letter-spacing:0.5px\">\n\n<h1 id=\"basics\" style=\"color:white; font-family:Times New Roman; padding:10px\"> \n    👨‍💻 | Getting Started: Let's Install All Necessary Libraries\n</h1>\n</div>","metadata":{}},{"cell_type":"code","source":"# Installing numpy 1.18.5\n!pip install numpy==1.18.5 --user\n# Installing PyCaret\n!pip install --ignore-installed pycaret --user\n# Installing english language model\n!python -m spacy download en_core_web_sm\n!python -m textblob.download_corpora","metadata":{"execution":{"iopub.status.busy":"2022-08-01T20:10:45.429054Z","iopub.execute_input":"2022-08-01T20:10:45.429474Z","iopub.status.idle":"2022-08-01T20:15:16.472764Z","shell.execute_reply.started":"2022-08-01T20:10:45.429441Z","shell.execute_reply":"2022-08-01T20:15:16.471430Z"},"_kg_hide-input":false,"_kg_hide-output":true,"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"color:white;\n           display:fill;\n           border-radius:5px;\n           background-color:#00008B;\n           font-size:100%;\n           font-family:Times New Roman;\n           letter-spacing:0.5px\">\n\n<h1 id=\"basics\" style=\"color:white; font-family:Times New Roman; padding:10px\"> \n    📈 | Getting to Know the Dataset\n</h1>\n</div>","metadata":{}},{"cell_type":"code","source":"# Importing libraries\nimport pandas as pd, plotly.express as px\nfrom plotly.offline import init_notebook_mode\ninit_notebook_mode(connected=True)\ntrain = pd.read_csv('../input/nlp-getting-started/train.csv')\ntest = pd.read_csv('../input/nlp-getting-started/test.csv')","metadata":{"execution":{"iopub.status.busy":"2022-08-02T14:38:09.817369Z","iopub.execute_input":"2022-08-02T14:38:09.817827Z","iopub.status.idle":"2022-08-02T14:38:09.865607Z","shell.execute_reply.started":"2022-08-02T14:38:09.817791Z","shell.execute_reply":"2022-08-02T14:38:09.864441Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Let's take a look at train dataframe\ntrain.head(10)","metadata":{"execution":{"iopub.status.busy":"2022-08-01T20:17:10.474435Z","iopub.execute_input":"2022-08-01T20:17:10.474985Z","iopub.status.idle":"2022-08-01T20:17:10.502468Z","shell.execute_reply.started":"2022-08-01T20:17:10.474936Z","shell.execute_reply":"2022-08-01T20:17:10.501629Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Printing the first few tweets\nprint(train['text'][:5])","metadata":{"execution":{"iopub.status.busy":"2022-08-01T20:24:54.754426Z","iopub.execute_input":"2022-08-01T20:24:54.754972Z","iopub.status.idle":"2022-08-01T20:24:54.767460Z","shell.execute_reply.started":"2022-08-01T20:24:54.754899Z","shell.execute_reply":"2022-08-01T20:24:54.766087Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Counting missing values\ntrain.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-08-02T13:19:06.101456Z","iopub.execute_input":"2022-08-02T13:19:06.101856Z","iopub.status.idle":"2022-08-02T13:19:06.121127Z","shell.execute_reply.started":"2022-08-02T13:19:06.101825Z","shell.execute_reply":"2022-08-02T13:19:06.118821Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Top 10 most frequent locations\ntrain.location.value_counts()[:10].plot(kind = 'bar')","metadata":{"execution":{"iopub.status.busy":"2022-08-02T13:38:14.865882Z","iopub.execute_input":"2022-08-02T13:38:14.866299Z","iopub.status.idle":"2022-08-02T13:38:15.106575Z","shell.execute_reply.started":"2022-08-02T13:38:14.866266Z","shell.execute_reply":"2022-08-02T13:38:15.105623Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Class balance\nfig = px.pie(train, names = 'target', title = 'Target Distribution')\nfig.update_traces(rotation=90, pull = [0.1], textinfo = \"percent+label\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-02T14:38:15.298080Z","iopub.execute_input":"2022-08-02T14:38:15.298621Z","iopub.status.idle":"2022-08-02T14:38:15.359601Z","shell.execute_reply.started":"2022-08-02T14:38:15.298578Z","shell.execute_reply":"2022-08-02T14:38:15.358264Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<span id=\"Introduction\" style=\"font-size:18px; font-family:Times New Roman\"> <b>0</b>: Not a real disaster<br><br>\n    <b>1</b>: Real disaster</span>","metadata":{}},{"cell_type":"markdown","source":"<div style=\"color:white;\n           display:fill;\n           border-radius:5px;\n           background-color:#00008B;\n           font-size:100%;\n           font-family:Times New Roman;\n           letter-spacing:0.5px\">\n\n<h1 id=\"basics\" style=\"color:white; font-family:Times New Roman; padding:10px\"> \n    🤖 | Using PyCaret's NLP Module\n</h1>\n</div>","metadata":{}},{"cell_type":"markdown","source":"<span id=\"Introduction\" style=\"font-size:18px; font-family:Times New Roman\">PyCaret's NLP module is an unsupervised machine learning module that's used to analyze the text data, divide text into different topics, and convert raw text into a format that machine learning algorithms can learn from<br><br>\nIt is also useful to automate text data treatment, such as removing numeric and special characters, word tokenization, bigram extraction, lemmatizing, and removing stopwords </span>","metadata":{}},{"cell_type":"code","source":"# Importing PyCaret's NLP module\nfrom pycaret.nlp import *","metadata":{"execution":{"iopub.status.busy":"2022-08-01T20:25:19.859822Z","iopub.execute_input":"2022-08-01T20:25:19.860259Z","iopub.status.idle":"2022-08-01T20:25:22.086061Z","shell.execute_reply.started":"2022-08-01T20:25:19.860224Z","shell.execute_reply":"2022-08-01T20:25:22.084825Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# First attempt at setting up texts\nsetup(train, # Defining dataframe\n      target = 'text', # Selecting target for treatment\n      session_id = 123 # Defining an id for reproducement)","metadata":{"execution":{"iopub.status.busy":"2022-08-01T20:26:34.093580Z","iopub.execute_input":"2022-08-01T20:26:34.094074Z","iopub.status.idle":"2022-08-01T20:27:08.570534Z","shell.execute_reply.started":"2022-08-01T20:26:34.094038Z","shell.execute_reply":"2022-08-01T20:27:08.569457Z"},"_kg_hide-input":false,"_kg_hide-output":true,"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Creating model\nlda = create_model('lda', # Choosing Algorithm\n                   num_topics = 3, # Setting 3 topics to divide texts into\n                   multi_core = True)","metadata":{"execution":{"iopub.status.busy":"2022-08-01T20:33:04.080122Z","iopub.execute_input":"2022-08-01T20:33:04.080554Z","iopub.status.idle":"2022-08-01T20:33:21.947294Z","shell.execute_reply.started":"2022-08-01T20:33:04.080522Z","shell.execute_reply":"2022-08-01T20:33:21.946414Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df = assign_model(lda) # Assigning model to dataframe","metadata":{"execution":{"iopub.status.busy":"2022-08-01T20:33:21.949414Z","iopub.execute_input":"2022-08-01T20:33:21.950023Z","iopub.status.idle":"2022-08-01T20:33:28.767261Z","shell.execute_reply.started":"2022-08-01T20:33:21.949987Z","shell.execute_reply":"2022-08-01T20:33:28.765997Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df # Observing results","metadata":{"execution":{"iopub.status.busy":"2022-08-01T20:36:06.969434Z","iopub.execute_input":"2022-08-01T20:36:06.970260Z","iopub.status.idle":"2022-08-01T20:36:06.995604Z","shell.execute_reply.started":"2022-08-01T20:36:06.970214Z","shell.execute_reply":"2022-08-01T20:36:06.994761Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_model(lda, plot = 'frequency') # Plotting word frequency","metadata":{"execution":{"iopub.status.busy":"2022-08-01T20:36:23.117748Z","iopub.execute_input":"2022-08-01T20:36:23.118154Z","iopub.status.idle":"2022-08-01T20:36:30.105098Z","shell.execute_reply.started":"2022-08-01T20:36:23.118120Z","shell.execute_reply":"2022-08-01T20:36:30.104025Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_model(lda, plot = 'topic_distribution') # Plotting topic distribution","metadata":{"execution":{"iopub.status.busy":"2022-08-01T20:37:59.443418Z","iopub.execute_input":"2022-08-01T20:37:59.443928Z","iopub.status.idle":"2022-08-01T20:38:06.440096Z","shell.execute_reply.started":"2022-08-01T20:37:59.443875Z","shell.execute_reply":"2022-08-01T20:38:06.439182Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_model(lda, plot = 'topic_model')","metadata":{"execution":{"iopub.status.busy":"2022-08-01T20:38:32.091056Z","iopub.execute_input":"2022-08-01T20:38:32.092148Z","iopub.status.idle":"2022-08-01T20:38:37.097694Z","shell.execute_reply.started":"2022-08-01T20:38:32.092107Z","shell.execute_reply":"2022-08-01T20:38:37.096511Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<span id=\"Introduction\" style=\"font-size:18px; font-family:Times New Roman\"> We can see that there are a high frequency of words that are irrelevant for training our model, such as <i>co</i> and <i>http</i>. We're gonna remove them by creating a list with custom stopwords, containing these words, and set up PyCaret's NLP module again</span>","metadata":{}},{"cell_type":"code","source":"stopwords = ['co','http','go','get'] # Creating a list with stopwords","metadata":{"execution":{"iopub.status.busy":"2022-08-01T20:39:52.837644Z","iopub.execute_input":"2022-08-01T20:39:52.839103Z","iopub.status.idle":"2022-08-01T20:39:52.844458Z","shell.execute_reply.started":"2022-08-01T20:39:52.839057Z","shell.execute_reply":"2022-08-01T20:39:52.843296Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Setting up nlp again\nsetup(train, target = 'text',session_id = 123, \n      custom_stopwords = stopwords # Informing stop words through custom_stopwords method)","metadata":{"execution":{"iopub.status.busy":"2022-08-01T20:40:21.246311Z","iopub.execute_input":"2022-08-01T20:40:21.246750Z","iopub.status.idle":"2022-08-01T20:40:54.160673Z","shell.execute_reply.started":"2022-08-01T20:40:21.246715Z","shell.execute_reply":"2022-08-01T20:40:54.159480Z"},"_kg_hide-input":false,"_kg_hide-output":true,"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"lda2 = create_model('lda')","metadata":{"execution":{"iopub.status.busy":"2022-08-01T20:41:18.196361Z","iopub.execute_input":"2022-08-01T20:41:18.196835Z","iopub.status.idle":"2022-08-01T20:41:34.995573Z","shell.execute_reply.started":"2022-08-01T20:41:18.196796Z","shell.execute_reply":"2022-08-01T20:41:34.994413Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_model(lda2, plot = 'frequency') # Plotting new word frequency","metadata":{"execution":{"iopub.status.busy":"2022-08-01T20:43:25.455935Z","iopub.execute_input":"2022-08-01T20:43:25.456372Z","iopub.status.idle":"2022-08-01T20:43:31.947960Z","shell.execute_reply.started":"2022-08-01T20:43:25.456336Z","shell.execute_reply":"2022-08-01T20:43:31.913314Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_model(lda2, plot = 'topic_distribution')","metadata":{"execution":{"iopub.status.busy":"2022-08-01T20:43:32.017626Z","iopub.execute_input":"2022-08-01T20:43:32.018170Z","iopub.status.idle":"2022-08-01T20:43:39.735659Z","shell.execute_reply.started":"2022-08-01T20:43:32.018114Z","shell.execute_reply":"2022-08-01T20:43:39.708563Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_model(lda, plot = 'topic_model')","metadata":{"execution":{"iopub.status.busy":"2022-08-01T20:44:13.457240Z","iopub.execute_input":"2022-08-01T20:44:13.458354Z","iopub.status.idle":"2022-08-01T20:44:15.257644Z","shell.execute_reply.started":"2022-08-01T20:44:13.458310Z","shell.execute_reply":"2022-08-01T20:44:15.256085Z"},"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df2 = assign_model(lda2)","metadata":{"execution":{"iopub.status.busy":"2022-08-01T20:42:53.611360Z","iopub.execute_input":"2022-08-01T20:42:53.611820Z","iopub.status.idle":"2022-08-01T20:42:59.934675Z","shell.execute_reply.started":"2022-08-01T20:42:53.611783Z","shell.execute_reply":"2022-08-01T20:42:59.909260Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df2","metadata":{"execution":{"iopub.status.busy":"2022-08-01T20:43:11.692471Z","iopub.execute_input":"2022-08-01T20:43:11.692952Z","iopub.status.idle":"2022-08-01T20:43:11.721549Z","shell.execute_reply.started":"2022-08-01T20:43:11.692900Z","shell.execute_reply":"2022-08-01T20:43:11.720224Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Importing PyCaret's classification module\nfrom pycaret.classification import *","metadata":{"execution":{"iopub.status.busy":"2022-08-01T20:45:31.147775Z","iopub.execute_input":"2022-08-01T20:45:31.148235Z","iopub.status.idle":"2022-08-01T20:45:31.455787Z","shell.execute_reply.started":"2022-08-01T20:45:31.148195Z","shell.execute_reply":"2022-08-01T20:45:31.454751Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Setting up data\nsetup(data = df2,target = 'target', train_size = 0.7,imputation_type = 'simple',\n     numeric_imputation = 'mean')","metadata":{"execution":{"iopub.status.busy":"2022-08-01T20:47:49.647060Z","iopub.execute_input":"2022-08-01T20:47:49.647520Z","iopub.status.idle":"2022-08-01T20:48:54.404029Z","shell.execute_reply.started":"2022-08-01T20:47:49.647484Z","shell.execute_reply":"2022-08-01T20:48:54.402854Z"},"_kg_hide-input":false,"_kg_hide-output":true,"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"compare_models(sort = 'F1', n_select = 3) # Sorting algorithm performance by F1 score. Selecting top 3 best models","metadata":{"execution":{"iopub.status.busy":"2022-08-01T20:49:46.805265Z","iopub.execute_input":"2022-08-01T20:49:46.805705Z","iopub.status.idle":"2022-08-01T21:19:46.658578Z","shell.execute_reply.started":"2022-08-01T20:49:46.805670Z","shell.execute_reply":"2022-08-01T21:19:46.657648Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"nb = create_model('nb')","metadata":{"execution":{"iopub.status.busy":"2022-08-01T21:20:59.272795Z","iopub.execute_input":"2022-08-01T21:20:59.273294Z","iopub.status.idle":"2022-08-01T21:21:03.865590Z","shell.execute_reply.started":"2022-08-01T21:20:59.273255Z","shell.execute_reply":"2022-08-01T21:21:03.864442Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"knn = create_model('knn')","metadata":{"execution":{"iopub.status.busy":"2022-08-01T21:21:11.123105Z","iopub.execute_input":"2022-08-01T21:21:11.123567Z","iopub.status.idle":"2022-08-01T21:21:39.785003Z","shell.execute_reply.started":"2022-08-01T21:21:11.123531Z","shell.execute_reply":"2022-08-01T21:21:39.783971Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"et = create_model('et')","metadata":{"execution":{"iopub.status.busy":"2022-08-01T21:21:39.787262Z","iopub.execute_input":"2022-08-01T21:21:39.788438Z","iopub.status.idle":"2022-08-01T21:23:27.948082Z","shell.execute_reply.started":"2022-08-01T21:21:39.788389Z","shell.execute_reply":"2022-08-01T21:23:27.946872Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"blendedmodels_1 = blend_models(estimator_list = [nb, knn, et],\n                              fold = 10, choose_better = True, optimize = 'F1',\n                              method = 'soft')","metadata":{"execution":{"iopub.status.busy":"2022-08-01T21:23:27.950185Z","iopub.execute_input":"2022-08-01T21:23:27.950560Z","iopub.status.idle":"2022-08-01T21:28:09.489642Z","shell.execute_reply.started":"2022-08-01T21:23:27.950526Z","shell.execute_reply":"2022-08-01T21:28:09.488429Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tuned_nb = tune_model(nb,optimize = 'F1',n_iter = 20,choose_better = True)","metadata":{"execution":{"iopub.status.busy":"2022-08-01T21:44:12.632788Z","iopub.execute_input":"2022-08-01T21:44:12.633307Z","iopub.status.idle":"2022-08-01T21:45:21.395811Z","shell.execute_reply.started":"2022-08-01T21:44:12.633267Z","shell.execute_reply":"2022-08-01T21:45:21.394610Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tuned_knn = tune_model(knn,optimize = 'F1',n_iter = 20,choose_better = True)","metadata":{"execution":{"iopub.status.busy":"2022-08-01T21:45:21.398147Z","iopub.execute_input":"2022-08-01T21:45:21.398605Z","iopub.status.idle":"2022-08-01T21:52:29.539774Z","shell.execute_reply.started":"2022-08-01T21:45:21.398573Z","shell.execute_reply":"2022-08-01T21:52:29.538393Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tuned_et = tune_model(et,optimize = 'F1',n_iter = 20,choose_better = True)","metadata":{"execution":{"iopub.status.busy":"2022-08-01T21:52:29.541294Z","iopub.execute_input":"2022-08-01T21:52:29.541625Z","iopub.status.idle":"2022-08-01T22:01:39.214458Z","shell.execute_reply.started":"2022-08-01T21:52:29.541595Z","shell.execute_reply":"2022-08-01T22:01:39.213187Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"blendedmodels_2 = blend_models(estimator_list = [tuned_nb, tuned_knn, tuned_et],\n                              fold = 10, choose_better = True, optimize = 'F1',\n                              method = 'soft')","metadata":{"execution":{"iopub.status.busy":"2022-08-01T22:01:39.216927Z","iopub.execute_input":"2022-08-01T22:01:39.217300Z","iopub.status.idle":"2022-08-01T22:03:24.767070Z","shell.execute_reply.started":"2022-08-01T22:01:39.217266Z","shell.execute_reply":"2022-08-01T22:03:24.765779Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tuned_blended_model = tune_model(blendedmodels_2,\n                                fold = 5,\n                                n_iter = 15,\n                                optimize = 'F1',\n                                choose_better = True)","metadata":{"execution":{"iopub.status.busy":"2022-08-01T22:07:45.719740Z","iopub.execute_input":"2022-08-01T22:07:45.720199Z","iopub.status.idle":"2022-08-01T22:13:19.817477Z","shell.execute_reply.started":"2022-08-01T22:07:45.720160Z","shell.execute_reply":"2022-08-01T22:13:19.816373Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<span id=\"Introduction\" style=\"font-size:18px; font-family:Times New Roman\">The best <b>F1 Score</b> was achieved by the <i>blendedmodels_2</i> model, which is the blending of the tuned versions of the original top 3 models.<br><br>\n    Now, lets use <i>predict_model</i> to see how our model performs on hold-out sample.</span>","metadata":{}},{"cell_type":"code","source":"predict_model(blendedmodels_2)","metadata":{"execution":{"iopub.status.busy":"2022-08-01T22:15:51.261485Z","iopub.execute_input":"2022-08-01T22:15:51.262077Z","iopub.status.idle":"2022-08-01T22:15:57.285227Z","shell.execute_reply.started":"2022-08-01T22:15:51.262039Z","shell.execute_reply":"2022-08-01T22:15:57.283847Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<span id=\"Introduction\" style=\"font-size:22px; font-family:Times New Roman\"><b>F1 Score on Hold-Out Sample</b>: 68.31%</span>","metadata":{}},{"cell_type":"code","source":"model = finalize_model(blendedmodels_2) # Finalizing Model","metadata":{"execution":{"iopub.status.busy":"2022-08-01T22:17:45.910323Z","iopub.execute_input":"2022-08-01T22:17:45.911275Z","iopub.status.idle":"2022-08-01T22:19:05.232291Z","shell.execute_reply.started":"2022-08-01T22:17:45.911227Z","shell.execute_reply":"2022-08-01T22:19:05.231139Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Now we must treat test text data with PyCaret's NLP module\nfrom pycaret.nlp import *\nsetup(test, target = 'text',session_id = 123, custom_stopwords = stopwords)","metadata":{"execution":{"iopub.status.busy":"2022-08-01T22:17:12.730423Z","iopub.execute_input":"2022-08-01T22:17:12.730852Z","iopub.status.idle":"2022-08-01T22:17:30.685952Z","shell.execute_reply.started":"2022-08-01T22:17:12.730818Z","shell.execute_reply":"2022-08-01T22:17:30.684647Z"},"_kg_hide-input":false,"_kg_hide-output":true,"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"lda = create_model('lda')","metadata":{"execution":{"iopub.status.busy":"2022-08-01T22:21:17.833387Z","iopub.execute_input":"2022-08-01T22:21:17.833891Z","iopub.status.idle":"2022-08-01T22:21:25.749762Z","shell.execute_reply.started":"2022-08-01T22:21:17.833856Z","shell.execute_reply":"2022-08-01T22:21:25.748609Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_model(lda, plot = 'topic_model')","metadata":{"execution":{"iopub.status.busy":"2022-08-01T22:21:47.519043Z","iopub.execute_input":"2022-08-01T22:21:47.519509Z","iopub.status.idle":"2022-08-01T22:21:48.322666Z","shell.execute_reply.started":"2022-08-01T22:21:47.519475Z","shell.execute_reply":"2022-08-01T22:21:48.321345Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df = assign_model(lda)","metadata":{"execution":{"iopub.status.busy":"2022-08-01T22:22:40.305079Z","iopub.execute_input":"2022-08-01T22:22:40.306275Z","iopub.status.idle":"2022-08-01T22:22:43.051694Z","shell.execute_reply.started":"2022-08-01T22:22:40.306228Z","shell.execute_reply":"2022-08-01T22:22:43.005600Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predictions = predict_model(model, data = test_df) # Using model to predict on test data","metadata":{"execution":{"iopub.status.busy":"2022-08-01T22:23:25.360724Z","iopub.execute_input":"2022-08-01T22:23:25.361266Z","iopub.status.idle":"2022-08-01T22:23:36.112472Z","shell.execute_reply.started":"2022-08-01T22:23:25.361218Z","shell.execute_reply":"2022-08-01T22:23:36.111216Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Seeing predictions dataframe\npredictions","metadata":{"execution":{"iopub.status.busy":"2022-08-01T22:23:44.264049Z","iopub.execute_input":"2022-08-01T22:23:44.264873Z","iopub.status.idle":"2022-08-01T22:23:44.293160Z","shell.execute_reply.started":"2022-08-01T22:23:44.264833Z","shell.execute_reply":"2022-08-01T22:23:44.292093Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predictions['target'] = predictions['Label']","metadata":{"execution":{"iopub.status.busy":"2022-08-01T22:25:28.206640Z","iopub.execute_input":"2022-08-01T22:25:28.207973Z","iopub.status.idle":"2022-08-01T22:25:28.215006Z","shell.execute_reply.started":"2022-08-01T22:25:28.207901Z","shell.execute_reply":"2022-08-01T22:25:28.213676Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission = predictions[['id','target']]","metadata":{"execution":{"iopub.status.busy":"2022-08-01T22:25:44.981823Z","iopub.execute_input":"2022-08-01T22:25:44.982302Z","iopub.status.idle":"2022-08-01T22:25:44.990980Z","shell.execute_reply.started":"2022-08-01T22:25:44.982266Z","shell.execute_reply":"2022-08-01T22:25:44.989666Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Submission dataframe\nsubmission","metadata":{"execution":{"iopub.status.busy":"2022-08-01T22:25:48.179001Z","iopub.execute_input":"2022-08-01T22:25:48.179444Z","iopub.status.idle":"2022-08-01T22:25:48.193624Z","shell.execute_reply.started":"2022-08-01T22:25:48.179410Z","shell.execute_reply":"2022-08-01T22:25:48.192219Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Saving submission to .csv file\nsubmission.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-01T22:26:30.935627Z","iopub.execute_input":"2022-08-01T22:26:30.936366Z","iopub.status.idle":"2022-08-01T22:26:30.958830Z","shell.execute_reply.started":"2022-08-01T22:26:30.936321Z","shell.execute_reply":"2022-08-01T22:26:30.957222Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Reading submission csv file\nsubmission = pd.read_csv('./submission.csv')\nsubmission","metadata":{"execution":{"iopub.status.busy":"2022-08-01T22:55:46.064697Z","iopub.execute_input":"2022-08-01T22:55:46.066104Z","iopub.status.idle":"2022-08-01T22:55:46.083658Z","shell.execute_reply.started":"2022-08-01T22:55:46.066039Z","shell.execute_reply":"2022-08-01T22:55:46.082345Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<span id=\"Introduction\" style=\"font-size:22px; font-family:Times New Roman\"><b>F1 Score on the Test Dataset</b>: 72.75%</span>","metadata":{}},{"cell_type":"markdown","source":"<div style=\"color:white;\n           display:fill;\n           border-radius:5px;\n           background-color:#00008B;\n           font-size:100%;\n           font-family:Times New Roman;\n           letter-spacing:0.5px\">\n\n<h1 id=\"basics\" style=\"color:white; font-family:Times New Roman; padding:10px\"> \n    👨‍💻 | Conclusion\n</h1>\n</div>","metadata":{}},{"cell_type":"markdown","source":"<span id=\"Introduction\" style=\"font-size:18px; font-family:Times New Roman\">Well, so this is my first attempt at working with machine learning for NLP.<br><br>\nOur F1 score on the test dataset was 72.75%, putting us among the top 91% of all competitors. Not only we can, but we must definitely try to improve this result later on. <br><br>\nAs a first attempt, I'm glad to have explored PyCaret's NLP module, which was completely new to me, even though I'm already very used to its classification and regression modules, and it was really fun to go through with this notebook.<br><br>\nAs I said before, we should definitely strive for better results, and I intend on coming back to this dataset and this competition to put into practice manually NLP processes for machine learning.<br><br><br>\nThank you for reading! Feel free to leave comments and suggestions.<br><br><br><br>\n    <i>Luís Fernando Torres</i></span>","metadata":{}}]}