{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Finding the best approach for each discourse type","metadata":{}},{"cell_type":"markdown","source":"![](https://blog.inkforall.com/wp-content/uploads/2019/07/myalt-207914158.png)\n","metadata":{}},{"cell_type":"markdown","source":"  <a id='top'></a>\n<div class=\"list-group\" id=\"list-tab\" role=\"tablist\">\n<p style=\"background-color:#4a8fdd;font-family:newtimeroman;color:#FFF9ED;font-size:300%;text-align:center;border-radius:9px 9px;\">TABLE OF CONTENTS</p>   \n    \n* [1. IMPORTING NECESARY LIBRARIES](#1)\n\n* [2. DATASET](#2)\n\n* [3. DATA VISUALIZATION](#3)\n    \n    * [3.1. Unique values](#3.1)   \n    * [3.2. Types repartition](#3.2)\n    \n    \n* [4. DATA CLEANING](#4)\n    \n* [5. DATA ANALYSIS](#5)\n    \n    * [5.1. Most frequent words](#5.1)   \n    * [5.2. Wordclouds](#5.2)\n    \n    \n* [6. VECTORIZATION](#6)\n    \n    * [6.1. Common functions](#6.1)   \n    * [6.2. TF-IDF](#6.2)\n    * [6.3. Universal Sentence Encoder](#6.3)   \n    * [6.4. BERT](#6.4)\n    \n    \n* [7. CLASSIFIERS COMPARISON](#7)\n    \n    * [7.1. Evaluation](#7.1)   \n    * [7.2. Classifiers list](#7.2)\n    * [7.3. Split Train/Test](#7.3)   \n    * [7.4. Baseline](#7.4)\n    * [7.5. Models comparison](#7.5)\n    * [7.6. Cross-Validation](#7.6)\n    \n    \n* [8. SPLITTING TYPES](#8)\n    \n    * [8.1. Models comparison](#8.1)\n    * [8.2. Cross-Validation](#8.2)\n    * [8.3. Overall performances](#8.3)\n    \n    \n* [9. MIXING THE MODELS](#9)\n    \n    * [9.1. Data subsets comparison](#9.1)\n    * [9.2. Performances](#9.2)\n    * [9.3. Saving the parameters](#9.3)\n    \n    \n* [10. MAKING THE PREDICTIONS](#10)\n    \n    * [10.1. Training the model](#10.1)\n    * [10.2. Vectorizing text data](#10.2)\n    * [10.3. Generating predictions](#10.3)","metadata":{}},{"cell_type":"markdown","source":"<a id=\"1\"></a>\n\n## <b>1 <span style='color:#4a8fdd'>|</span> IMPORTING NECESARY LIBRARIES</b>\n\n","metadata":{}},{"cell_type":"markdown","source":"We'll use <i>numpy</i> and <i>pandas</i> to manipulate arrays and dataframes. <i>Pyplot</i>, <i>plotly</i> and <i>seaborn</i> will help for visualizations, while <i>string</i> and <i>re</i> are useful for text cleaning via regular expressions.","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\n\nimport matplotlib.pyplot as plt\nimport plotly.express as px\nimport seaborn as sns\n\nimport string\nimport re","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-07-09T10:39:54.325184Z","iopub.execute_input":"2022-07-09T10:39:54.325687Z","iopub.status.idle":"2022-07-09T10:39:56.246443Z","shell.execute_reply.started":"2022-07-09T10:39:54.32557Z","shell.execute_reply":"2022-07-09T10:39:56.24527Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"2\"></a>\n\n## <b>2 <span style='color:#4a8fdd'>|</span> DATASET</b>","metadata":{}},{"cell_type":"markdown","source":"Three files are available:\n - <i>train.csv</i> contains the training data,\n - <i>test.csv</i> contains testing data for submission purposes,\n - <i>sample_submission.csv</i> is an exemple of submission's format.","metadata":{}},{"cell_type":"code","source":"train = pd.read_csv('../input/feedback-prize-effectiveness/train.csv')\nprint(train.shape)\ntrain.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T10:39:56.248888Z","iopub.execute_input":"2022-07-09T10:39:56.249322Z","iopub.status.idle":"2022-07-09T10:39:56.534645Z","shell.execute_reply.started":"2022-07-09T10:39:56.24928Z","shell.execute_reply":"2022-07-09T10:39:56.533495Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test = pd.read_csv('../input/feedback-prize-effectiveness/test.csv')\nprint(test.shape)\ntest.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:30:28.756749Z","iopub.execute_input":"2022-07-09T15:30:28.757181Z","iopub.status.idle":"2022-07-09T15:30:28.777848Z","shell.execute_reply.started":"2022-07-09T15:30:28.757149Z","shell.execute_reply":"2022-07-09T15:30:28.776748Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submit = pd.read_csv('../input/feedback-prize-effectiveness/sample_submission.csv')\nprint(submit.shape)\nsubmit.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T10:39:56.55776Z","iopub.execute_input":"2022-07-09T10:39:56.558651Z","iopub.status.idle":"2022-07-09T10:39:56.575724Z","shell.execute_reply.started":"2022-07-09T10:39:56.55861Z","shell.execute_reply":"2022-07-09T10:39:56.574641Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"3\"></a>\n\n## <b>3 <span style='color:#4a8fdd'>|</span> DATA VISUALIZATION</b>\n","metadata":{}},{"cell_type":"markdown","source":"<a id=\"3.1\"></a>\n\n#### <b>3.1 <span style='color:#4a8fdd'>|</span> Unique values</b>\n\nLet's start with counting the unique values for each variables, and listing the different <i>discourse_type</i> available.","metadata":{}},{"cell_type":"code","source":"print(train.shape[0], \"entries\")\ntrain.nunique()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T10:39:56.57819Z","iopub.execute_input":"2022-07-09T10:39:56.578471Z","iopub.status.idle":"2022-07-09T10:39:56.671654Z","shell.execute_reply.started":"2022-07-09T10:39:56.578444Z","shell.execute_reply":"2022-07-09T10:39:56.670474Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"discourse_types = set(train[\"discourse_type\"].values)\nprint(discourse_types)","metadata":{"execution":{"iopub.status.busy":"2022-07-09T10:39:56.676672Z","iopub.execute_input":"2022-07-09T10:39:56.679396Z","iopub.status.idle":"2022-07-09T10:39:56.692621Z","shell.execute_reply.started":"2022-07-09T10:39:56.679346Z","shell.execute_reply":"2022-07-09T10:39:56.691527Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"3.2\"></a>\n\n#### <b>3.2 <span style='color:#4a8fdd'>|</span> Types repartition</b>\n\nWe have seven different types of discourse, and three <i>discourse_effectiveness</i> possible values. Let's check the repartition of these variables in our dataset.","metadata":{}},{"cell_type":"code","source":"def proportions_pie(data, col, title):\n    counts = data[col].value_counts()\n    plt.pie(counts, labels=counts.index, autopct='%1.1f%%', radius=2)\n    plt.axis('equal')\n    plt.title(title)\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T10:39:56.697288Z","iopub.execute_input":"2022-07-09T10:39:56.69798Z","iopub.status.idle":"2022-07-09T10:39:56.708211Z","shell.execute_reply.started":"2022-07-09T10:39:56.697936Z","shell.execute_reply":"2022-07-09T10:39:56.707093Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"proportions_pie(train, 'discourse_type', 'Discourse type proportions')\nproportions_pie(train, 'discourse_effectiveness', 'Discourse effectiveness proportions')","metadata":{"execution":{"iopub.status.busy":"2022-07-09T10:39:56.713377Z","iopub.execute_input":"2022-07-09T10:39:56.71409Z","iopub.status.idle":"2022-07-09T10:39:57.044085Z","shell.execute_reply.started":"2022-07-09T10:39:56.714047Z","shell.execute_reply":"2022-07-09T10:39:57.042769Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"counts = train.groupby(by='discourse_type')[\"discourse_effectiveness\"].value_counts(normalize=True).unstack(fill_value=0)\ncounts[['Effective', 'Adequate', 'Ineffective']].plot.bar(stacked=True, color=['tab:green', 'tab:cyan', 'tab:red'])\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T10:39:57.049652Z","iopub.execute_input":"2022-07-09T10:39:57.052472Z","iopub.status.idle":"2022-07-09T10:39:57.361732Z","shell.execute_reply.started":"2022-07-09T10:39:57.052421Z","shell.execute_reply":"2022-07-09T10:39:57.360577Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"4\"></a>\n\n## <b>4 <span style='color:#4a8fdd'>|</span> DATA CLEANING</b>\n\nWe make a first cleaning by filling the <i>NaN</i> values before removing the <i>HTML</i> tags, the <i>digits</i> and the unnecessary <i>whitespaces</i>.","metadata":{}},{"cell_type":"code","source":"def cleaning(text_data):\n    text_data = text_data.fillna(\"\")\n    text_data = text_data.apply(lambda x: re.sub(r'\\<[^\\>]*\\>', ' ', x))\n    text_data = text_data.apply(lambda x: re.sub(r'\\d+', '', x))\n    text_data = text_data.apply(lambda x: re.sub(rf\"[{re.escape(string.whitespace)}]+\", ' ', x))\n    return text_data","metadata":{"execution":{"iopub.status.busy":"2022-07-09T10:39:57.366733Z","iopub.execute_input":"2022-07-09T10:39:57.367202Z","iopub.status.idle":"2022-07-09T10:39:57.37211Z","shell.execute_reply.started":"2022-07-09T10:39:57.367173Z","shell.execute_reply":"2022-07-09T10:39:57.371397Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data = train[['discourse_text', 'discourse_type', 'discourse_effectiveness']]\ndata['clean_text'] = cleaning(data['discourse_text'])\ndata.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T10:39:57.373147Z","iopub.execute_input":"2022-07-09T10:39:57.373636Z","iopub.status.idle":"2022-07-09T10:39:58.31698Z","shell.execute_reply.started":"2022-07-09T10:39:57.373605Z","shell.execute_reply":"2022-07-09T10:39:58.316251Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Some vectorization techniques require further cleaning, namely removing the <i>punctuations</i>, the <i>uppercase</i> characters and the <i>stopwords</i>. We shall alos <b>lemmatize</b> the words thanks to the package <i>nltk</i>.","metadata":{}},{"cell_type":"code","source":"import nltk\nfrom nltk.corpus import stopwords\nfrom nltk.stem import WordNetLemmatizer\n\ndef more_cleaning(text_data):\n    punctuations = '.,!?:;-=...\"@#_()' + \"'\"\n    text_data = text_data.apply(lambda x: re.sub(rf'[{punctuations}]', ' ', x))\n    text_data = text_data.apply(lambda x: x.lower())\n    stop = set(stopwords.words('english'))\n    lemmatizer = WordNetLemmatizer()\n    text_data = text_data.apply(lambda x: \" \".join([lemmatizer.lemmatize(word) for word in x.split(\" \") if lemmatizer.lemmatize(word) not in stop]))\n    text_data = text_data.apply(lambda x: re.sub(rf\"[{re.escape(string.whitespace)}]+\", ' ', x))\n    text_data = text_data.fillna(\"\")\n    return text_data","metadata":{"execution":{"iopub.status.busy":"2022-07-09T10:39:58.318264Z","iopub.execute_input":"2022-07-09T10:39:58.318812Z","iopub.status.idle":"2022-07-09T10:39:59.105974Z","shell.execute_reply.started":"2022-07-09T10:39:58.318782Z","shell.execute_reply":"2022-07-09T10:39:59.104563Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data['cleaner_text'] = more_cleaning(data[\"clean_text\"])\ndata.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T10:39:59.107454Z","iopub.execute_input":"2022-07-09T10:39:59.107926Z","iopub.status.idle":"2022-07-09T10:40:10.838402Z","shell.execute_reply.started":"2022-07-09T10:39:59.107877Z","shell.execute_reply":"2022-07-09T10:40:10.837549Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"5\"></a>\n\n## <b>5 <span style='color:#4a8fdd'>|</span> DATA ANALYSIS</b>\n","metadata":{}},{"cell_type":"markdown","source":"<a id=\"5.1\"></a>\n\n#### <b>5.1 <span style='color:#4a8fdd'>|</span> Most frequent words</b>\n\nLet's see the 50 most frequent words in our dataset, with both cleaning techniques.","metadata":{}},{"cell_type":"code","source":"def top_words(text_data, normalize=False):\n    pattern = (rf\"((\\w)[{string.punctuation}](?:\\B|$)|(?:^|\\B)[{string.punctuation}](\\w))\")\n    return (text_data.apply(lambda x: x.replace(pattern, r\"\\2 \\3\").split()).explode().value_counts(normalize=normalize))\n\ndef frequency_barplot(df, nr_top_words=50):\n    fig, ax = plt.subplots(1,1,figsize=(20,5))\n    df.sort_values(ascending=False, inplace = True)\n    sns.barplot(list(range(nr_top_words)), df.values[:nr_top_words], palette='hls', ax=ax)\n    ax.set_xticks(list(range(nr_top_words)))\n    ax.set_xticklabels(df.index[:nr_top_words], fontsize=14, rotation=90)\n    return ax","metadata":{"execution":{"iopub.status.busy":"2022-07-09T10:40:10.839702Z","iopub.execute_input":"2022-07-09T10:40:10.840195Z","iopub.status.idle":"2022-07-09T10:40:10.848868Z","shell.execute_reply.started":"2022-07-09T10:40:10.840162Z","shell.execute_reply":"2022-07-09T10:40:10.84802Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"nbr_words = 50\n\nwords_org = top_words(data['clean_text']).head(nbr_words)\nwords_cln = top_words(data['cleaner_text']).head(nbr_words)\n\nax = frequency_barplot(words_org)\nax.set_title(\"Words Frequencies (cleaned texts) - 50 most frequent\", fontsize=16);\nax = frequency_barplot(words_cln)\nax.set_title(\"Words Frequencies (lemmatized texts) - 50 most frequent\", fontsize=16);","metadata":{"execution":{"iopub.status.busy":"2022-07-09T10:40:10.850256Z","iopub.execute_input":"2022-07-09T10:40:10.850647Z","iopub.status.idle":"2022-07-09T10:40:12.920942Z","shell.execute_reply.started":"2022-07-09T10:40:10.850617Z","shell.execute_reply":"2022-07-09T10:40:12.919748Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"5.2\"></a>\n\n#### <b>5.2 <span style='color:#4a8fdd'>|</span> Wordclouds</b>\n\n<i>Wordclouds</i> are a great way to visualize the most frequent word and their relative importance.","metadata":{}},{"cell_type":"code","source":"#from collections import Counter\nfrom wordcloud import WordCloud\n\ndef draw_wordcloud(text_data, max_words=200, colormap='viridis', figsize=(20,10)):\n    words = top_words(text_data)\n    wordcloud = WordCloud(width=800, height=400, max_words=max_words, min_font_size=4, background_color=\"WHITE\",\n        max_font_size=None, relative_scaling='auto', colormap=colormap).generate_from_frequencies(words)\n    fig, ax = plt.subplots(figsize=figsize)\n    ax.imshow(wordcloud, interpolation=\"bilinear\")\n    ax.axis(\"off\")","metadata":{"execution":{"iopub.status.busy":"2022-07-09T10:40:12.92257Z","iopub.execute_input":"2022-07-09T10:40:12.923394Z","iopub.status.idle":"2022-07-09T10:40:12.957156Z","shell.execute_reply.started":"2022-07-09T10:40:12.923355Z","shell.execute_reply":"2022-07-09T10:40:12.95606Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"draw_wordcloud(data['clean_text'], max_words=100, colormap='viridis', figsize=(15,8))\ndraw_wordcloud(data['cleaner_text'], max_words=100, colormap='plasma', figsize=(15,8))","metadata":{"execution":{"iopub.status.busy":"2022-07-09T10:40:12.958533Z","iopub.execute_input":"2022-07-09T10:40:12.959107Z","iopub.status.idle":"2022-07-09T10:40:15.312752Z","shell.execute_reply.started":"2022-07-09T10:40:12.959073Z","shell.execute_reply":"2022-07-09T10:40:15.311634Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"6\"></a>\n\n## <b>6 <span style='color:#4a8fdd'>|</span> VECTORIZATION</b>","metadata":{}},{"cell_type":"markdown","source":"<a id=\"6.1\"></a>\n\n#### <b>6.1 <span style='color:#4a8fdd'>|</span> Common functions</b>\n\nFor every vectorization we shall implement, we will need to visualize the data via a <i>scatterplot</i> after a <i>dimensional reduction</i>. The latter will be obtained by a <i>t-SNE</i> algorithm.","metadata":{}},{"cell_type":"code","source":"from sklearn.manifold import TSNE\n\nlist_vector = []\n\ndef scatterplot(data, col, color = None, hover_name = None, hover_data = None, title = \"\"):\n    plot_values = np.stack(data[col], axis=1)\n    dimension = len(plot_values)\n    if dimension < 2 or dimension > 3:\n        print(\"The column you want to visualize has dimension < 2 or dimension > 3.\")\n        print(\" The function can only visualize 2- and 3-dimensional data.\")\n    else:\n        if dimension == 2:\n            x, y = plot_values[0], plot_values[1]\n            fig = px.scatter(data, x=x, y=y, color=color, hover_data=hover_data, title=title, hover_name=hover_name)\n        else:\n            x, y, z = plot_values[0], plot_values[1], plot_values[2]\n            fig = px.scatter_3d(data, x=x, y=y, z=z, color=color, hover_data=hover_data, title=title, hover_name=hover_name)\n        fig.show()\n        \ndef tsne(data, n_components=2, perplexity=30.0, learning_rate=200.0, n_iter=1000, random_state=None, n_jobs=-1):\n    tsne = TSNE(n_components=n_components, perplexity=perplexity, learning_rate=learning_rate, n_iter=n_iter,\n        random_state=random_state, n_jobs=n_jobs)\n    if isinstance(data, pd.DataFrame):\n        data_coo = data.sparse.to_coo()\n        data_for_vectorization = data_coo.astype(\"float64\")\n    else:\n        data_for_vectorization = list(data)\n    return pd.Series(list(tsne.fit_transform(data_for_vectorization)), index=data.index)\n","metadata":{"execution":{"iopub.status.busy":"2022-07-09T10:40:15.314324Z","iopub.execute_input":"2022-07-09T10:40:15.315271Z","iopub.status.idle":"2022-07-09T10:40:15.404364Z","shell.execute_reply.started":"2022-07-09T10:40:15.315236Z","shell.execute_reply":"2022-07-09T10:40:15.403445Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"6.2\"></a>\n\n#### <b>6.2 <span style='color:#4a8fdd'>|</span> TF-IDF</b>\n\nEach variable will correspond to a word in the corpus, and each text will have coefficients proportional to the frequence of such word in said text <b>but also</b> to the inverse frequency in the complete corpus, making common words less important than specific ones.","metadata":{}},{"cell_type":"code","source":"from sklearn.feature_extraction.text import TfidfVectorizer, CountVectorizer\n\ndef tokenize(text_data):\n    punct = string.punctuation.replace(\"_\", \"\")\n    pattern = rf\"((\\w)([{punct}])(?:\\B|$)|(?:^|\\B)([{punct}])(\\w))\"\n    return text_data.apply(lambda x: x.replace(pattern, r\"\\2 \\3 \\4 \\5\").split())\n\ndef tfidf(text_data, max_features=None, min_df=1, max_df=1.0):\n    if not isinstance(text_data.iloc[0], list):\n        text_data = tokenize(text_data)\n    tfidf = TfidfVectorizer(use_idf=True, max_features=max_features, min_df=min_df, max_df=max_df,\n        tokenizer=lambda x: x, preprocessor=lambda x: x, norm=None)\n    tfidf_vectors_csr = tfidf.fit_transform(text_data)\n    features = tfidf.get_feature_names()\n    return (pd.Series(tfidf_vectors_csr.todense().tolist(), text_data.index), features, tfidf)","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:30:39.164272Z","iopub.execute_input":"2022-07-09T15:30:39.164804Z","iopub.status.idle":"2022-07-09T15:30:39.17874Z","shell.execute_reply.started":"2022-07-09T15:30:39.164758Z","shell.execute_reply":"2022-07-09T15:30:39.176835Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"(data['TF-IDF'], text_features, tfidf_vectorizer) = tfidf(data['cleaner_text'], max_features=1000)\nlist_vector.append('TF-IDF')\ndata.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:30:43.293753Z","iopub.execute_input":"2022-07-09T15:30:43.294134Z","iopub.status.idle":"2022-07-09T15:30:43.330436Z","shell.execute_reply.started":"2022-07-09T15:30:43.294105Z","shell.execute_reply":"2022-07-09T15:30:43.328908Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data['TF-IDF 2D'] = tsne(data['TF-IDF'])\nscatterplot(data, col='TF-IDF 2D', title=\"TF-IDF visualization via t-SNE\", color='discourse_type')\nscatterplot(data, col='TF-IDF 2D', title=\"TF-IDF visualization via t-SNE\", color='discourse_effectiveness')","metadata":{"execution":{"iopub.status.busy":"2022-07-09T10:40:18.376402Z","iopub.execute_input":"2022-07-09T10:40:18.377077Z","iopub.status.idle":"2022-07-09T10:45:38.417071Z","shell.execute_reply.started":"2022-07-09T10:40:18.377041Z","shell.execute_reply":"2022-07-09T10:45:38.415966Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"6.3\"></a>\n\n#### <b>6.3 <span style='color:#4a8fdd'>|</span> Universal Sentence Encoder</b>\n\nThe <i>Universal Sentence Encoder</i> is a word embedding which represent words and sentences as real vectors of size 512, while maintaining semantic meaning as algebraic properties.","metadata":{}},{"cell_type":"code","source":"import tensorflow\nimport tensorflow_hub as hub\n\nuse_vectorizer = hub.load(\"https://tfhub.dev/google/universal-sentence-encoder/4\")\ndata['USE'] = data['clean_text'].apply(lambda x: use_vectorizer([x]).numpy()[0])\nlist_vector.append('USE')\ndata.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T10:45:38.418957Z","iopub.execute_input":"2022-07-09T10:45:38.419271Z","iopub.status.idle":"2022-07-09T10:47:42.465767Z","shell.execute_reply.started":"2022-07-09T10:45:38.41924Z","shell.execute_reply":"2022-07-09T10:47:42.464523Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data['USE 2D'] = tsne(data['USE'])\nscatterplot(data, col='USE 2D', title=\"USE visualization via t-SNE\", color='discourse_type')\nscatterplot(data, col='USE 2D', title=\"USE visualization via t-SNE\", color='discourse_effectiveness')","metadata":{"execution":{"iopub.status.busy":"2022-07-09T10:47:42.467518Z","iopub.execute_input":"2022-07-09T10:47:42.468245Z","iopub.status.idle":"2022-07-09T10:52:17.63551Z","shell.execute_reply.started":"2022-07-09T10:47:42.468196Z","shell.execute_reply":"2022-07-09T10:52:17.63451Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"6.4\"></a>\n\n#### <b>6.4 <span style='color:#4a8fdd'>|</span> BERT</b>\n\nThe <i>BERT</i> algorithm is another word embedding, whose vectorizations are real vectors of length 768.","metadata":{}},{"cell_type":"code","source":"import torch\nfrom transformers import BertTokenizer, BertModel\n\ndef format_sentences_BERT(text):\n    phrases = re.findall(r'[^\\.\\!\\?]*[\\.\\!\\?]', text)\n    return [\"[CLS] \" + sentence + \" [SEP]\" for sentence in phrases]\n\ndef preprocess_sentences_BERT(phrase, tokenizer, MAX_LEN=512):\n    temp = tokenizer.tokenize(phrase)\n    ids = [tokenizer.convert_tokens_to_ids(token) for token in temp]\n    if len(ids) > MAX_LEN:\n        ids = ids[:MAX_LEN]\n    return ids\n\ndef vectorize_sentence_BERT(input_ids, model):\n    tokens_tensor = torch.tensor([input_ids])\n    segments_tensors = torch.tensor([[1] * len(input_ids)])\n    with torch.no_grad():\n        outputs = model(tokens_tensor, segments_tensors)\n        hidden_states = outputs[2]\n    sentence_embedding = torch.mean(hidden_states[-2][0], dim=0)\n    return sentence_embedding.numpy()\n\ndef vectorize_text_BERT(text, model, tokenizer):\n    phrases = format_sentences_BERT(text)\n    vectors = [vectorize_sentence_BERT(preprocess_sentences_BERT(phrase, tokenizer), model) for phrase in phrases]\n    if len(vectors) > 0:\n        result = np.mean(vectors, axis=0)\n    else:\n        result = np.zeros(768)\n    return result","metadata":{"execution":{"iopub.status.busy":"2022-07-09T10:52:17.636942Z","iopub.execute_input":"2022-07-09T10:52:17.637252Z","iopub.status.idle":"2022-07-09T10:52:19.366674Z","shell.execute_reply.started":"2022-07-09T10:52:17.637224Z","shell.execute_reply":"2022-07-09T10:52:19.365684Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tokenizer = BertTokenizer.from_pretrained('bert-base-cased', do_lower_case=True)\nmodel = BertModel.from_pretrained('bert-base-cased', output_hidden_states = True)\nmodel.eval()\ndata['BERT'] = data['clean_text'].apply(lambda x: vectorize_text_BERT(x, model, tokenizer))\nlist_vector.append('BERT')\ndata.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T10:52:19.36816Z","iopub.execute_input":"2022-07-09T10:52:19.369086Z","iopub.status.idle":"2022-07-09T12:19:15.177539Z","shell.execute_reply.started":"2022-07-09T10:52:19.36904Z","shell.execute_reply":"2022-07-09T12:19:15.176481Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data['BERT 2D'] = tsne(data['BERT'])\nscatterplot(data, col='BERT 2D', title=\"BERT visualization via t-SNE\", color='discourse_type')\nscatterplot(data, col='BERT 2D', title=\"BERT visualization via t-SNE\", color='discourse_effectiveness')","metadata":{"execution":{"iopub.status.busy":"2022-07-09T12:19:15.178843Z","iopub.execute_input":"2022-07-09T12:19:15.179541Z","iopub.status.idle":"2022-07-09T12:24:56.669946Z","shell.execute_reply.started":"2022-07-09T12:19:15.179511Z","shell.execute_reply":"2022-07-09T12:24:56.668706Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"7\"></a>\n\n## <b>7 <span style='color:#4a8fdd'>|</span> CLASSIFIERS COMPARISON</b>","metadata":{}},{"cell_type":"markdown","source":"<a id=\"7.1\"></a>\n\n#### <b>7.1 <span style='color:#4a8fdd'>|</span> Evaluation</b>\n\nWe will use the <i>log loss error</i> to compare different classifier in order to select the most relevant one.","metadata":{}},{"cell_type":"code","source":"from sklearn import metrics\nimport timeit\n\ndef performances(model, X_train, y_train, X_test, y_test):\n    y_simul = model.predict_proba(X_train)\n    y_pred = model.predict_proba(X_test)    \n    logloss_train = metrics.log_loss(y_train, y_simul)\n    logloss_test = metrics.log_loss(y_test, y_pred)\n    return (logloss_train, logloss_test)\n\ndef default_classifier_perf(classif, X_train, y_train, X_test, y_test):\n    start_time = timeit.default_timer()\n    classif.fit(X_train, y_train)\n    elapsed = timeit.default_timer() - start_time\n    scores = performances(classif, X_train, y_train, X_test, y_test)\n    return (scores[0], scores[1], elapsed)","metadata":{"execution":{"iopub.status.busy":"2022-07-09T12:24:56.671707Z","iopub.execute_input":"2022-07-09T12:24:56.672089Z","iopub.status.idle":"2022-07-09T12:24:56.680393Z","shell.execute_reply.started":"2022-07-09T12:24:56.672054Z","shell.execute_reply":"2022-07-09T12:24:56.679418Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"7.2\"></a>\n\n#### <b>7.2 <span style='color:#4a8fdd'>|</span> Classifiers list</b>\n\nWe will use four different classifiers: <i>Logistic Regression</i>, <i>Random Forest</i> classifier, <i>Multi-Layer Perceptron</i> and <i>XGBoost</i> classifier. Each of them will be used on <b>every</b> possible vectorization, for a total of 12 combinations.","metadata":{}},{"cell_type":"code","source":"from sklearn.svm import LinearSVC\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.ensemble import RandomForestClassifier\nfrom sklearn.neural_network import MLPClassifier\nfrom xgboost import XGBClassifier\n\ndict_classifiers = {'Logistic Regression' : (LogisticRegression, {'random_state':42, 'n_jobs':-1}),\n                    'Random Forest' : (RandomForestClassifier, {'random_state':42, 'n_jobs':-1}),\n                    'MLP' : (MLPClassifier, {'random_state':42, 'early_stopping':True}),\n                    'XGBoost' : (XGBClassifier, {'objective':'multi:softproba', 'eval_metric':'merror',\n                                                            'num_class':3, 'seed':42, 'n_jobs':-1})}","metadata":{"execution":{"iopub.status.busy":"2022-07-09T12:24:56.685852Z","iopub.execute_input":"2022-07-09T12:24:56.687047Z","iopub.status.idle":"2022-07-09T12:24:56.906709Z","shell.execute_reply.started":"2022-07-09T12:24:56.686994Z","shell.execute_reply":"2022-07-09T12:24:56.905547Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"7.3\"></a>\n\n#### <b>7.3 <span style='color:#4a8fdd'>|</span> Split Train/Test</b>\n\nWe will split our data into two disjoint sets: one for training the model (80% of the entries) and the other for testing purposes (the remaining 20%).","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split\n\nlist_vector = ['TF-IDF', 'USE', 'BERT']\ndata_train, data_test = train_test_split(data, stratify=data['discourse_type'], test_size=.2, random_state=42)","metadata":{"execution":{"iopub.status.busy":"2022-07-09T12:46:15.20546Z","iopub.execute_input":"2022-07-09T12:46:15.205879Z","iopub.status.idle":"2022-07-09T12:46:15.318528Z","shell.execute_reply.started":"2022-07-09T12:46:15.205838Z","shell.execute_reply":"2022-07-09T12:46:15.317429Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"7.4\"></a>\n\n#### <b>7.4 <span style='color:#4a8fdd'>|</span> Baseline</b>\n\nWe need a dummy baseline to check if our different approches are relevant. We will use the <i>TF-IDF</i> vectorization along with a <i>Multinomial Naive Bayes</i> classifier.","metadata":{}},{"cell_type":"code","source":"from sklearn.naive_bayes import MultinomialNB\nfrom sklearn.preprocessing import LabelEncoder\n\nperfs = pd.DataFrame(index = ['train log loss', 'test log loss', 'computation time'])\n\nX_train = pd.DataFrame(data_train['TF-IDF'].dropna().tolist())\nX_test = pd.DataFrame(data_test['TF-IDF'].dropna().tolist())\nLE = LabelEncoder()\ny_train = LE.fit_transform(data_train['discourse_effectiveness'].iloc[X_train.index])\ny_test = LE.fit_transform(data['discourse_effectiveness'].iloc[X_test.index])\n\nclassif = MultinomialNB()\nresults = default_classifier_perf(classif, X_train, y_train, X_test, y_test)\nperfs['Baseline'] = [results[0], results[1], results[2]]\nprint(perfs)","metadata":{"execution":{"iopub.status.busy":"2022-07-09T12:47:36.585519Z","iopub.execute_input":"2022-07-09T12:47:36.585912Z","iopub.status.idle":"2022-07-09T12:47:51.022537Z","shell.execute_reply.started":"2022-07-09T12:47:36.585882Z","shell.execute_reply":"2022-07-09T12:47:51.021359Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"7.5\"></a>\n\n#### <b>7.5 <span style='color:#4a8fdd'>|</span> Models comparison</b>\n\nWe will compare the basic performances of every couple <i>vectorization/classifier</i> on our dataset, based on the <i>log loss error</i> and the required <i>computation time</i>. We shall then select the most relevant model for optimization and submission.","metadata":{}},{"cell_type":"code","source":"def compare_classifiers(data_train, data_test, col_y, list_vector, dict_classifiers):\n    \n    perform = pd.DataFrame(index = ['train log loss', 'test log loss', 'computation time'])\n    \n    for vect in list_vector:\n        \n        X_train =  pd.DataFrame(data_train[vect].dropna().tolist())\n        X_test = pd.DataFrame(data_test[vect].dropna().tolist())\n        \n        LE = LabelEncoder()\n        y_train = LE.fit_transform(data_train[col_y].iloc[X_train.index])\n        y_test = LE.transform(data_test[col_y].iloc[X_test.index])\n    \n    \n        for key_c, item_c in dict_classifiers.items():\n            classif = item_c[0](**item_c[1])\n            results = default_classifier_perf(classif, X_train, y_train, X_test, y_test)\n            \n            perform[vect + ' + ' + key_c] = results\n        \n    return perform\n\ndef plot_performances(perform, figsize=(12,7)):\n    perform.transpose().plot(kind='bar',secondary_y='computation time', figsize=figsize)\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T12:51:30.626251Z","iopub.execute_input":"2022-07-09T12:51:30.626708Z","iopub.status.idle":"2022-07-09T12:51:30.637541Z","shell.execute_reply.started":"2022-07-09T12:51:30.62667Z","shell.execute_reply":"2022-07-09T12:51:30.636586Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"perfs = compare_classifiers(data_train, data_test, \"discourse_effectiveness\", list_vector, dict_classifiers)\nplot_performances(perfs)","metadata":{"execution":{"iopub.status.busy":"2022-07-09T12:51:32.847875Z","iopub.execute_input":"2022-07-09T12:51:32.848879Z","iopub.status.idle":"2022-07-09T13:16:44.534632Z","shell.execute_reply.started":"2022-07-09T12:51:32.848839Z","shell.execute_reply":"2022-07-09T13:16:44.532235Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(perfs.transpose()['test log loss'].sort_values())","metadata":{"execution":{"iopub.status.busy":"2022-07-09T13:16:44.536522Z","iopub.execute_input":"2022-07-09T13:16:44.536812Z","iopub.status.idle":"2022-07-09T13:16:44.546599Z","shell.execute_reply.started":"2022-07-09T13:16:44.536785Z","shell.execute_reply":"2022-07-09T13:16:44.54542Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"7.6\"></a>\n\n#### <b>7.6 <span style='color:#4a8fdd'>|</span> Cross-Validation</b>\n\nWe will optimise some hyperparameters of the selected model via the <i>Cross-Validation</i> method. We also list ranges of parameters for other models for further works.","metadata":{}},{"cell_type":"code","source":"parameters_MLP = {\n    'hidden_layer_sizes' : [(50,50,50), (50,100,50), (100,)],\n    'activation' : ['logistic', 'relu'],\n    'solver': ['adam'],\n    'alpha' : [0.0001, 0.05],\n    'random_state' : [42],\n    'early_stopping' : [True],\n    'validation_fraction' : [0.1],\n}\n\nparameters_LR = {\n    'penalty' : ['l2'],\n    'class_weight' : [None, 'balanced'],\n    'random_state' : [42],\n    'solver' : ['saga', 'lbfgs'],\n    'multi_class' : ['auto'],\n    'n_jobs' : [-1],\n}\n\nparameters_RF = {\n    'n_estimators' : [50, 100, 200],\n    'max_features' : [\"sqrt\", \"log2\"],\n    'n_jobs' : [-1],\n    'random_state' : [42],\n    'class_weight' : [None, 'balanced'],\n}","metadata":{"execution":{"iopub.status.busy":"2022-07-09T13:17:29.50523Z","iopub.execute_input":"2022-07-09T13:17:29.505657Z","iopub.status.idle":"2022-07-09T13:17:29.513465Z","shell.execute_reply.started":"2022-07-09T13:17:29.505621Z","shell.execute_reply":"2022-07-09T13:17:29.512558Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.model_selection import GridSearchCV\n\ndef select_data(data_train, data_test, col_y, vect):\n    X_train = pd.DataFrame(data_train[vect].dropna().tolist())\n    X_test = pd.DataFrame(data_test[vect].dropna().tolist())\n    LE = LabelEncoder()\n    y_train = LE.fit_transform(data_train[col_y].iloc[X_train.index])\n    y_test = LE.transform(data_test[col_y].iloc[X_test.index])\n    return (X_train, y_train, X_test, y_test)\n\ndef cross_validation(X_train, y_train, classif, params):\n    cross_val = GridSearchCV(classif(), params, scoring='neg_log_loss', n_jobs=-1)\n    cross_val.fit(X_train, y_train)\n    return cross_val.best_params_","metadata":{"execution":{"iopub.status.busy":"2022-07-09T13:17:32.631688Z","iopub.execute_input":"2022-07-09T13:17:32.632122Z","iopub.status.idle":"2022-07-09T13:17:32.640275Z","shell.execute_reply.started":"2022-07-09T13:17:32.632084Z","shell.execute_reply":"2022-07-09T13:17:32.639069Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"(X_train, y_train, X_test, y_test) = select_data(data_train, data_test, 'discourse_effectiveness', 'USE')\nglobal_model_best_params = cross_validation(X_train, y_train, MLPClassifier, parameters_MLP)","metadata":{"execution":{"iopub.status.busy":"2022-07-09T13:17:39.222199Z","iopub.execute_input":"2022-07-09T13:17:39.222625Z","iopub.status.idle":"2022-07-09T13:22:48.01052Z","shell.execute_reply.started":"2022-07-09T13:17:39.22259Z","shell.execute_reply":"2022-07-09T13:22:48.008656Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we simply construct the <i>optimal</i> global model, and check its performances.","metadata":{}},{"cell_type":"code","source":"def optimal_model(X_train, y_train, X_test, y_test, classif, params):\n    model = classif(**params)\n    model.fit(X_train, y_train)\n    perform = performances(model, X_train, y_train, X_test, y_test)\n    return (model, perform)","metadata":{"execution":{"iopub.status.busy":"2022-07-09T13:27:56.377798Z","iopub.execute_input":"2022-07-09T13:27:56.378329Z","iopub.status.idle":"2022-07-09T13:27:56.385621Z","shell.execute_reply.started":"2022-07-09T13:27:56.378284Z","shell.execute_reply":"2022-07-09T13:27:56.384366Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"(global_model, scores) = optimal_model(X_train, y_train, X_test, y_test, MLPClassifier, global_model_best_params)\nprint(\"The global model gets a test log loss of {:.5f}\".format(scores[1]))","metadata":{"execution":{"iopub.status.busy":"2022-07-09T13:29:05.50318Z","iopub.execute_input":"2022-07-09T13:29:05.503607Z","iopub.status.idle":"2022-07-09T13:30:01.896848Z","shell.execute_reply.started":"2022-07-09T13:29:05.503569Z","shell.execute_reply":"2022-07-09T13:30:01.895431Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"8\"></a>\n\n## <b>8 <span style='color:#4a8fdd'>|</span> SPLITTING TYPES</b>\n\nThe texts we have to classify belong to different types (<i>opening</i>, <i>claim</i>, etc) and the optimal approach may vary depending on the considered discourse type.","metadata":{}},{"cell_type":"markdown","source":"<a id=\"8.1\"></a>\n\n#### <b>8.1 <span style='color:#4a8fdd'>|</span> Models comparison</b>\n\nWe will split the dataset into the seven different subset based on the <i>discourse_type</i> variable, before comparing once again the basic performances of every couple <i>vectorization/classifier</i> on <b>each data subset</b>!","metadata":{}},{"cell_type":"code","source":"perfs = pd.DataFrame(columns = discourse_types)\nfor i in discourse_types:\n    data_train_temp = data_train[data_train['discourse_type'] == i]\n    data_test_temp = data_test[data_test['discourse_type'] == i]\n    perfs[i] = compare_classifiers(data_train_temp, data_test_temp, 'discourse_effectiveness',\n                                   list_vector, dict_classifiers).loc['test log loss']\nperfs.plot(kind='bar', figsize=(12,7))\nperfs.transpose().plot(kind='bar', figsize=(12,7))","metadata":{"execution":{"iopub.status.busy":"2022-07-09T13:30:01.899216Z","iopub.execute_input":"2022-07-09T13:30:01.907922Z","iopub.status.idle":"2022-07-09T13:53:18.371961Z","shell.execute_reply.started":"2022-07-09T13:30:01.90785Z","shell.execute_reply":"2022-07-09T13:53:18.370912Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i in discourse_types:\n    print(i, \":\", perfs[i].sort_values().index[0])\n    print(\"    log loss: {:.5f}\".format(perfs[i].sort_values().iloc[0]))","metadata":{"execution":{"iopub.status.busy":"2022-07-09T13:57:46.291624Z","iopub.execute_input":"2022-07-09T13:57:46.292112Z","iopub.status.idle":"2022-07-09T13:57:46.303361Z","shell.execute_reply.started":"2022-07-09T13:57:46.292073Z","shell.execute_reply":"2022-07-09T13:57:46.302024Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"8.2\"></a>\n\n#### <b>8.2 <span style='color:#4a8fdd'>|</span> Cross-Validation</b>\n\nAgain we optimise the hyperparameters of the selected models via the <i>Cross-Validation</i> method, for <b>every</b> discourse type.","metadata":{}},{"cell_type":"code","source":"dict_best_combinations = {\n    'Evidence' : ('USE', LogisticRegression, parameters_LR),\n    'Lead' : ('BERT', RandomForestClassifier, parameters_RF),\n    'Position' : ('USE', LogisticRegression, parameters_LR),\n    'Counterclaim' : ('USE', LogisticRegression, parameters_LR),\n    'Claim' : ('USE', MLPClassifier, parameters_MLP),\n    'Rebuttal' : ('USE', LogisticRegression, parameters_LR),\n    'Concluding Statement' : ('USE', LogisticRegression, parameters_LR),\n}","metadata":{"execution":{"iopub.status.busy":"2022-07-09T14:11:10.496923Z","iopub.execute_input":"2022-07-09T14:11:10.498337Z","iopub.status.idle":"2022-07-09T14:11:10.504481Z","shell.execute_reply.started":"2022-07-09T14:11:10.498294Z","shell.execute_reply":"2022-07-09T14:11:10.503379Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"split_parameters = {}\nsplit_models = {}\nsplit_scores = {}\n\nfor i in discourse_types:\n    data_train_temp = data_train[data_train['discourse_type'] == i]\n    data_test_temp = data_test[data_test['discourse_type'] == i]\n    (X_train, y_train, X_test, y_test) = select_data(data_train_temp, data_test_temp,\n                                                     'discourse_effectiveness', dict_best_combinations[i][0])\n    split_parameters[i] = cross_validation(X_train, y_train, dict_best_combinations[i][1], dict_best_combinations[i][2])\n    (split_models[i], split_scores[i]) = optimal_model(X_train, y_train, X_test, y_test,\n                                    dict_best_combinations[i][1], split_parameters[i])\n    print(\"The split model for \", i, \" gets a test log loss of {:.5f}\".format(split_scores[i][1]))\n","metadata":{"execution":{"iopub.status.busy":"2022-07-09T14:41:29.630543Z","iopub.execute_input":"2022-07-09T14:41:29.631426Z","iopub.status.idle":"2022-07-09T14:44:32.489Z","shell.execute_reply.started":"2022-07-09T14:41:29.631383Z","shell.execute_reply":"2022-07-09T14:44:32.487586Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"8.3\"></a>\n\n#### <b>8.3 <span style='color:#4a8fdd'>|</span> Overall performances</b>\n\nIn order to compare them with the global model, we need to compute the log loss error <b>over the entire dataset</b> when we select the prediction model depending on the discourse type.","metadata":{}},{"cell_type":"code","source":"def split_models_predict_proba(X, col_type, dict_combinations, trained_models):\n    vect_method = dict_combinations[X[col_type]][0]\n    prediction = trained_models[X[col_type]].predict_proba([X[vect_method]])[0]\n    return prediction","metadata":{"execution":{"iopub.status.busy":"2022-07-09T14:44:32.492032Z","iopub.execute_input":"2022-07-09T14:44:32.492846Z","iopub.status.idle":"2022-07-09T14:44:32.500571Z","shell.execute_reply.started":"2022-07-09T14:44:32.492797Z","shell.execute_reply":"2022-07-09T14:44:32.499394Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_pred_split = data_test.apply(lambda x : split_models_predict_proba(x, 'discourse_type', dict_best_combinations, split_models), axis=1)\n\nLE = LabelEncoder()\ny_train = LE.fit_transform(data_train['discourse_effectiveness'])\ny_test = LE.transform(data_test['discourse_effectiveness'])\n\nsplit_overall_test_score = metrics.log_loss(y_test, np.array(y_pred_split.tolist()))\nprint(\"The split model gets an overall test log loss of {:.5f}\".format(split_overall_test_score))","metadata":{"execution":{"iopub.status.busy":"2022-07-09T14:44:32.502682Z","iopub.execute_input":"2022-07-09T14:44:32.503434Z","iopub.status.idle":"2022-07-09T14:45:23.049029Z","shell.execute_reply.started":"2022-07-09T14:44:32.503391Z","shell.execute_reply":"2022-07-09T14:45:23.047816Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The split model's loss is lower than the global model's one (0.77919) which means we get higher predictive power by splitting the training sets depending on the discourse types.","metadata":{}},{"cell_type":"markdown","source":"<a id=\"9\"></a>\n\n## <b>9 <span style='color:#4a8fdd'>|</span> MIXING THE MODELS</b>\n\n<a id=\"9.1\"></a>\n\n#### <b>9.1 <span style='color:#4a8fdd'>|</span> Data subsets comparison</b>\n\nWe will check if the split model is better than the global model on its own restricted subset. If not, we will replace it by the global model (trained with <b>all</b> discourse types).","metadata":{}},{"cell_type":"code","source":"for i in discourse_types:\n    data_train_temp = data_train[data_train['discourse_type'] == i]\n    data_test_temp = data_test[data_test['discourse_type'] == i]\n    (X_train, y_train, X_test, y_test) = select_data(data_train_temp, data_test_temp, 'discourse_effectiveness', 'USE')\n    y_pred_temp = global_model.predict_proba(X_test)\n    temp_score = metrics.log_loss(y_test, y_pred_temp)\n    print(\"The global model restricted to \", i, \" data gets a test log loss of {:.5f}\".format(temp_score))\n    if temp_score < split_scores[i][1]:\n        print(\"      The corresponding split model had {:.5f} which is GREATER\".format(split_scores[i][1]))\n    else:\n        print(\"      The corresponding split model had {:.5f}\".format(split_scores[i][1]))","metadata":{"execution":{"iopub.status.busy":"2022-07-09T14:45:23.052058Z","iopub.execute_input":"2022-07-09T14:45:23.052935Z","iopub.status.idle":"2022-07-09T14:45:31.429742Z","shell.execute_reply.started":"2022-07-09T14:45:23.052886Z","shell.execute_reply":"2022-07-09T14:45:31.426269Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"9.2\"></a>\n\n#### <b>9.2 <span style='color:#4a8fdd'>|</span> Performances</b>\n\nWe will use the global model to predict the effectiveness of <i>Rebuttal</i> texts, and use the specific optimal algorithm for the other discourse types, before computing the log loss error on our test dataset for this brand new model.","metadata":{}},{"cell_type":"code","source":"def mixed_model_predict_proba(X, col_type, list_types_global, trained_models, dict_combinations):\n    discourse_type = X[col_type]\n    if discourse_type in list_types_global:\n        discourse_type = 'Global'\n    vect_method = dict_combinations[discourse_type][0]\n    prediction = trained_models[discourse_type].predict_proba([X[vect_method]])[0]\n    return prediction","metadata":{"execution":{"iopub.status.busy":"2022-07-09T14:50:19.362813Z","iopub.execute_input":"2022-07-09T14:50:19.363254Z","iopub.status.idle":"2022-07-09T14:50:19.370232Z","shell.execute_reply.started":"2022-07-09T14:50:19.363221Z","shell.execute_reply":"2022-07-09T14:50:19.369066Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"list_types_global = ['Rebuttal']\nsplit_models['Global'] = global_model\ndict_best_combinations['Global'] = ('USE', MLPClassifier, parameters_MLP)\n\ny_pred_mixed = data_test.apply(lambda x : mixed_model_predict_proba(x, 'discourse_type', list_types_global,\n                                                                    split_models, dict_best_combinations), axis=1)\n\nLE = LabelEncoder()\ny_train = LE.fit_transform(data_train['discourse_effectiveness'])\ny_test = LE.transform(data_test['discourse_effectiveness'])\n\nmixed_overall_test_score = metrics.log_loss(y_test, np.array(y_pred_mixed.tolist()))\nprint(\"The split model gets an overall test log loss of {:.5f}\".format(mixed_overall_test_score))","metadata":{"execution":{"iopub.status.busy":"2022-07-09T14:53:17.207702Z","iopub.execute_input":"2022-07-09T14:53:17.208615Z","iopub.status.idle":"2022-07-09T14:54:07.443361Z","shell.execute_reply.started":"2022-07-09T14:53:17.208574Z","shell.execute_reply":"2022-07-09T14:54:07.439708Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The mixed model has a log loss error of 0.75140, which is lower than the global model's error 0.77919 and the split model's 0.75253.","metadata":{}},{"cell_type":"markdown","source":"<a id=\"9.3\"></a>\n\n#### <b>9.3 <span style='color:#4a8fdd'>|</span> Saving the parameters</b>\n\nWe will save the optimal parameters into a dictionary for the next section.","metadata":{}},{"cell_type":"code","source":"mixed_model_parameters = {}\nmixed_model_parameters['Global'] = (True, dict_best_combinations['Global'][0],\n                                    dict_best_combinations['Global'][1], global_model_best_params)\n\nfor i in discourse_types:\n    if i in list_types_global:\n        mixed_model_parameters[i] = mixed_model_parameters['Global']\n    else:\n        is_global = False\n        vectorization_method = dict_best_combinations[i][0]\n        classifier = dict_best_combinations[i][1]\n        optimal_parameters = split_parameters[i]\n        mixed_model_parameters[i] = (is_global, vectorization_method, classifier, optimal_parameters)","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:11:58.08207Z","iopub.execute_input":"2022-07-09T15:11:58.082952Z","iopub.status.idle":"2022-07-09T15:11:58.095541Z","shell.execute_reply.started":"2022-07-09T15:11:58.082899Z","shell.execute_reply":"2022-07-09T15:11:58.093593Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"10\"></a>\n\n## <b>10 <span style='color:#4a8fdd'>|</span> MAKING THE PREDICTIONS</b>\n\n<a id=\"10.1\"></a>\n\n#### <b>10.1 <span style='color:#4a8fdd'>|</span> Training the model</b>\n\nNow we need to train our mixed model over all available data, according to the required subset splits.","metadata":{}},{"cell_type":"code","source":"dict_trained_models = {}\n    \nX_train_global = pd.DataFrame(data[mixed_model_parameters['Global'][1]].dropna().tolist())\nLE = LabelEncoder()\ny_train_global = LE.fit_transform(data['discourse_effectiveness'].iloc[X_train_global.index])\n    \ndict_trained_models['Global'] = mixed_model_parameters['Global'][2](**mixed_model_parameters['Global'][3]).fit(X_train_global, y_train_global)\n    \nfor i in discourse_types:\n    if mixed_model_parameters[i][0]:\n        dict_trained_models[i] = dict_trained_models['Global']\n    else :\n        X_train_local = pd.DataFrame(data[mixed_model_parameters[i][1]].dropna().tolist())\n        y_train_local = LE.transform(data['discourse_effectiveness'].iloc[X_train_local.index])\n        \n        dict_trained_models[i] = mixed_model_parameters[i][2](**mixed_model_parameters[i][3]).fit(X_train_local, y_train_local)","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:39:45.707581Z","iopub.execute_input":"2022-07-09T15:39:45.708254Z","iopub.status.idle":"2022-07-09T15:45:51.716902Z","shell.execute_reply.started":"2022-07-09T15:39:45.70821Z","shell.execute_reply":"2022-07-09T15:45:51.715339Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"10.2\"></a>\n\n#### <b>10.2 <span style='color:#4a8fdd'>|</span> Vectorizing test data</b>\n","metadata":{}},{"cell_type":"code","source":"test['clean_text'] = cleaning(test['discourse_text'])\ntest['cleaner_text'] = more_cleaning(test[\"clean_text\"])\n\ndef tfidf_transform(tfidf_vectorizer, text_data):\n    if not isinstance(text_data.iloc[0], list):\n        text_data = tokenize(text_data)\n    tfidf_vectors_csr = tfidf_vectorizer.transform(text_data)\n    return (pd.Series(tfidf_vectors_csr.todense().tolist(), text_data.index))\n\ntest['TF-IDF'] = tfidf_transform(tfidf_vectorizer, test['cleaner_text'])\ntest['USE'] = test['clean_text'].apply(lambda x: use_vectorizer([x]).numpy()[0])\ntest['BERT'] = test['clean_text'].apply(lambda x: vectorize_text_BERT(x, model, tokenizer))\n\ntest.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:33:57.374918Z","iopub.execute_input":"2022-07-09T15:33:57.375489Z","iopub.status.idle":"2022-07-09T15:33:59.420877Z","shell.execute_reply.started":"2022-07-09T15:33:57.375444Z","shell.execute_reply":"2022-07-09T15:33:59.419768Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"10.3\"></a>\n\n#### <b>10.3 <span style='color:#4a8fdd'>|</span> Generating predictions</b>\n\nWe predict the discourse effectiveness of the test dataset.","metadata":{}},{"cell_type":"code","source":"def trained_mixed_model_predict_proba(X, col_type, mixed_model_parameters, dict_trained_models):\n    discourse_type = X[col_type]\n    vect_method = mixed_model_parameters[discourse_type][1]\n    prediction = dict_trained_models[discourse_type].predict_proba([X[vect_method]])[0]\n    return prediction\n\ntest['prediction'] = test.apply(lambda x : trained_mixed_model_predict_proba(x, 'discourse_type', mixed_model_parameters,\n                                                                     dict_trained_models), axis=1)\n\ntest.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:45:52.32077Z","iopub.execute_input":"2022-07-09T15:45:52.321614Z","iopub.status.idle":"2022-07-09T15:45:52.494577Z","shell.execute_reply.started":"2022-07-09T15:45:52.321552Z","shell.execute_reply":"2022-07-09T15:45:52.493336Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"labels = LE.inverse_transform([0,1,2])\npredictions = pd.DataFrame(test['prediction'].tolist(), columns = labels, index = test.index)\npredictions['discourse_id'] = test['discourse_id']\npredictions.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:51:41.395889Z","iopub.execute_input":"2022-07-09T15:51:41.396876Z","iopub.status.idle":"2022-07-09T15:51:41.411523Z","shell.execute_reply.started":"2022-07-09T15:51:41.396834Z","shell.execute_reply":"2022-07-09T15:51:41.410401Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predictions[['discourse_id', 'Ineffective', 'Adequate', 'Effective']].to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-09T15:55:28.453085Z","iopub.execute_input":"2022-07-09T15:55:28.453546Z","iopub.status.idle":"2022-07-09T15:55:28.470049Z","shell.execute_reply.started":"2022-07-09T15:55:28.453482Z","shell.execute_reply":"2022-07-09T15:55:28.468677Z"},"trusted":true},"execution_count":null,"outputs":[]}]}