{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# **Notebook: Starting with EDA + Logistic model**\n\nThe objectives of the notebook are:\n\n* Have a better understanding of the **data** we are working with\n* Understand better the **interaction** between the different **variables** and the **discourse effectiveness**\n* Explore new **features** that can be included in a model to identify the **effectiveness of the discourse**.\n* Fit different Logistic Regressions\n* Pick the best model\n* A guide of Conclusions and Next Steps for ameliorating the model \n* Make a submission","metadata":{}},{"cell_type":"markdown","source":"# Import Libraries","metadata":{}},{"cell_type":"code","source":"import os\nfrom os.path import join \n\nimport pandas as pd\nimport numpy as np\n\n\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n","metadata":{"execution":{"iopub.status.busy":"2022-07-20T10:26:23.416103Z","iopub.execute_input":"2022-07-20T10:26:23.416464Z","iopub.status.idle":"2022-07-20T10:26:23.423097Z","shell.execute_reply.started":"2022-07-20T10:26:23.416434Z","shell.execute_reply":"2022-07-20T10:26:23.421469Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Import Data\n\n* import train data for training model\n* import test data for submit predictions","metadata":{}},{"cell_type":"code","source":"TRAIN_ESSAYS_PATH = \"../input/feedback-prize-effectiveness/train\"\nTEST_ESSAYS_PATH = \"../input/feedback-prize-effectiveness/test\"\n\ntrain = pd.read_csv(\"../input/feedback-prize-effectiveness/train.csv\")\ntest = pd.read_csv(\"../input/feedback-prize-effectiveness/test.csv\")\n","metadata":{"execution":{"iopub.status.busy":"2022-07-20T10:26:24.164783Z","iopub.execute_input":"2022-07-20T10:26:24.165519Z","iopub.status.idle":"2022-07-20T10:26:24.345712Z","shell.execute_reply.started":"2022-07-20T10:26:24.165479Z","shell.execute_reply":"2022-07-20T10:26:24.344602Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def get_essay(essay_id, train=True):\n    if train:\n        PATH = TRAIN_ESSAYS_PATH\n    else:\n        PATH = TEST_ESSAYS_PATH\n    essay_path = join(PATH, f\"{essay_id}.txt\")\n    essay_text = open(essay_path, 'r').read()\n    return essay_text","metadata":{"execution":{"iopub.status.busy":"2022-07-20T10:26:24.547473Z","iopub.execute_input":"2022-07-20T10:26:24.547876Z","iopub.status.idle":"2022-07-20T10:26:24.554455Z","shell.execute_reply.started":"2022-07-20T10:26:24.547841Z","shell.execute_reply":"2022-07-20T10:26:24.553153Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* **Goal of the competition**: Create a model that given a discourse text and a discourse type is able to predict if the discourse is **ineffective, adequate** or **effective**\n\n* What is a **discourse type**?: The data contains argumentative essays, each argumentative essay has different discourse elements commonly found in argumentative writing. The discourse types  are the different categories to which the discourse element can belong : \n    * Lead - an introduction that begins with a statistic, a quotation, a description, or some other device to grab the reader’s attention and point toward the thesis\n    * Position - an opinion or conclusion on the main question\n    * Claim - a claim that supports the position\n    * Counterclaim - a claim that refutes another claim or gives an opposing reason to the position\n    * Rebuttal - a claim that refutes a counterclaim\n    * Evidence - ideas or examples that support claims, counterclaims, or rebuttals.\n    * Concluding Statement - a concluding statement that restates the claims\n    \n* What is a **discourse text**? : A discourse text is an extract of the essay that matches one of the different discourse type categories.\n","metadata":{}},{"cell_type":"markdown","source":"# Exploratory Data Analysis","metadata":{}},{"cell_type":"markdown","source":"## 1. Describe columns of dataframe","metadata":{}},{"cell_type":"code","source":"train.head(5)","metadata":{"execution":{"iopub.status.busy":"2022-07-20T10:26:26.147362Z","iopub.execute_input":"2022-07-20T10:26:26.148405Z","iopub.status.idle":"2022-07-20T10:26:26.161881Z","shell.execute_reply.started":"2022-07-20T10:26:26.148359Z","shell.execute_reply":"2022-07-20T10:26:26.160515Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('the trainset has', train.shape[0], 'rows')\nprint('the trainset has', train.shape[1], 'columns')\nprint('\\n')\n\ntrain.describe()","metadata":{"execution":{"iopub.status.busy":"2022-07-20T10:26:26.911658Z","iopub.execute_input":"2022-07-20T10:26:26.912523Z","iopub.status.idle":"2022-07-20T10:26:27.006293Z","shell.execute_reply.started":"2022-07-20T10:26:26.912474Z","shell.execute_reply":"2022-07-20T10:26:27.005510Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* There are only 4191 different essays and 36691 different discourses (9 times more)\n* Few discourse text are repeated since there are 36765 rows and 36691 texts which means that we use the same discourse text for different discourse type","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"markdown","source":"## 2. Distribution of Variables","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport seaborn as sns\nsns.set()\nfig, ax = plt.subplots(nrows=1, ncols=2, figsize=(24, 6))\n\n\nsns.countplot(x=\"discourse_effectiveness\", data=train, order = ['Ineffective', 'Adequate', 'Effective'],\n              palette = ['lightcoral', 'lightyellow', 'lightgreen'], ax=ax[0])\n\nsns.countplot(x=\"discourse_type\", data=train,ax=ax[1])\n\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2022-07-20T10:23:32.375621Z","iopub.execute_input":"2022-07-20T10:23:32.376071Z","iopub.status.idle":"2022-07-20T10:23:32.975956Z","shell.execute_reply.started":"2022-07-20T10:23:32.376030Z","shell.execute_reply":"2022-07-20T10:23:32.974644Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* The target (i.e. discourse effectiveness) is highly unbalanced\n    * There are 6462 ineffective discourses\n    * There are 20977 adequate discourses\n    * There are 9326 effective discourses\n* The same happens for the discourse type\n    * There are 6 times more **Claims** or **Evidence** than **Counterclaims** discourses","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots(nrows=1, ncols=1, figsize=(20, 8))\n\nsns.countplot(x = 'discourse_effectiveness',\n            hue = 'discourse_type', order = ['Ineffective', 'Adequate', 'Effective'] ,  data = train)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-20T10:49:36.832799Z","iopub.execute_input":"2022-07-20T10:49:36.833203Z","iopub.status.idle":"2022-07-20T10:49:37.133206Z","shell.execute_reply.started":"2022-07-20T10:49:36.833167Z","shell.execute_reply":"2022-07-20T10:49:37.131800Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The distribution of **discouse types** varies for different **discourse effectiveness**:\n* For the Adequate and Effective it follows a similar distribution but in proportion there are more position discourses that are adequate\n* Few Claims are Ineffective \n","metadata":{}},{"cell_type":"markdown","source":"### Top expressions per discourse type and based on discourse efficiency\n\nLet's study the most common expressions depending of the **effectiveness and type** of the **discourse** ","metadata":{}},{"cell_type":"code","source":"from sklearn.feature_extraction.text import TfidfVectorizer\nimport spacy\nstop_words = list(spacy.load(\"en_core_web_sm\").Defaults.stop_words)\n\ndiscourses = np.unique(train.discourse_type.values)\n\n\nfor discourse in discourses:\n    print('For the discourse:', discourse,'\\n')\n    print('the top unigrams and bigrams are: \\n')\n    \n    title = '|Discourse effectiveness|Top unigrams|Top Bigrams|'\n    subtitle = '|:---:|:-----:|:---:|'\n    print(title)\n    print(subtitle)\n\n\n    for effectiveness in ['Ineffective', 'Adequate', 'Effective']:\n        vect1 = TfidfVectorizer(ngram_range=(1, 1)) #, stop_words=stop_words)\n        vect2 = TfidfVectorizer(ngram_range=(2, 2)) \n        docs = train.loc[(train.loc[:,'discourse_type'] ==discourse) & (train.loc[:,'discourse_effectiveness'] ==effectiveness),'discourse_text'].values\n        top_words = []\n        for vect in [vect1, vect2]:\n            tfidf = vect.fit_transform(docs)\n            sparse_row = vect.transform(['. '.join(docs)])\n            scores = dict(zip(vect.vocabulary_, np.array(sparse_row.sum(axis=0))[0]))\n            sorted_score = sorted(scores.items(), key=lambda x: x[1], reverse=True)\n            top_words.append(', '.join([x[0] for x in sorted_score[:10]]))\n        print(f'|{effectiveness}|{top_words[0]}|{top_words[1]}|')\n    print('\\n')\n    print('----------------------------------')\n","metadata":{"execution":{"iopub.status.busy":"2022-07-20T10:23:33.357045Z","iopub.execute_input":"2022-07-20T10:23:33.359864Z","iopub.status.idle":"2022-07-20T10:23:43.563214Z","shell.execute_reply.started":"2022-07-20T10:23:33.359811Z","shell.execute_reply":"2022-07-20T10:23:43.561757Z"},"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"-------------------------------------**Claim**-------------------------------------\n\n\n\nFor the discourse: **Claim**\n\nthe top unigrams and bigrams are: \n\n|Discourse effectiveness|Top unigrams|Top Bigrams|\n|:---:|:-----:|:---:|\n|Ineffective|easyer, suceceed, opportunitie, talks, hadle, wealth, affects, kitchen, werid, veterans|planet ourselves, likely to, future we, from what, show your, posters they, doing maybe, bound because, they tumbled, get material|\n|Adequate|hassle, socialization, convine, unfair, random, club, facal, tome, concluded, recks|the topic, key to, out states, assignment rather, ride would, go though, since it, choice will, when adding, phrase bigger|\n|Effective|lingering, dumb, widely, by, preparedness, proud, anywhere, employment, reports, senators|that attend, found student, but even, us who, easy transportation, point across, person teacher, help all, and succeed, schools everyone|\n\n\n-------------------------------------**Concluding Statement**-------------------------------------\n\nFor the discourse: **Concluding Statement**\n\nthe top unigrams and bigrams are: \n\n|Discourse effectiveness|Top unigrams|Top Bigrams|\n|:---:|:-----:|:---:|\n|Ineffective|upcoming, avantage, put, bother, staying, advancen, protection, deep, ansers, back|seconds is, educated but, fine by, some differents, clean up, have real, will win, my favor, are deficult, think all|\n|Adequate|irresponisble, arugued, dangerous, self, variable, dixoide, brought, eventhey, stick, issues|thing like, understand people, task by, now could, you grow, just driving, people understood, that much, but need, why most|\n|Effective|notifications, separate, falls, teaches, lonely, check, unappreciated, depends, creations, planned|learning would, forced into, the relevancy, home allows, unfair system, cut down, somebody asked, base on, emissions keep, friends he|\n\n\n-------------------------------------**Counterclaim**-------------------------------------\n\n\n\nFor the discourse: **Counterclaim**\n\nthe top unigrams and bigrams are: \n\n|Discourse effectiveness|Top unigrams|Top Bigrams|\n|:---:|:-----:|:---:|\n|Ineffective|read, everyones, helping, landforms, side, potentail, president, say, doing, hvsicallv|they brung, what he, in your, from home, states big, college they, usually stops, the school, have won, that was|\n|Adequate|across, claiming, struggle, candiates, appered, council, wont, choos, young, recived|geology and, make everyday, many other, or lower, have lack, could stop, that their, someone who, and don, its setup|\n|Effective|worked, significantly, tremendous, single, if, written, communicate, natural, risky, taken|can put, cars which, habit and, so student, higher grades, and certainty, in front, can convice, playing video, be seen|\n\n\n\n-------------------------------------**Evidence**-------------------------------------\n\n\nFor the discourse: **Evidence**\n\nthe top unigrams and bigrams are: \n\n|Discourse effectiveness|Top unigrams|Top Bigrams|\n|:---:|:-----:|:---:|\n|Ineffective|innovember, thna, episodes, familiar, lifeform, airplans, oberserved, advantageous, potatoes, skids|are thousands, anyone has, ho not, vinci used, is keep, that me, expanded ot, or working, unlikely nixon, food then|\n|Adequate|enforce, promgran, angery, hometown, hourly, themselfs, supposedlly, anddesine, ill, suited|out during, public transportaion, to adapt, could the, also texting, poor donating, people above, emissions behind, more depressed, without vehicle|\n|Effective|siphon, sloppiness, perception, overview, businesses, dimmer, professionally, sided, rougher, equivalent|places would, particularly knowing, losing papers, others lacking, darn morning, father lastly, bring you, allowing human, allows educators, see something|\n\n\n\n-------------------------------------**Lead**-------------------------------------\n\n\n\nFor the discourse: **Lead** \n\nthe top unigrams and bigrams are: \n\n|Discourse effectiveness|Top unigrams|Top Bigrams|\n|:---:|:-----:|:---:|\n|Ineffective|briveles, development, rocky, ready, cowboys, wight, communication, deal, scared, re|planet reaches, used to, also called, this in, in home, is waste, complete summer, terms venus, states over, comes from|\n|Adequate|id, conservative, similar, may, learnig, result, dreads, problaby, fare, exist|as option, off time, circle the, photos of, author talkes, fathers they, were about, community in, be rocks, because shadows|\n|Effective|carbon, differing, adhd, aren, giving, figment, somehow, drops, ponder, relief|one word, where ever, sticking with, of fatal, read the, two days, people coming, allows them, material before, determined by|\n\n\n\n-------------------------------------**Lead**-------------------------------------\n\n\n\nFor the discourse: **Position** \n\nthe top unigrams and bigrams are: \n\n|Discourse effectiveness|Top unigrams|Top Bigrams|\n|:---:|:-----:|:---:|\n|Ineffective|anachronism, simply, replace, road, widely, out, threw, get, belief, mentally|state president, technology read, scared there, here all, the changes, won want, citizens should, changed anything, it as, is wrong|\n|Adequate|fan, riding, promoting, trips, nose, astroid, recently, first, madeby, into|cars has, longer necessary, vote method, since our, impossible but, our teachers, benefit to, figure is, advice seeking, impossible with|\n|Effective|lunchtime, inaccurate, disastrous, teacher, enforcing, used, worse, backed, state, utilize|depicting who, corrosive to, phone use, not the, amazing in, start electing, with grade, skewed to, the entire, smart choices|\n\n\n-------------------------------------**Rebuttal**-------------------------------------\n\n\n\n\nFor the discourse: **Rebuttal** \n\nthe top unigrams and bigrams are: \n\n|Discourse effectiveness|Top unigrams|Top Bigrams|\n|:---:|:-----:|:---:|\n|Ineffective|hope, lunch, potential, counting, job, whether, cause, aliens, play, stand|about themselves, thee program, something like, interesting face, we like, isn some, that important, at use, of america, know their|\n|Adequate|uniforms, worthless, democrat, higher, material, americans, age, chosse, people, acient|can change, as well, formed in, beive the, have much, have problems, faulty system, then you, not least, and only|\n|Effective|widely, ties, typically, totally, handled, everyday, valid, agree, resolution, elect|class it, problems ahead, precise rather, prove meaningless, as students, not permitted, fabric can, doesnt mean, be tie, each have|\n\n\n----------------------------------\n\n\n\n**GENERAL OBSERVATIONS**:\n\n* Four all kinds of discourse type there is a clear difference between the choice of unigrams and bigrams depending on the effectiveness. The more effective is the discourse, the more sophisticated are the top words. \n* We do not see vocabulary specific to the discourse type\n* Regarding the discourse efficiency we observe:\n    * For the **Ineffective** discourse: \n        * we have many mispelling issues f.ex: *avantage,suceceed, easyer*\n        * demonstratives on bigrams: *that, those, these*\n        * adverbs\n        * some bigrams do not make sense\n    * For the **Adequate** discourse:\n        * some mispelling\n        * majority of the bigrams make sense\n    * For the **Effective** dicourse:\n        * sophisticated expressions presence of words like: *rather, depicting*\n        * More active verbs finishing in : *-ing*\n\n\n","metadata":{}},{"cell_type":"markdown","source":"## Variables Creation\n\nLet's study other particularities of the discourses that may have different distributions depending on the **effectiveness**","metadata":{}},{"cell_type":"code","source":"from nltk.corpus import words\nenglish_dict = dict.fromkeys(words.words(), None) #dictionnary of english words \n","metadata":{"execution":{"iopub.status.busy":"2022-07-20T10:54:27.917733Z","iopub.execute_input":"2022-07-20T10:54:27.918132Z","iopub.status.idle":"2022-07-20T10:54:28.047752Z","shell.execute_reply.started":"2022-07-20T10:54:27.918099Z","shell.execute_reply":"2022-07-20T10:54:28.046682Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import re\ntrain.loc[:,'len_discourse'] = train.loc[:,'discourse_text'].apply(lambda x: len(x))\ntrain.loc[:,'num_words'] = train.loc[:,'discourse_text'].apply(lambda x : len(x.split()))\ntrain.loc[:,'num_long_words'] = train.loc[:,'discourse_text'].apply(lambda x: len([word for word in x.split() if len(word)>6]))\ntrain.loc[:,'num_sentences'] = train.loc[:,'discourse_text'].apply(lambda x :len(x.replace('?','.').replace('!','.').split('.')))\ntrain.loc[:,'num_punctuation'] = train.loc[:,'discourse_text'].apply(lambda x: len(re.findall(r'[^\\w\\s]', x)))\ntrain.loc[:,'num_exlamations'] = train.loc[:,'discourse_text'].apply(lambda x :len(re.findall(r'\\!',x)))\ntrain.loc[:,'num_interrogations'] = train.loc[:,'discourse_text'].apply(lambda x :len(re.findall(r'\\?',x)))\ntrain.loc[:,'num_commas'] = train.loc[:,'discourse_text'].apply(lambda x :len(re.findall(',',x)))\ntrain.loc[:,'num_doublepoints'] = train.loc[:,'discourse_text'].apply(lambda x :len(re.findall(r'\\:',x)))\ntrain.loc[:,'num_digits'] = train.loc[:,'discourse_text'].apply(lambda x :len(re.findall('\\d+',x)))\ntrain.loc[:,'num_mispelling'] = train.loc[:,'discourse_text'].apply(lambda x: len([word for word in x.split() if word not in english_dict]))\n#train.loc[:,'ratio_mispelling'] = train.loc[:,'num_mispelling']/train.loc[:,'num_words']\nvariables = train.columns[5:]","metadata":{"execution":{"iopub.status.busy":"2022-07-20T10:54:31.047766Z","iopub.execute_input":"2022-07-20T10:54:31.048264Z","iopub.status.idle":"2022-07-20T10:54:32.462241Z","shell.execute_reply.started":"2022-07-20T10:54:31.048218Z","shell.execute_reply":"2022-07-20T10:54:32.461041Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(nrows=len(variables), ncols=1, figsize=(16, 36))\n\nfor k, variable in enumerate(variables):\n    sns.boxplot(data = train, \n            y = variable, \n            x='discourse_type', \n            hue='discourse_effectiveness', \n            hue_order = ['Ineffective', 'Adequate', 'Effective'], \n            palette = ['lightcoral', 'lightyellow', 'lightgreen'],\n            ax=ax[k])\n","metadata":{"execution":{"iopub.status.busy":"2022-07-20T10:54:36.476131Z","iopub.execute_input":"2022-07-20T10:54:36.476512Z","iopub.status.idle":"2022-07-20T10:54:42.146393Z","shell.execute_reply.started":"2022-07-20T10:54:36.476481Z","shell.execute_reply":"2022-07-20T10:54:42.145379Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Variables Normalization\n\nWe observe that many of the variables are correlated, let's uncorrelate them and chose the ones with different distributions","metadata":{}},{"cell_type":"code","source":"train.loc[:,'ratio_long_words'] = train.loc[:,'num_long_words']/train.loc[:,'num_words']\ntrain.loc[:,'ratio_sentences'] = train.loc[:,'num_sentences']/train.loc[:,'num_words']\ntrain.loc[:,'ratio_mispelling'] = train.loc[:,'num_mispelling']/train.loc[:,'num_words']\ntrain.loc[:,'ratio_punctuation'] = train.loc[:,'num_punctuation']/train.loc[:,'num_words']","metadata":{"execution":{"iopub.status.busy":"2022-07-20T10:56:13.133261Z","iopub.execute_input":"2022-07-20T10:56:13.133747Z","iopub.status.idle":"2022-07-20T10:56:13.148551Z","shell.execute_reply.started":"2022-07-20T10:56:13.133704Z","shell.execute_reply":"2022-07-20T10:56:13.147117Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def feature_creation(data_):\n    data = data_\n    data.loc[:,'len_discourse'] = data.loc[:,'discourse_text'].apply(lambda x: len(x))\n    data.loc[:,'num_words'] = data.loc[:,'discourse_text'].apply(lambda x : len(x.split()))\n    data.loc[:,'num_long_words'] = data.loc[:,'discourse_text'].apply(lambda x: len([word for word in x.split() if len(word)>6]))\n    data.loc[:,'num_sentences'] = data.loc[:,'discourse_text'].apply(lambda x :len(x.replace('?','.').replace('!','.').split('.')))\n    data.loc[:,'num_punctuation'] = data.loc[:,'discourse_text'].apply(lambda x: len(re.findall(r'[^\\w\\s]', x)))\n    data.loc[:,'num_exlamations'] = data.loc[:,'discourse_text'].apply(lambda x :len(re.findall(r'\\!',x)))\n    data.loc[:,'num_interrogations'] = data.loc[:,'discourse_text'].apply(lambda x :len(re.findall(r'\\?',x)))\n    data.loc[:,'num_commas'] = data.loc[:,'discourse_text'].apply(lambda x :len(re.findall(',',x)))\n    data.loc[:,'num_doublepoints'] = data.loc[:,'discourse_text'].apply(lambda x :len(re.findall(r'\\:',x)))\n    data.loc[:,'num_digits'] = data.loc[:,'discourse_text'].apply(lambda x :len(re.findall('\\d+',x)))\n    data.loc[:,'num_mispelling'] = data.loc[:,'discourse_text'].apply(lambda x: len([word for word in x.split() if word not in english_dict]))\n    data.loc[:,'ratio_long_words'] = data.loc[:,'num_long_words']/train.loc[:,'num_words']\n    data.loc[:,'ratio_sentences'] = data.loc[:,'num_sentences']/train.loc[:,'num_words']\n    data.loc[:,'ratio_mispelling'] = data.loc[:,'num_mispelling']/train.loc[:,'num_words']\n    data.loc[:,'ratio_punctuation'] = data.loc[:,'num_punctuation']/train.loc[:,'num_words']\n    return data","metadata":{"execution":{"iopub.status.busy":"2022-07-20T10:23:51.786643Z","iopub.execute_input":"2022-07-20T10:23:51.787524Z","iopub.status.idle":"2022-07-20T10:23:51.806244Z","shell.execute_reply.started":"2022-07-20T10:23:51.787481Z","shell.execute_reply":"2022-07-20T10:23:51.805270Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ratio_variables = ['ratio_long_words', 'ratio_sentences', 'ratio_mispelling', 'ratio_punctuation']\nfig, ax = plt.subplots(nrows=len(ratio_variables), ncols=1, figsize=(16, 25))\n\nfor k, variable in enumerate(ratio_variables):\n    sns.boxplot(data = train, \n            y = variable, \n            x='discourse_type', \n            hue='discourse_effectiveness', \n            hue_order = ['Ineffective', 'Adequate', 'Effective'], \n            palette = ['lightcoral', 'lightyellow', 'lightgreen'],\n            ax=ax[k])\n","metadata":{"execution":{"iopub.status.busy":"2022-07-20T10:56:23.423614Z","iopub.execute_input":"2022-07-20T10:56:23.424091Z","iopub.status.idle":"2022-07-20T10:56:26.077968Z","shell.execute_reply.started":"2022-07-20T10:56:23.424054Z","shell.execute_reply":"2022-07-20T10:56:26.076728Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Observations:\n* On distributions:\n    * The variable **num_long_words** distinguish very well the different effectiveness, we see clearly if the greater is the effectiveness the greater are the number of long words.\n    * The ratio of sentences also looks to be different, there are less sentences for more effective discourses.\n    * For the mispelling and punctuation there are \n* On outliers: \n    * Regarding punctuation, we observe that some discourses  have a ratio of 1 or 2 that means that there is 1 or more punctuation sign per word, which does not make sense.\n    * Regardind mispelling: We observe the same, some discourses are only made of mispelled words. Even for effective discourses\n    \n","metadata":{}},{"cell_type":"markdown","source":"# Logistic Regression for Model Interpretability","metadata":{}},{"cell_type":"code","source":"from sklearn.pipeline import Pipeline\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.compose import ColumnTransformer\nfrom sklearn.preprocessing import StandardScaler, Normalizer, OneHotEncoder\n\n############Only text\ncols_trans_v1 = ColumnTransformer([\n    ('vect_text', TfidfVectorizer(ngram_range=(1,3), min_df=0.001), 'discourse_text'),\n], remainder='drop')\n\n\n############Only features\ncols_trans_v2 = ColumnTransformer([\n    ('features', StandardScaler(), ratio_variables),\n], remainder='drop')\n\ncols_trans_v3 = ColumnTransformer([\n     ('vect_text', TfidfVectorizer(ngram_range=(1,3), min_df=0.001), 'discourse_text'),\n    ('features', StandardScaler(), ratio_variables),\n], remainder='drop')\n\ncols_trans_v4 = ColumnTransformer([\n     ('vect_text', TfidfVectorizer(ngram_range=(1,3), min_df=0.001), 'discourse_text'),\n    ('features', StandardScaler(), ratio_variables),\n    ('cat', OneHotEncoder(), ['discourse_type'])\n], remainder='drop')\n\n\nclfs = {}\n\n\nclfs['LOG_v1'] = Pipeline([ ####################################Best model for bio's\n    ('trans', cols_trans_v1),\n    ('clf', LogisticRegression(random_state=42, multi_class='multinomial', class_weight='balanced', solver='newton-cg'))])\n\n\nclfs['LOG_v2'] = Pipeline([ ####################################Best model for bio's\n    ('trans', cols_trans_v2),\n    ('clf', LogisticRegression(random_state=42, multi_class='multinomial', class_weight='balanced', solver='newton-cg'))])\n\nclfs['LOG_v3'] = Pipeline([ ####################################Best model for bio's\n    ('trans', cols_trans_v3),\n    ('clf', LogisticRegression(random_state=42, multi_class='multinomial', class_weight='balanced', solver='newton-cg'))])\n\nclfs['LOG_v4'] = Pipeline([ ####################################Best model for bio's\n    ('trans', cols_trans_v4),\n    ('clf', LogisticRegression(random_state=42, multi_class='multinomial', class_weight='balanced', solver='newton-cg'))])\n\n","metadata":{"execution":{"iopub.status.busy":"2022-07-20T11:04:51.915824Z","iopub.execute_input":"2022-07-20T11:04:51.916732Z","iopub.status.idle":"2022-07-20T11:04:51.936712Z","shell.execute_reply.started":"2022-07-20T11:04:51.916693Z","shell.execute_reply":"2022-07-20T11:04:51.935208Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from tqdm import tqdm\n\nmapping_target = {effectiveness:k for k,effectiveness in enumerate(['Ineffective', 'Adequate', 'Effective'])}\n\ntrain['target'] = train['discourse_effectiveness'].apply(lambda x: mapping_target[x])","metadata":{"execution":{"iopub.status.busy":"2022-07-20T11:04:53.377293Z","iopub.execute_input":"2022-07-20T11:04:53.378222Z","iopub.status.idle":"2022-07-20T11:04:53.401119Z","shell.execute_reply.started":"2022-07-20T11:04:53.378168Z","shell.execute_reply":"2022-07-20T11:04:53.399801Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Separate your training data\n\nWe separate our train data into a nre reduce train set and a dev set that will serve to chose the best model","metadata":{}},{"cell_type":"code","source":"from sklearn.metrics import confusion_matrix\ndef metrics_report_cat3(target, preds, cats):\n    indexes = [f'{cat}' for cat in cats]\n    cm = confusion_matrix(target, preds)\n    cm_df = pd.DataFrame(cm, index=indexes, columns=indexes)\n    plt.figure(figsize=(5,4))\n    sns.heatmap(cm_df, annot=True, cmap=\"Blues\")\n    plt.title('Confusion Matrix')\n    plt.ylabel('Actual Values')\n    plt.xlabel('Predicted Values')\n    plt.show()    \n    accuracy = np.sum(np.diag(cm))/np.sum(cm)\n    recall = np.diag(cm) / np.sum(cm, axis = 1)\n    precision = np.diag(cm) / np.sum(cm, axis = 0)\n    f1 = 2*recall*precision/(recall+precision)\n\n    print('Results model with username where accuracy is: ', accuracy)\n    print('|Age group|Precision |Recall|F1|')\n    print('|:--:|:---:|:-:|:--:|')\n    for k in range(recall.shape[0]):\n        print(f'|{cm_df.index[k]}|{precision[k]:.02f}|{recall[k]:.02f}|{f1[k]:.02f}|')\n    return None","metadata":{"execution":{"iopub.status.busy":"2022-07-20T11:05:52.042314Z","iopub.execute_input":"2022-07-20T11:05:52.043335Z","iopub.status.idle":"2022-07-20T11:05:52.054303Z","shell.execute_reply.started":"2022-07-20T11:05:52.043268Z","shell.execute_reply":"2022-07-20T11:05:52.053043Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split\nX_train,X_dev, y_train, y_dev = train_test_split(train, train['target'].values, test_size=0.33, random_state=13)","metadata":{"execution":{"iopub.status.busy":"2022-07-20T11:05:59.744791Z","iopub.execute_input":"2022-07-20T11:05:59.745158Z","iopub.status.idle":"2022-07-20T11:05:59.774046Z","shell.execute_reply.started":"2022-07-20T11:05:59.745127Z","shell.execute_reply":"2022-07-20T11:05:59.772763Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for key, clf in clfs.items():\n    print(f'---------------{key}--------------------')\n    clf.fit(X_train, y_train)\n    metrics_report_cat3(y_dev, clf.predict(X_dev), ['Ineffective', 'Adequate', 'Effective'])\n    print('')","metadata":{"execution":{"iopub.status.busy":"2022-07-20T11:06:08.716938Z","iopub.execute_input":"2022-07-20T11:06:08.717338Z","iopub.status.idle":"2022-07-20T11:07:00.764165Z","shell.execute_reply.started":"2022-07-20T11:06:08.717305Z","shell.execute_reply":"2022-07-20T11:07:00.762984Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We see that the best model is **LOG_v4**. \n\nThe results are: \n\nResults model with username where accuracy is:  0.5829555757026292\n\n\n|Age group|Precision |Recall|F1|\n|:--:|:---:|:-:|:--:|\n|Ineffective|0.38|0.56|0.45|\n|Adequate|0.72|0.53|0.61|\n|Effective|0.58|0.72|0.64|\n\nThis model:\n\n* Takes into accout unigrams, bigrams and trigrams\n* Takes into account the features we study previously\n* Differentiates between discourse type","metadata":{}},{"cell_type":"markdown","source":"### Study Top Coefficients","metadata":{}},{"cell_type":"code","source":"feature_names = clfs['LOG_v4'].named_steps[\"trans\"].get_feature_names_out()\ncoefs = clfs['LOG_v4'].named_steps[\"clf\"].coef_\n\nfor coef, name in zip(coefs, ['Ineffective', 'Adequate', 'Effective']):\n\n# Zip coefficients and names together and make a DataFrame\n    zipped = zip(feature_names, coef)\n    df = pd.DataFrame(zipped, columns=[\"feature\", \"value\"])\n# Sort the features by the absolute value of their coefficient\n    df[\"abs_value\"] = df[\"value\"].apply(lambda x: abs(x))\n    df[\"colors\"] = df[\"value\"].apply(lambda x: \"green\" if x > 0 else \"red\")\n    df = df.sort_values(\"abs_value\", ascending=False)\n\n    import seaborn as sns\n    fig, ax = plt.subplots(1, 1, figsize=(12, 7))\n    sns.barplot(x=\"feature\",\n            y=\"value\",\n            data=df.head(20),\n           palette=df.head(20)[\"colors\"])\n    ax.set_xticklabels(ax.get_xticklabels(), rotation=90, fontsize=20)\n    ax.set_title(f\"Top 20 Features contributing or penalizing to the {name} class\", fontsize=25)\n    ax.set_ylabel(\"Coef\", fontsize=22)\n    ax.set_xlabel(\"Feature Name\", fontsize=22)","metadata":{"execution":{"iopub.status.busy":"2022-07-20T11:07:55.512244Z","iopub.execute_input":"2022-07-20T11:07:55.512610Z","iopub.status.idle":"2022-07-20T11:07:56.721536Z","shell.execute_reply.started":"2022-07-20T11:07:55.512581Z","shell.execute_reply":"2022-07-20T11:07:56.717508Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"All of the top Coefficients are unigrams or bigrams","metadata":{}},{"cell_type":"markdown","source":"### Look at Results specific to each discourse type","metadata":{}},{"cell_type":"code","source":"discourses = np.unique(train.loc[:,'discourse_type'])\nfor discourse in discourses:\n    print(f'---------------{discourse}--------------------')\n    X_dev.loc[:,'target'] = y_dev\n    X_dev_discourse = X_dev.loc[X_dev.loc[:,'discourse_type']==discourse,]\n    metrics_report_cat3(X_dev_discourse['target'].values, clfs['LOG_v4'].predict(X_dev_discourse), ['Ineffective', 'Adequate', 'Effective'])\n    print('')","metadata":{"execution":{"iopub.status.busy":"2022-07-20T11:09:01.511982Z","iopub.execute_input":"2022-07-20T11:09:01.512414Z","iopub.status.idle":"2022-07-20T11:09:05.119001Z","shell.execute_reply.started":"2022-07-20T11:09:01.512377Z","shell.execute_reply":"2022-07-20T11:09:05.117889Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Conclusions and Next Steps:\n\n## Conclusions:\n* We have a highly unbalanced train dataset.\n* We have seen that features like: number of long words or ratio ofsentences are clear indicators of the effectiveness\n* With the logistic regression we managed to pass from 33% (random) of Accuracy to 58%\n\n## Next Steps:\n* Study other type of features for distinguishing the efectiveness\n* Try different approaches to solve the **data unbalanced** issue: data augmentation on **Ineffective** and **Effective** or **Undersampling** or some sort of **Penalization** term.\n* Same for **discourse type**\n* Try a non linear classifier: **Random Forest, XGBoost**\n* Other types of embeddings: **Fasttext, BERT, RNN** that account for more things.\n","metadata":{}},{"cell_type":"markdown","source":"## Make Submission","metadata":{}},{"cell_type":"code","source":"submission = pd.DataFrame(columns=['discourse_id','Ineffective','Adequate','Effective'])\ntest_ = feature_creation(test)\nprobas = clfs['LOG_v4'].predict_proba(test_)\nsubmission['discourse_id'] = test['discourse_id']\nsubmission['Ineffective'] = probas[:,0]\nsubmission['Adequate'] = probas[:,1]\nsubmission['Effective'] = probas[:,2]\n","metadata":{"execution":{"iopub.status.busy":"2022-07-20T11:22:37.884732Z","iopub.execute_input":"2022-07-20T11:22:37.885259Z","iopub.status.idle":"2022-07-20T11:22:37.928999Z","shell.execute_reply.started":"2022-07-20T11:22:37.885218Z","shell.execute_reply":"2022-07-20T11:22:37.928083Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission.to_csv('submission.csv',index=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-20T11:22:38.238388Z","iopub.execute_input":"2022-07-20T11:22:38.239474Z","iopub.status.idle":"2022-07-20T11:22:38.248354Z","shell.execute_reply.started":"2022-07-20T11:22:38.239431Z","shell.execute_reply":"2022-07-20T11:22:38.247250Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}