{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Simple EDA for beginners with `seaborn`","metadata":{}},{"cell_type":"markdown","source":"The following notebook consists of a simple EDA of the dataset.\n\nAll the graphs are visualised using `seaborn`\n\nThe conclusions from the EDA is presented in the last cell\n\nThanks to https://www.kaggle.com/code/lextoumbourou/feedback-prize-the-complete-overview for inspiration\n\nPlease **upvote** and feel free to suggest changes and provide feedback. Thank you :)","metadata":{}},{"cell_type":"code","source":"import os\n\nimport pandas as pd\nimport seaborn as sns","metadata":{"execution":{"iopub.status.busy":"2022-07-28T11:03:32.387491Z","iopub.execute_input":"2022-07-28T11:03:32.388903Z","iopub.status.idle":"2022-07-28T11:03:33.691911Z","shell.execute_reply.started":"2022-07-28T11:03:32.388746Z","shell.execute_reply":"2022-07-28T11:03:33.690044Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Read the data","metadata":{"tags":[]}},{"cell_type":"code","source":"path = \"train\"\ndef get_essay(essay_id):\n    essay_path = os.path.join(\"../input/feedback-prize-effectiveness/\"+path+\"/\", f\"{essay_id}.txt\")\n    essay_text = open(essay_path, 'r').read()\n    return essay_text","metadata":{"execution":{"iopub.status.busy":"2022-07-28T11:03:33.694837Z","iopub.execute_input":"2022-07-28T11:03:33.695253Z","iopub.status.idle":"2022-07-28T11:03:33.700650Z","shell.execute_reply.started":"2022-07-28T11:03:33.695216Z","shell.execute_reply":"2022-07-28T11:03:33.699553Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df = pd.read_csv(\"../input/feedback-prize-effectiveness/train.csv\")\ndf['essay_text'] = df['essay_id'].apply(get_essay)\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-28T11:03:33.702330Z","iopub.execute_input":"2022-07-28T11:03:33.702751Z","iopub.status.idle":"2022-07-28T11:04:04.328608Z","shell.execute_reply.started":"2022-07-28T11:03:33.702715Z","shell.execute_reply":"2022-07-28T11:04:04.327347Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"path=\"test\"\ntest_df = pd.read_csv(\"../input/feedback-prize-effectiveness/test.csv\", )\ntest_df['essay_text'] = test_df['essay_id'].apply(get_essay)\ntest_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-28T11:04:04.330502Z","iopub.execute_input":"2022-07-28T11:04:04.330924Z","iopub.status.idle":"2022-07-28T11:04:04.361078Z","shell.execute_reply.started":"2022-07-28T11:04:04.330888Z","shell.execute_reply":"2022-07-28T11:04:04.359664Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.columns","metadata":{"execution":{"iopub.status.busy":"2022-07-28T11:04:04.364126Z","iopub.execute_input":"2022-07-28T11:04:04.364599Z","iopub.status.idle":"2022-07-28T11:04:04.374100Z","shell.execute_reply.started":"2022-07-28T11:04:04.364556Z","shell.execute_reply":"2022-07-28T11:04:04.372750Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Visualise","metadata":{"tags":[]}},{"cell_type":"code","source":"g = sns.catplot(data=df, x=\"discourse_type\",\n            height=6,\n            aspect=3,\n            kind=\"count\")","metadata":{"execution":{"iopub.status.busy":"2022-07-28T11:04:04.376268Z","iopub.execute_input":"2022-07-28T11:04:04.376906Z","iopub.status.idle":"2022-07-28T11:04:04.769797Z","shell.execute_reply.started":"2022-07-28T11:04:04.376862Z","shell.execute_reply":"2022-07-28T11:04:04.768165Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"g = sns.catplot(\n    data=df,\n    x=\"discourse_effectiveness\",\n    col=\"discourse_type\",\n    kind=\"count\",\n    height = 3,\n)","metadata":{"execution":{"iopub.status.busy":"2022-07-28T11:04:04.771642Z","iopub.execute_input":"2022-07-28T11:04:04.772671Z","iopub.status.idle":"2022-07-28T11:04:05.780895Z","shell.execute_reply.started":"2022-07-28T11:04:04.772620Z","shell.execute_reply":"2022-07-28T11:04:05.779570Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"g = sns.catplot(\n    data=df,\n    x=\"discourse_effectiveness\",\n    kind=\"count\",\n)","metadata":{"execution":{"iopub.status.busy":"2022-07-28T11:04:05.783024Z","iopub.execute_input":"2022-07-28T11:04:05.783422Z","iopub.status.idle":"2022-07-28T11:04:06.080603Z","shell.execute_reply.started":"2022-07-28T11:04:05.783387Z","shell.execute_reply":"2022-07-28T11:04:06.079621Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# A general trend is to pass the essay + dicourse_text + discourse_type as input to the model\n# Here we form a different column [total_word_count] which consists of the length of this concatenated text. \n\ndef get_word_count(x):\n    return len(x.split())\n\ndf[\"total_word_count\"] = (df[\"essay_text\"]+df[\"discourse_type\"]+df[\"discourse_text\"]).apply(get_word_count)\ndf[\"discourse_word_count\"] = (df[\"discourse_text\"]).apply(get_word_count)\n\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-28T11:04:06.081796Z","iopub.execute_input":"2022-07-28T11:04:06.083163Z","iopub.status.idle":"2022-07-28T11:04:07.360003Z","shell.execute_reply.started":"2022-07-28T11:04:06.083109Z","shell.execute_reply":"2022-07-28T11:04:07.358690Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"g = sns.displot(df[\"total_word_count\"], height=5, aspect=4)","metadata":{"execution":{"iopub.status.busy":"2022-07-28T11:04:07.361930Z","iopub.execute_input":"2022-07-28T11:04:07.362530Z","iopub.status.idle":"2022-07-28T11:04:07.865719Z","shell.execute_reply.started":"2022-07-28T11:04:07.362450Z","shell.execute_reply":"2022-07-28T11:04:07.863942Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"g = sns.catplot(data = df, x = 'total_word_count', kind = \"box\", height = 5, aspect = 4)","metadata":{"execution":{"iopub.status.busy":"2022-07-28T11:04:07.867766Z","iopub.execute_input":"2022-07-28T11:04:07.869235Z","iopub.status.idle":"2022-07-28T11:04:08.209078Z","shell.execute_reply.started":"2022-07-28T11:04:07.869167Z","shell.execute_reply":"2022-07-28T11:04:08.207974Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(df['total_word_count'].mean())\nprint(df['total_word_count'].max())\nprint(df['total_word_count'].min())","metadata":{"execution":{"iopub.status.busy":"2022-07-28T11:04:08.210994Z","iopub.execute_input":"2022-07-28T11:04:08.211789Z","iopub.status.idle":"2022-07-28T11:04:08.221159Z","shell.execute_reply.started":"2022-07-28T11:04:08.211735Z","shell.execute_reply":"2022-07-28T11:04:08.219843Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Number of data samples that will be too short for any standard transformer\nn = len(df[df['total_word_count']>512])\nprint(\"Number of samples with total word count more than 512: \", n)\nprint(\"Percentage of samples that have total word count more than 512: \", (n/len(df))*100)","metadata":{"execution":{"iopub.status.busy":"2022-07-28T11:04:08.223009Z","iopub.execute_input":"2022-07-28T11:04:08.223806Z","iopub.status.idle":"2022-07-28T11:04:08.240156Z","shell.execute_reply.started":"2022-07-28T11:04:08.223753Z","shell.execute_reply":"2022-07-28T11:04:08.238632Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"g = sns.displot(df[\"discourse_word_count\"], height=5, aspect=4)","metadata":{"execution":{"iopub.status.busy":"2022-07-28T11:04:08.244955Z","iopub.execute_input":"2022-07-28T11:04:08.245834Z","iopub.status.idle":"2022-07-28T11:04:09.344355Z","shell.execute_reply.started":"2022-07-28T11:04:08.245780Z","shell.execute_reply":"2022-07-28T11:04:09.343165Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"g = sns.catplot(data = df, x = 'discourse_word_count', kind = \"box\", height = 5, aspect = 4)","metadata":{"execution":{"iopub.status.busy":"2022-07-28T11:04:09.345988Z","iopub.execute_input":"2022-07-28T11:04:09.346630Z","iopub.status.idle":"2022-07-28T11:04:09.623777Z","shell.execute_reply.started":"2022-07-28T11:04:09.346589Z","shell.execute_reply":"2022-07-28T11:04:09.622494Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(df['discourse_word_count'].mean())\nprint(df['discourse_word_count'].max())\nprint(df['discourse_word_count'].min())","metadata":{"execution":{"iopub.status.busy":"2022-07-28T11:04:09.626361Z","iopub.execute_input":"2022-07-28T11:04:09.626800Z","iopub.status.idle":"2022-07-28T11:04:09.635321Z","shell.execute_reply.started":"2022-07-28T11:04:09.626760Z","shell.execute_reply":"2022-07-28T11:04:09.633841Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Number of data samples that will be too short for any standard transformer\nn = len(df[df['discourse_word_count']>512])\nprint(\"Number of samples with total word count more than 512: \", n)\nprint(\"Percentage of samples that have total word count more than 512: \", (n/len(df))*100)","metadata":{"execution":{"iopub.status.busy":"2022-07-28T11:04:19.374532Z","iopub.execute_input":"2022-07-28T11:04:19.374966Z","iopub.status.idle":"2022-07-28T11:04:19.385180Z","shell.execute_reply.started":"2022-07-28T11:04:19.374933Z","shell.execute_reply":"2022-07-28T11:04:19.383813Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Conclusions","metadata":{"tags":[]}},{"cell_type":"markdown","source":"1. There is imbalance in the dataset between the three output classes [Adequate, Effective, Ineffective]. (Use class weights technique)\n2. The number of Adequate samples are relatively higher in each discourse type. (Should the class weights depend on the discourse type as well?)\n3. Average Number of words of total (essay+discourse+type) is 502. (Architectures handling 512 length can be used)\n4. [If we use essay + discourse_text as input] Most of the notebooks use 512 as max length for token. Here around 39% of the data has input length of more than 512 words. Which means we need to truncate them, and this could cause loss of information(context) and \"might\" affect the performance of the model.\n5. Using just discourse_text and maybe discourse type as input, only 2% of the data will have number of words more than 512. So less loss of information by truncating compared to the method specified in the previous point.","metadata":{}}]}