{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# <h><center>Feedback Prize - Predicting Effective Arguments</center></h>\n\n<img src='https://www.incimages.com/uploaded_files/image/1920x1080/getty_506903004_200013332000928076_348061.jpg'>\n\n\n\n# <center>**NLP Based Problem: Classification + TextData.....**</center>\n\n## <center>**Goal of the Competition:**</center> \n\nThe goal of this competition is to classify argumentative elements in student writing as \"effective,\" \"adequate,\" or \"ineffective.\" You will create a model trained on data that is representative of the 6th-12th grade population in the United States in order to minimize bias. Models derived from this competition will help pave the way for students to receive enhanced feedback on their argumentative writing. With automated guidance, students can complete more assignments and ultimately become more confident, proficient writers.","metadata":{}},{"cell_type":"markdown","source":"## ***Continue the US pattern matching competition approach:***\n\n1. https://www.kaggle.com/code/venkatkumar001/u-s-p-p-baseline-eda-dataprep\n2. Starter1: https://www.kaggle.com/code/venkatkumar001/nlp-starter1-almost-all-basic-concept\n3. Starter2: https://www.kaggle.com/code/venkatkumar001/nlp-starter2-hf-pretrain-finetune\n4. https://www.kaggle.com/code/venkatkumar001/transformeranatomy-encoder\n\n## ***Now, This feedback price competiton onward!***\n\n1. Starter3: https://www.kaggle.com/venkatkumar001/nlpstarter3-baseline-approach\n\n\n## ***So, I am trying to build baseline Approach of Competition data***\n\n# **Steps:**\n\n### **1. Import Necessary Library**\n\n### **2. Load and analysis the data**\n\n### **3. Preprocessing**\n\n### **4. Feature selection**\n\n### **5. Build the Model**\n\n### **6. Predict Output**\n\n### **7. Generate Submission file**\n\n","metadata":{}},{"cell_type":"markdown","source":"# <center>**Import Necessary Library**</center>","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport re\nimport os\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport time\nimport datetime\nfrom scipy import sparse\n\n\nfrom sklearn.feature_extraction.text import TfidfVectorizer\nfrom sklearn.model_selection import KFold\nfrom sklearn.preprocessing import OneHotEncoder,LabelEncoder\nfrom sklearn.metrics import log_loss\n\nfrom sklearn import svm\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.ensemble import RandomForestClassifier\nfrom sklearn.naive_bayes import BernoulliNB,MultinomialNB\nfrom sklearn.model_selection import cross_val_score\nfrom sklearn.metrics import classification_report, confusion_matrix, accuracy_score, f1_score, precision_score,recall_score","metadata":{"execution":{"iopub.status.busy":"2022-07-20T14:47:39.744292Z","iopub.execute_input":"2022-07-20T14:47:39.744817Z","iopub.status.idle":"2022-07-20T14:47:40.924860Z","shell.execute_reply.started":"2022-07-20T14:47:39.744702Z","shell.execute_reply":"2022-07-20T14:47:40.923117Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import nltk\nfrom nltk.corpus import stopwords\nfrom nltk.stem import SnowballStemmer\nfrom string import punctuation\nfrom nltk.stem.wordnet import WordNetLemmatizer\nfrom tqdm import tqdm\n\n%matplotlib inline\n","metadata":{"execution":{"iopub.status.busy":"2022-07-20T14:47:43.850479Z","iopub.execute_input":"2022-07-20T14:47:43.850895Z","iopub.status.idle":"2022-07-20T14:47:44.352718Z","shell.execute_reply.started":"2022-07-20T14:47:43.850862Z","shell.execute_reply":"2022-07-20T14:47:44.351442Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <center>**Load and Analysis the data**</center>","metadata":{}},{"cell_type":"code","source":"!ls '../input/feedback-prize-effectiveness'","metadata":{"execution":{"iopub.status.busy":"2022-07-20T14:47:47.427112Z","iopub.execute_input":"2022-07-20T14:47:47.427616Z","iopub.status.idle":"2022-07-20T14:47:48.202460Z","shell.execute_reply.started":"2022-07-20T14:47:47.427577Z","shell.execute_reply":"2022-07-20T14:47:48.200429Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train = pd.read_csv('../input/feedback-prize-effectiveness/train.csv')\ntest = pd.read_csv('../input/feedback-prize-effectiveness/test.csv')\nsample = pd.read_csv('../input/feedback-prize-effectiveness/sample_submission.csv')\nprint(f'Train_Shape: {train.shape},Test_Shape: {test.shape},Sample_Shape: {sample.shape}')\ndisplay(train.sample(2))\ndisplay(test.sample(2))\ndisplay(sample.sample(2))","metadata":{"execution":{"iopub.status.busy":"2022-07-20T14:48:37.959058Z","iopub.execute_input":"2022-07-20T14:48:37.959498Z","iopub.status.idle":"2022-07-20T14:48:38.182743Z","shell.execute_reply.started":"2022-07-20T14:48:37.959448Z","shell.execute_reply":"2022-07-20T14:48:38.180798Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.info()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.describe(include='object')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train['discourse_type'].value_counts()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train['discourse_effectiveness'].value_counts()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(15,10))\nsns.countplot(x='discourse_id',data=train)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(15,10))\nsns.countplot(train['discourse_type'])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(15,10))\nsns.countplot(train['discourse_effectiveness'])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(15,10))\nsns.countplot(train['essay_id'])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.discourse_type.unique()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **Identify the imbalance of data - discourse_type**","metadata":{}},{"cell_type":"code","source":"#identify the imbalance one\n\ndef count_target(target_list):\n    target_dict = {}\n    for x in target_list:\n        count = len(train['discourse_type'] == x)\n        dict_t = dict({x:count})\n        target_dict.update(dict_t)\n    return target_dict\n\ntarget_list = ['Lead', 'Position', 'Claim', 'Evidence', 'Counterclaim',\n       'Rebuttal', 'Concluding Statement']\ncount_target(target_list)\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from lightgbm import LGBMRegressor,LGBMClassifier\ntrain['kfold'] = -1\nkfold = KFold(n_splits=10, shuffle=True,random_state=100)\nfor fold, (train_ind, valid_ind) in enumerate(kfold.split(X = train)):\n    train.loc[valid_ind,'kfold'] = fold\n# train.to_csv(\"trainfold_10.csv\",index=False)   \ndef preprocess_text(text):\n    text = re.sub('[^A-Za-z]',' ',text)\n    text=text.lower()\n    return text\ntrain['preprocessed_discourse_text'] = train['discourse_text'].apply(preprocess_text)\ntrain['discourse_effectiveness'] = train['discourse_effectiveness'].map({'Ineffective' : 0, 'Adequate':1,'Effective':2})\n# train['discourse_effectiveness'] = train['discourse_effectiveness'].to_categorical()\ndef data_prep(df):\n    tf = TfidfVectorizer(input='content',\n        ngram_range=(1,1),\n        use_idf=True,\n        smooth_idf=True)\n    X = tf.fit_transform(df['preprocessed_discourse_text'])\n    ohe = OneHotEncoder(sparse=False)\n    one_hot = ohe.fit_transform(df['discourse_type'].values.reshape(-1, 1))\n    tf_idf = sparse.hstack((X,one_hot))\n    return tf_idf, ohe, tf\ntrain_X, ohe, tf = data_prep(train)\ndef build_lgbm_model(X,y):\n    for k in range(10):\n        print(f\"fold:{k}\")\n        lgbm_param = {\"boosting_type\": \"gbdt\",\n        \"objective\": \"multiclass\",\n        'learning_rate': 0.051635,\n        \"max_depth\": 6,\n        'random_state':12,\n        'n_estimators':1000}\n        lgbm_param['metric']='multi_logloss'\n        model = LGBMClassifier(**lgbm_param)\n        model.fit(X,y)\n    return model    \nmodel = build_lgbm_model(train_X,train['discourse_effectiveness'])   ","metadata":{"execution":{"iopub.status.busy":"2022-07-20T14:48:51.424839Z","iopub.execute_input":"2022-07-20T14:48:51.425308Z","iopub.status.idle":"2022-07-20T15:03:44.249720Z","shell.execute_reply.started":"2022-07-20T14:48:51.425272Z","shell.execute_reply":"2022-07-20T15:03:44.248180Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test['preprocessed_discourse_text'] = test['discourse_text'].apply(preprocess_text)\nX = tf.transform(test['preprocessed_discourse_text'])\none_hot = ohe.transform(test['discourse_type'].values.reshape(-1, 1))\ntest_X = sparse.hstack((X,one_hot))\nprediction = model.predict_proba(test_X)\nsub = pd.read_csv(\"../input/feedback-prize-effectiveness/sample_submission.csv\")\nsub['Ineffective'] = prediction[:,0]\nsub['Adequate'] = prediction[:,1]\nsub['Effective'] = prediction[:,2]\nsub.to_csv(\"submission.csv\", index=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-20T15:07:29.693836Z","iopub.execute_input":"2022-07-20T15:07:29.694292Z","iopub.status.idle":"2022-07-20T15:07:29.729239Z","shell.execute_reply.started":"2022-07-20T15:07:29.694257Z","shell.execute_reply":"2022-07-20T15:07:29.728299Z"},"trusted":true},"execution_count":null,"outputs":[]}]}