{"metadata":{"kernelspec":{"name":"python3","display_name":"Python 3","language":"python"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"colab":{"provenance":[],"gpuType":"T4"},"accelerator":"GPU","kaggle":{"accelerator":"none","dataSources":[{"sourceId":10737,"databundleVersionId":290346,"sourceType":"competition"},{"sourceId":10228339,"sourceType":"datasetVersion","datasetId":6323872}],"dockerImageVersionId":30805,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"from IPython.display import Image\n\n\nImage(filename='/kaggle/input/lstm-cell-image/LSTM.jpg')\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-17T16:33:17.647626Z","iopub.execute_input":"2024-12-17T16:33:17.647971Z","iopub.status.idle":"2024-12-17T16:33:17.697975Z","shell.execute_reply.started":"2024-12-17T16:33:17.647940Z","shell.execute_reply":"2024-12-17T16:33:17.697238Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"---\n\n# **Project Proposal: Quora Insincere Question Classification**\n\n---\n\n#### **Introduction:**\nThe aim of this project is to build a machine learning model that can predict whether a question asked on Quora is sincere or insincere. The dataset consists of user-submitted questions, and the goal is to classify each question as either sincere or insincere based on its content. An insincere question is one that aims to make a statement or provoke a reaction rather than seek a genuine answer.\n\n---\n\n#### **Project Objective:**\nThe objectives of this project are:\n- Build a predictive model to classify Quora questions as sincere or insincere.\n- Explore the relationships between question features (such as text length, use of specific words, tone, etc.) and their sincerity.\n- Provide a machine learning solution that can be used to automatically classify future questions in real-time.\n\n---\n\n#### **Dataset:**\nThe project relies on the **Quora Insincere Question Dataset**, which contains the following columns:\n- **qid:** Unique question identifier.\n- **question_text:** The text of the question.\n- **target:** The target variable (1 = Insincere, 0 = Sincere).\n\nThe dataset is divided into:\n- **train.csv:** Training data with labeled questions.\n- **test.csv:** Test data to make predictions on.\n\n---\n\n#### **Methodology:**\n\n1. **Stage 1: Data Exploration and Preprocessing:**\n   - **Exploratory Data Analysis (EDA):** Analyze the distribution of the target variable (`target`), explore text length, frequency of words, and identify any imbalances or biases in the dataset.\n   - **Handling Missing Values:** Check for any missing values and handle them accordingly (e.g., drop or impute missing entries).\n   - **Text Preprocessing:** Apply common natural language processing (NLP) techniques such as:\n     - Convert text to lowercase.\n     - Remove special characters, numbers, and punctuation.\n     - Remove stop words that don’t contribute meaningful information.\n     - Apply stemming or lemmatization to reduce words to their root form.\n     - Remove extra spaces, URLs, and irrelevant text.\n\n2. **Stage 2: Text Tokenization and Sequence Padding:**\n   - **Tokenization:** Convert the text data into sequences of tokens (words) that the machine learning model can understand.\n   - **Sequence Padding:** Ensure all input sequences have the same length by padding them to a fixed length (e.g., 80 words).\n\n3. **Stage 3: Model Building:**\n   - **Model Architecture:** Build a **Long Short-Term Memory (LSTM)** model, a type of Recurrent Neural Network (RNN) well-suited for sequence data like text.\n     - Embedding layer: To convert words into dense vectors.\n     - LSTM layer: To capture long-term dependencies in the text.\n     - Dense layer: To predict whether the question is sincere or insincere.\n     - Use **Dropout** layers to prevent overfitting.\n\n4. **Stage 4: Model Training and Evaluation:**\n   - **Training:** Train the model using the training data (`train.csv`).\n   - **Validation:** Validate the model’s performance on a validation set.\n   - **Early Stopping:** Use early stopping to prevent overfitting during training.\n   - **Evaluation Metrics:** Evaluate the model using metrics like **accuracy**, **precision**, **recall**, and **F1-score** to assess its performance in classification tasks.\n\n5. **Stage 5: Prediction on Test Data:**\n   - **Test Data Processing:** Preprocess the test data in the same way as the training data.\n   - **Prediction:** Use the trained model to predict whether the questions in the test dataset are sincere or insincere.\n   - **Result Preparation:** Convert predictions into binary labels (0 or 1) and prepare a submission file in the required format.\n\n---\n\n#### **Proposed Models:**\n- **LSTM (Long Short-Term Memory):** A deep learning model suitable for sequence data such as text. It is effective in capturing long-term dependencies and understanding contextual relationships in text.\n- **Naive Bayes (NB):** A probabilistic model that can serve as a baseline. It is fast and simple, ideal for comparing performance against more complex models like LSTM.\n- **Logistic Regression (LR):** A traditional classification model for binary outcomes, which will also be tested as a baseline for comparison.\n\n---\n\n#### **Techniques and Tools:**\n- **Programming Language:** Python\n- **Libraries Used:**\n  - **Pandas** (for data manipulation)\n  - **NumPy** (for numerical operations)\n  - **Keras** (for deep learning model building)\n  - **TensorFlow** (for model training)\n  - **NLTK** (for natural language processing)\n  - **Matplotlib** and **Seaborn** (for data visualization)\n- **Machine Learning Techniques:** **Deep Learning**, **Text Classification**, **Natural Language Processing (NLP)**\n\n---\n\n#### **Expected Outcomes:**\n- The model is expected to achieve a high classification accuracy for distinguishing between sincere and insincere questions.\n- The model’s performance is anticipated to range between **85% and 90%** accuracy, based on the complexity of the text data and the effectiveness of preprocessing.\n\n---\n\n#### **Challenges:**\n- **Imbalanced Data:** There may be an imbalance between sincere and insincere questions, which can affect the model’s performance. We will apply techniques such as **SMOTE** (Synthetic Minority Over-sampling Technique) or class weighting to address this issue.\n- **Text Complexity:** The complexity of the questions (e.g., sarcasm, ambiguous tone) may make it challenging for the model to classify certain questions accurately.\n- **Overfitting:** Given the large number of features in the text, there is a risk of overfitting. **Dropout** and **Early Stopping** techniques will help mitigate this.\n\n---\n\n#### **Expected Results:**\n- A highly accurate model capable of classifying questions as sincere or insincere with high precision.\n- Insights into the linguistic features and tone used in insincere questions.\n- A functional system for automatic classification of questions on Quora or similar platforms.\n\n---\n\n### **Conclusion:**\nThis project aims to develop an effective machine learning model that can accurately classify Quora questions as sincere or insincere. By applying advanced **text preprocessing** techniques, leveraging deep learning models like **LSTM**, and carefully tuning the model, we aim to achieve a robust and efficient solution for detecting insincere content online.\n\n---","metadata":{"id":"5a_Y7gq7sBS3"}},{"cell_type":"markdown","source":"# 1. **Importing Libraries**","metadata":{"id":"g8nFclmtOVwf"}},{"cell_type":"markdown","source":"- In this step, we import the essential libraries we need to work with:\n  - `numpy`: A library for mathematical operations and numerical analysis.\n  - `pandas`: A library for data manipulation, such as reading data from files, cleaning, and analysis.\n  - `seaborn`: A library for advanced plotting and data visualization.\n  - `matplotlib`: A library for creating basic plots.\n","metadata":{"id":"tU-3_9N_NaT-"}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport seaborn as sns\nimport matplotlib.pyplot as plt","metadata":{"id":"pdvoJh4V0x7m","trusted":true,"execution":{"iopub.status.busy":"2024-12-17T16:33:17.699632Z","iopub.execute_input":"2024-12-17T16:33:17.700054Z","iopub.status.idle":"2024-12-17T16:33:18.660721Z","shell.execute_reply.started":"2024-12-17T16:33:17.700016Z","shell.execute_reply":"2024-12-17T16:33:18.659849Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"---\n\n# 2. **Loading The Data**","metadata":{"id":"qSQA9gCCObOl"}},{"cell_type":"markdown","source":"- In this step, we read CSV `data` files\n\n- Then, we display the dataset using `head()` to review the content","metadata":{"id":"21ZLvvWSNlJw"}},{"cell_type":"code","source":"data = pd.read_csv('/kaggle/input/quora-insincere-questions-classification/train.csv')\nprint(\"Data loaded successfully.\")\ndata.head(5)","metadata":{"id":"vp8Hkd4b08_z","outputId":"312f7043-1534-4095-f022-57dd3f5c65f3","trusted":true,"execution":{"iopub.status.busy":"2024-12-17T16:33:18.661745Z","iopub.execute_input":"2024-12-17T16:33:18.662102Z","iopub.status.idle":"2024-12-17T16:33:22.196083Z","shell.execute_reply.started":"2024-12-17T16:33:18.662076Z","shell.execute_reply":"2024-12-17T16:33:22.195155Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n\n---\n\n# 3. **Preprocessing the Text Data**","metadata":{"id":"4t9xIgpo1s5e"}},{"cell_type":"markdown","source":"**Objective:**  \n- Improve the quality and consistency of textual data.  \n- Prepare the data for further processing and analysis by applying operations such as:  \n  - Converting text to lowercase  \n  - Removing special characters  \n  - Correcting words, and more.\n\n","metadata":{"id":"N5asCYPuUypg"}},{"cell_type":"markdown","source":"## **📍Convert Text to Lowercase**","metadata":{"id":"odkdav8WIvLy"}},{"cell_type":"markdown","source":"- **Objective:** Convert text to lowercase for consistency.\n  - Apply the `to_lowercase` function to convert text in the `question_text` column to lowercase.\n\n","metadata":{"id":"zIGX04qXVSPg"}},{"cell_type":"code","source":"def to_lowercase(text):\n  return text.lower()\n\ndata['question_text'] = data['question_text'].apply(to_lowercase)","metadata":{"id":"jhHsBSaF1oQO","trusted":true,"execution":{"iopub.status.busy":"2024-12-17T16:33:22.197178Z","iopub.execute_input":"2024-12-17T16:33:22.197433Z","iopub.status.idle":"2024-12-17T16:33:22.688261Z","shell.execute_reply.started":"2024-12-17T16:33:22.197408Z","shell.execute_reply":"2024-12-17T16:33:22.687533Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **📍 Remove Special Characters**","metadata":{"id":"cVkkkmIsIwgj"}},{"cell_type":"markdown","source":"- **Objective:** Remove special characters like numbers and punctuation marks.\n  - Apply the `remove_special_chars` function to remove non-alphabetical characters such as numbers and symbols.\n\n","metadata":{"id":"lTuxtzOyVZ14"}},{"cell_type":"code","source":"import re\n\ndef remove_special_chars(text):\n  return re.sub(r'[^a-zA-Z\\s]', '', text)\n\ndata['question_text'] = data['question_text'].apply(remove_special_chars)","metadata":{"id":"qfZ9HLUdB9V-","trusted":true,"execution":{"iopub.status.busy":"2024-12-17T16:33:22.690309Z","iopub.execute_input":"2024-12-17T16:33:22.690576Z","iopub.status.idle":"2024-12-17T16:33:25.197547Z","shell.execute_reply.started":"2024-12-17T16:33:22.690551Z","shell.execute_reply":"2024-12-17T16:33:25.196852Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **📍  Remove Stop Words**","metadata":{"id":"HOHZGu8ZIxdv"}},{"cell_type":"markdown","source":"- **Objective:** Remove common words that don't add valuable information for analysis, such as \"$the$\", \"$is$\", \"$in$\".\n  - Load the stop words set using the `nltk` library.\n  - Apply the `remove_stopwords` function to remove stop words from the text.\n","metadata":{"id":"BAzWCL-HVde2"}},{"cell_type":"code","source":"from nltk.corpus import stopwords\nstop_words = set(stopwords.words('english'))\n\ndef remove_stopwords(text):\n  return ' '.join([word for word in text.split() if word not in stop_words])\n\ndata['question_text'] = data['question_text'].apply(remove_stopwords)","metadata":{"id":"R3A5fCtOCK4p","outputId":"657dc794-528a-46f7-a84d-1a838f1e3d53","trusted":true,"execution":{"iopub.status.busy":"2024-12-17T16:33:25.198474Z","iopub.execute_input":"2024-12-17T16:33:25.198728Z","iopub.status.idle":"2024-12-17T16:33:28.746577Z","shell.execute_reply.started":"2024-12-17T16:33:25.198703Z","shell.execute_reply":"2024-12-17T16:33:28.745644Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **📍 Remove Numbers and URLs**","metadata":{"id":"3caSTjSPIzXh"}},{"cell_type":"markdown","source":"- **Objective:** Remove numbers and URLs from the text.\n  - Apply the `remove_numbers_and_urls` function to remove numbers and URLs.\n","metadata":{"id":"FEu2e8_5VnFL"}},{"cell_type":"code","source":"def remove_numbers_and_urls(text):\n  text = re.sub(r'\\d+', '', text)\n  text = re.sub(r'http\\S+', '', text)\n  return text\n\ndata['question_text'] = data['question_text'].apply(remove_numbers_and_urls)\n","metadata":{"id":"27g694unIhCZ","trusted":true,"execution":{"iopub.status.busy":"2024-12-17T16:33:28.747733Z","iopub.execute_input":"2024-12-17T16:33:28.748078Z","iopub.status.idle":"2024-12-17T16:33:31.812576Z","shell.execute_reply.started":"2024-12-17T16:33:28.748049Z","shell.execute_reply":"2024-12-17T16:33:31.811844Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **📍 Remove Extra Spaces**","metadata":{"id":"O1OezfY1I1ZH"}},{"cell_type":"markdown","source":"- **Objective:** Remove unnecessary spaces between words.\n  - Apply the `remove_extra_spaces` function to remove any extra spaces between words.","metadata":{"id":"GbsORGyfVqV4"}},{"cell_type":"code","source":"def remove_extra_spaces(text):\n  return \" \".join(text.split())\n\ndata['question_text'] = data['question_text'].apply(remove_extra_spaces)","metadata":{"id":"It6b1OicI11X","trusted":true,"execution":{"iopub.status.busy":"2024-12-17T16:33:31.813555Z","iopub.execute_input":"2024-12-17T16:33:31.813829Z","iopub.status.idle":"2024-12-17T16:33:32.784952Z","shell.execute_reply.started":"2024-12-17T16:33:31.813787Z","shell.execute_reply":"2024-12-17T16:33:32.784256Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **📍 Data verification after processing**","metadata":{"id":"yLDy9sTPyAHw"}},{"cell_type":"code","source":"data.head()","metadata":{"id":"K5LPDBzCKUEy","outputId":"969b4bcd-c9a6-412f-a03f-ce0b6079ed33","trusted":true,"execution":{"iopub.status.busy":"2024-12-17T16:33:32.786095Z","iopub.execute_input":"2024-12-17T16:33:32.786684Z","iopub.status.idle":"2024-12-17T16:33:32.795365Z","shell.execute_reply.started":"2024-12-17T16:33:32.786647Z","shell.execute_reply":"2024-12-17T16:33:32.794480Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n\n---\n\n# 4. **Preparing the Data for Modeling**","metadata":{"id":"Lp7YVofvP817"}},{"cell_type":"markdown","source":"- Separate the **Features** and the **Target**","metadata":{"id":"ZmIDIDQdWbwD"}},{"cell_type":"code","source":"X = data['question_text']\ny = data['target']","metadata":{"id":"dM-llM0HkkJw","trusted":true,"execution":{"iopub.status.busy":"2024-12-17T16:33:32.796482Z","iopub.execute_input":"2024-12-17T16:33:32.796945Z","iopub.status.idle":"2024-12-17T16:33:32.805177Z","shell.execute_reply.started":"2024-12-17T16:33:32.796918Z","shell.execute_reply":"2024-12-17T16:33:32.804430Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n\n---\n\n# 5. **Text Tokenization and Padding**","metadata":{"id":"msF7HTIImQBL"}},{"cell_type":"markdown","source":"## **📍 Text Tokenization**","metadata":{"id":"w-M8IiLol8lU"}},{"cell_type":"markdown","source":"- Convert text into a sequence of numbers.","metadata":{"id":"pNucjYY4XYK2"}},{"cell_type":"code","source":"from tensorflow.keras.preprocessing.text import Tokenizer\n\ntokenizer = Tokenizer(num_words=5000)\ntokenizer.fit_on_texts(X)","metadata":{"id":"Kzj6IutWP-PT","trusted":true,"execution":{"iopub.status.busy":"2024-12-17T16:33:32.806353Z","iopub.execute_input":"2024-12-17T16:33:32.807185Z","iopub.status.idle":"2024-12-17T16:33:57.322467Z","shell.execute_reply.started":"2024-12-17T16:33:32.807146Z","shell.execute_reply":"2024-12-17T16:33:57.321613Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## **📍 Padding the Sequences**","metadata":{"id":"Cg8noKuqlwaP"}},{"cell_type":"markdown","source":"- Convert text sequences into fixed lengths for easier processing.","metadata":{"id":"QS2VJatOXhWp"}},{"cell_type":"code","source":"from tensorflow.keras.preprocessing.sequence import pad_sequences\n\nX_seq = tokenizer.texts_to_sequences(X)\nX_seq = pad_sequences(X_seq, maxlen=80)","metadata":{"id":"iFNJwAVChApL","trusted":true,"execution":{"iopub.status.busy":"2024-12-17T16:33:57.323607Z","iopub.execute_input":"2024-12-17T16:33:57.324329Z","iopub.status.idle":"2024-12-17T16:34:12.427942Z","shell.execute_reply.started":"2024-12-17T16:33:57.324288Z","shell.execute_reply":"2024-12-17T16:34:12.427000Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(f\"Number of words in the dictionary : {len(tokenizer.word_index)}\")","metadata":{"id":"2ixikKmg2NDT","outputId":"c897adfc-87b7-486b-dd34-4665011fb54f","trusted":true,"execution":{"iopub.status.busy":"2024-12-17T16:34:12.429094Z","iopub.execute_input":"2024-12-17T16:34:12.429384Z","iopub.status.idle":"2024-12-17T16:34:12.433997Z","shell.execute_reply.started":"2024-12-17T16:34:12.429359Z","shell.execute_reply":"2024-12-17T16:34:12.433109Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n\n---\n\n# 6. **Encoding the Target Variable**","metadata":{"id":"a_KXyiRVmTTt"}},{"cell_type":"markdown","source":"❐ **Objective:** Improve the performance of the machine learning model by representing categorical values in the target variable **y** as numbers using **Label Encoding**.","metadata":{"id":"k58a22UDYyFF"}},{"cell_type":"code","source":"from sklearn.preprocessing import LabelEncoder\n\nLB = LabelEncoder()\ny_enc = LB.fit_transform(y)","metadata":{"id":"IcW2awpihSO9","trusted":true,"execution":{"iopub.status.busy":"2024-12-17T16:34:12.438858Z","iopub.execute_input":"2024-12-17T16:34:12.439268Z","iopub.status.idle":"2024-12-17T16:34:12.591322Z","shell.execute_reply.started":"2024-12-17T16:34:12.439242Z","shell.execute_reply":"2024-12-17T16:34:12.590594Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n\n---\n\n# 7. **Splitting the Data**","metadata":{"id":"okPHhuvsmzc0"}},{"cell_type":"markdown","source":"\n❐ **Objective** : Separate the dataset into features (**X**) and target (**y**) for model preparation.\n\n❐ **Purpose** : Isolate independent variables (**X**) and the dependent variable (**y**) to streamline the modeling process.","metadata":{"id":"P9Mms41HY4jf"}},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split\n\nX_train, X_val, y_train, y_val = train_test_split(X_seq, y_enc, test_size=0.2, random_state=42)","metadata":{"id":"AKtbbM3bm03v","trusted":true,"execution":{"iopub.status.busy":"2024-12-17T16:34:12.592351Z","iopub.execute_input":"2024-12-17T16:34:12.592627Z","iopub.status.idle":"2024-12-17T16:34:12.840386Z","shell.execute_reply.started":"2024-12-17T16:34:12.592601Z","shell.execute_reply":"2024-12-17T16:34:12.839665Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n\n---\n\n# 8. **Building an LSTM Model**","metadata":{"id":"TwnRgFd2nB_L"}},{"cell_type":"markdown","source":"- We are creating a `Sequential` neural network model using `Keras` to perform regression.\n","metadata":{"id":"XBMFFWFnb6b8"}},{"cell_type":"markdown","source":"**1- Importing Layers and Necessary Libraries**","metadata":{"id":"tK5Z4xXgibo_"}},{"cell_type":"code","source":"import tensorflow as tf\nfrom tensorflow.keras.layers import LSTM\nfrom tensorflow.keras.layers import Dense\nfrom tensorflow.keras import regularizers\nfrom tensorflow.keras.layers import Dropout\nfrom tensorflow.keras.optimizers import Adam\nfrom tensorflow.keras.layers import Embedding\nfrom tensorflow.keras.models import Sequential\nfrom tensorflow.keras.preprocessing.text import Tokenizer","metadata":{"id":"uN-H49fiGoa3","trusted":true,"execution":{"iopub.status.busy":"2024-12-17T16:34:12.841326Z","iopub.execute_input":"2024-12-17T16:34:12.841590Z","iopub.status.idle":"2024-12-17T16:34:12.851244Z","shell.execute_reply.started":"2024-12-17T16:34:12.841549Z","shell.execute_reply":"2024-12-17T16:34:12.850438Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**2- Preparing Word Embeddings and Tokenization**","metadata":{"id":"TZNG6YEf_9Sx"}},{"cell_type":"code","source":"embedding_dim = 100\nword_index = tokenizer.word_index\nnum_words = min(len(word_index) + 1, 10000)","metadata":{"id":"FeUKm0WW1G_h","trusted":true,"execution":{"iopub.status.busy":"2024-12-17T16:34:12.852254Z","iopub.execute_input":"2024-12-17T16:34:12.852500Z","iopub.status.idle":"2024-12-17T16:34:12.865486Z","shell.execute_reply.started":"2024-12-17T16:34:12.852476Z","shell.execute_reply":"2024-12-17T16:34:12.864833Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**3- Building the Model**","metadata":{"id":"X79CYtZBr80v"}},{"cell_type":"code","source":"model = Sequential()\n\nmodel.add(Embedding(input_dim=num_words, output_dim=embedding_dim, input_length=80))\n\nmodel.add(LSTM(128, kernel_regularizer=regularizers.l2(0.01)))\nmodel.add(Dropout(0.3))\n\nmodel.add(Dense(64, activation='relu', kernel_regularizer=regularizers.l2(0.01)))\nmodel.add(Dropout(0.3))\n\nmodel.add(Dense(1, activation='sigmoid'))","metadata":{"id":"S0o-Ckx3rxOz","outputId":"14c10ed9-3f57-45f0-aec0-5331235acdb1","trusted":true,"execution":{"iopub.status.busy":"2024-12-17T16:34:12.866445Z","iopub.execute_input":"2024-12-17T16:34:12.866697Z","iopub.status.idle":"2024-12-17T16:34:13.586911Z","shell.execute_reply.started":"2024-12-17T16:34:12.866673Z","shell.execute_reply":"2024-12-17T16:34:13.586225Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**4- Compiling the Model**","metadata":{"id":"NkY9W6TTtesA"}},{"cell_type":"code","source":"my_optimizer = Adam(learning_rate=0.0001)\n\nmodel.compile(loss='binary_crossentropy',\n              optimizer=my_optimizer,\n              metrics=['accuracy'])","metadata":{"id":"8c8dTHcTteSm","outputId":"9cffe74e-3467-43fa-b501-11b0e1ef8ad2","trusted":true,"execution":{"iopub.status.busy":"2024-12-17T16:34:13.587838Z","iopub.execute_input":"2024-12-17T16:34:13.588070Z","iopub.status.idle":"2024-12-17T16:34:13.604936Z","shell.execute_reply.started":"2024-12-17T16:34:13.588046Z","shell.execute_reply":"2024-12-17T16:34:13.604338Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**5- Setting Up EarlyStopping and Learning Rate Reduction Callbacks**","metadata":{"id":"zDQCZzodAW5w"}},{"cell_type":"code","source":"from tensorflow.keras.callbacks import EarlyStopping\nfrom tensorflow.keras.callbacks import ReduceLROnPlateau\n\nearly_stopping = EarlyStopping(monitor='val_loss',\n                               patience=3,\n                               restore_best_weights=True)\n\nreduce_lr = ReduceLROnPlateau(monitor='val_loss', factor=0.5, patience=3, min_lr=0.0001)","metadata":{"id":"6NLQ6CD4AQ9z","trusted":true,"execution":{"iopub.status.busy":"2024-12-17T16:34:13.605740Z","iopub.execute_input":"2024-12-17T16:34:13.606015Z","iopub.status.idle":"2024-12-17T16:34:13.611779Z","shell.execute_reply.started":"2024-12-17T16:34:13.605991Z","shell.execute_reply":"2024-12-17T16:34:13.611134Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**6- Training the Model**","metadata":{"id":"5rdjnfKpt882"}},{"cell_type":"code","source":"history = model.fit(X_train, y_train,\n                    validation_data=(X_val, y_val),\n                    epochs=10,\n                    batch_size=64,\n                    callbacks=[early_stopping,reduce_lr])","metadata":{"id":"0W-Z6et0t9y4","outputId":"1b4ed04b-a071-4267-b6cf-6ab603736ba6","trusted":true,"execution":{"iopub.status.busy":"2024-12-17T16:34:13.612573Z","iopub.execute_input":"2024-12-17T16:34:13.612871Z","iopub.status.idle":"2024-12-17T16:58:17.988856Z","shell.execute_reply.started":"2024-12-17T16:34:13.612831Z","shell.execute_reply":"2024-12-17T16:58:17.988080Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**7- Model Evaluation: Loss and Accuracy**","metadata":{"id":"XhSduDHwdze0"}},{"cell_type":"code","source":"loss, accuracy = model.evaluate(X_val, y_val)\nprint(f\"Test Accuracy: {accuracy * 100:.2f}%\")","metadata":{"id":"kvCjJpuw3Qh4","trusted":true,"execution":{"iopub.status.busy":"2024-12-17T16:58:17.990003Z","iopub.execute_input":"2024-12-17T16:58:17.990243Z","iopub.status.idle":"2024-12-17T16:58:45.924380Z","shell.execute_reply.started":"2024-12-17T16:58:17.990219Z","shell.execute_reply":"2024-12-17T16:58:45.923507Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**8- Model Evaluation: F1-Score**","metadata":{"id":"XwJSXrD-t6oE"}},{"cell_type":"code","source":"y_val_pred = model.predict(X_val)\ny_val_pred = (y_val_pred > 0.5).astype(int).flatten()","metadata":{"id":"Fg2RA0b6snOl","trusted":true,"execution":{"iopub.status.busy":"2024-12-17T16:58:45.925465Z","iopub.execute_input":"2024-12-17T16:58:45.925735Z","iopub.status.idle":"2024-12-17T16:59:09.983769Z","shell.execute_reply.started":"2024-12-17T16:58:45.925709Z","shell.execute_reply":"2024-12-17T16:59:09.983038Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from sklearn.metrics import f1_score\n\nf1_val = f1_score(y_val, y_val_pred)\nprint(f\"F1-Score on Validation Data: {f1_val:.2f}\")","metadata":{"id":"tAatxs_On5Zs","trusted":true,"execution":{"iopub.status.busy":"2024-12-17T16:59:09.985594Z","iopub.execute_input":"2024-12-17T16:59:09.986374Z","iopub.status.idle":"2024-12-17T16:59:10.148851Z","shell.execute_reply.started":"2024-12-17T16:59:09.986330Z","shell.execute_reply":"2024-12-17T16:59:10.148012Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**9- Confusion Matrix**","metadata":{"id":"Qy1u2lNA3u8x"}},{"cell_type":"code","source":"from sklearn.metrics import confusion_matrix\n\ncm = confusion_matrix(y_val, y_val_pred)\nprint(\"Confusion Matrix:\\n\", cm)","metadata":{"id":"339yB10n3lnh","trusted":true,"execution":{"iopub.status.busy":"2024-12-17T16:59:10.150179Z","iopub.execute_input":"2024-12-17T16:59:10.150836Z","iopub.status.idle":"2024-12-17T16:59:10.221874Z","shell.execute_reply.started":"2024-12-17T16:59:10.150776Z","shell.execute_reply":"2024-12-17T16:59:10.221069Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**10-  Plotting Training Metrics**","metadata":{"id":"ZNRRQL77d5bx"}},{"cell_type":"markdown","source":"\n\n---\n\n# 9.**Loading Test Data**","metadata":{"id":"vWVZmXhfH3W-"}},{"cell_type":"code","source":"data_test = pd.read_csv('/kaggle/input/quora-insincere-questions-classification/test.csv')\ndata_test.head(5)","metadata":{"id":"94knEBpMIH1w","trusted":true,"execution":{"iopub.status.busy":"2024-12-17T16:59:10.222814Z","iopub.execute_input":"2024-12-17T16:59:10.223102Z","iopub.status.idle":"2024-12-17T16:59:11.264869Z","shell.execute_reply.started":"2024-12-17T16:59:10.223077Z","shell.execute_reply":"2024-12-17T16:59:11.263972Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n\n---\n\n# 10. **Preprocessing Test Data**","metadata":{"id":"22pZX2O7I4wv"}},{"cell_type":"markdown","source":"- 📍 Convert Text to Lowercase","metadata":{"id":"cLgkBn6CI7qd"}},{"cell_type":"code","source":"data_test['question_text'] = data_test['question_text'].apply(to_lowercase)","metadata":{"id":"mCsbsHtdI6e0","trusted":true,"execution":{"iopub.status.busy":"2024-12-17T16:59:11.266159Z","iopub.execute_input":"2024-12-17T16:59:11.266877Z","iopub.status.idle":"2024-12-17T16:59:11.414408Z","shell.execute_reply.started":"2024-12-17T16:59:11.266825Z","shell.execute_reply":"2024-12-17T16:59:11.413700Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"- 📍 Remove Special Characters\n","metadata":{"id":"2P7DBeQXI_ES"}},{"cell_type":"code","source":"data_test['question_text'] = data_test['question_text'].apply(remove_special_chars)","metadata":{"id":"1Kvx5eT9I_2D","trusted":true,"execution":{"iopub.status.busy":"2024-12-17T16:59:11.415392Z","iopub.execute_input":"2024-12-17T16:59:11.415654Z","iopub.status.idle":"2024-12-17T16:59:12.143838Z","shell.execute_reply.started":"2024-12-17T16:59:11.415628Z","shell.execute_reply":"2024-12-17T16:59:12.142918Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"- 📍 Remove Stop Words\n","metadata":{"id":"1hahzyMBJBos"}},{"cell_type":"code","source":"data_test['question_text'] = data_test['question_text'].apply(remove_stopwords)","metadata":{"id":"eEBpRozZJCLR","trusted":true,"execution":{"iopub.status.busy":"2024-12-17T16:59:12.145186Z","iopub.execute_input":"2024-12-17T16:59:12.145812Z","iopub.status.idle":"2024-12-17T16:59:13.019654Z","shell.execute_reply.started":"2024-12-17T16:59:12.145753Z","shell.execute_reply":"2024-12-17T16:59:13.018968Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"- 📍 Remove Numbers and URLs\n","metadata":{"id":"wg8qXqRuJGmu"}},{"cell_type":"code","source":"data_test['question_text'] = data_test['question_text'].apply(remove_numbers_and_urls)","metadata":{"id":"34sSEC7eWlZz","trusted":true,"execution":{"iopub.status.busy":"2024-12-17T16:59:13.020672Z","iopub.execute_input":"2024-12-17T16:59:13.021025Z","iopub.status.idle":"2024-12-17T16:59:13.875645Z","shell.execute_reply.started":"2024-12-17T16:59:13.020988Z","shell.execute_reply":"2024-12-17T16:59:13.874929Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"- 📍 Remove Extra Spaces\n","metadata":{"id":"k_iwVzQlJKdp"}},{"cell_type":"code","source":"data_test['question_text'] = data_test['question_text'].apply(remove_extra_spaces)","metadata":{"id":"jQnNjp0TWnav","trusted":true,"execution":{"iopub.status.busy":"2024-12-17T16:59:13.876591Z","iopub.execute_input":"2024-12-17T16:59:13.876835Z","iopub.status.idle":"2024-12-17T16:59:14.148518Z","shell.execute_reply.started":"2024-12-17T16:59:13.876811Z","shell.execute_reply":"2024-12-17T16:59:14.147847Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n\n---\n\n# 11. **Tokenizing and Padding Test Data**","metadata":{"id":"rygTZf6vJeK9"}},{"cell_type":"code","source":"X_test = data_test['question_text']\nX_test_seq = tokenizer.texts_to_sequences(X_test)\n\nX_test_seq = pad_sequences(X_test_seq, maxlen=80)","metadata":{"id":"iFdXD-0WWpXO","trusted":true,"execution":{"iopub.status.busy":"2024-12-17T16:59:14.149543Z","iopub.execute_input":"2024-12-17T16:59:14.149853Z","iopub.status.idle":"2024-12-17T16:59:18.212268Z","shell.execute_reply.started":"2024-12-17T16:59:14.149811Z","shell.execute_reply":"2024-12-17T16:59:18.211066Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n\n---\n\n# 12. **Model Predictions on Test Data**","metadata":{"id":"nwh5yHkCKHiL"}},{"cell_type":"code","source":"y_test_pred = model.predict(X_test_seq)","metadata":{"id":"IrsRzP0QWrdE","trusted":true,"execution":{"iopub.status.busy":"2024-12-17T16:59:18.213507Z","iopub.execute_input":"2024-12-17T16:59:18.213902Z","iopub.status.idle":"2024-12-17T16:59:52.882606Z","shell.execute_reply.started":"2024-12-17T16:59:18.213855Z","shell.execute_reply":"2024-12-17T16:59:52.881857Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"- Converting Predictions to Binary Labels","metadata":{"id":"F5zD3BWfKTkP"}},{"cell_type":"code","source":"y_test_pred = (y_test_pred > 0.5).astype(int)","metadata":{"id":"H3bOAwQVWtAj","trusted":true,"execution":{"iopub.status.busy":"2024-12-17T16:59:52.883968Z","iopub.execute_input":"2024-12-17T16:59:52.884262Z","iopub.status.idle":"2024-12-17T16:59:52.889117Z","shell.execute_reply.started":"2024-12-17T16:59:52.884228Z","shell.execute_reply":"2024-12-17T16:59:52.888283Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"y_test_pred","metadata":{"id":"KWTuA15DWvD2","trusted":true,"execution":{"iopub.status.busy":"2024-12-17T16:59:52.890181Z","iopub.execute_input":"2024-12-17T16:59:52.890603Z","iopub.status.idle":"2024-12-17T16:59:52.901530Z","shell.execute_reply.started":"2024-12-17T16:59:52.890564Z","shell.execute_reply":"2024-12-17T16:59:52.900765Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"- Reshaping Predictions to 1D","metadata":{"id":"iBvNAk7IMIjP"}},{"cell_type":"code","source":"y_test_pred = y_test_pred.flatten()","metadata":{"id":"XVknGeGaMJsf","trusted":true,"execution":{"iopub.status.busy":"2024-12-17T16:59:52.902479Z","iopub.execute_input":"2024-12-17T16:59:52.902833Z","iopub.status.idle":"2024-12-17T16:59:52.911684Z","shell.execute_reply.started":"2024-12-17T16:59:52.902768Z","shell.execute_reply":"2024-12-17T16:59:52.910841Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"y_test_pred.shape","metadata":{"id":"icDVIH6FNeta","trusted":true,"execution":{"iopub.status.busy":"2024-12-17T16:59:52.912671Z","iopub.execute_input":"2024-12-17T16:59:52.912929Z","iopub.status.idle":"2024-12-17T16:59:52.923022Z","shell.execute_reply.started":"2024-12-17T16:59:52.912905Z","shell.execute_reply":"2024-12-17T16:59:52.922271Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\n\n---\n\n# 13. **Preparing the Submission File**","metadata":{"id":"0u68E7LUMKca"}},{"cell_type":"code","source":"submission = pd.read_csv('/kaggle/input/quora-insincere-questions-classification/sample_submission.csv')\nsubmission","metadata":{"id":"JFewEqd17pJ6","trusted":true,"execution":{"iopub.status.busy":"2024-12-17T16:59:52.923807Z","iopub.execute_input":"2024-12-17T16:59:52.924034Z","iopub.status.idle":"2024-12-17T16:59:53.256903Z","shell.execute_reply.started":"2024-12-17T16:59:52.924011Z","shell.execute_reply":"2024-12-17T16:59:53.255924Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"submission.info()","metadata":{"id":"Wqq1fy8zBWVD","trusted":true,"execution":{"iopub.status.busy":"2024-12-17T16:59:53.258003Z","iopub.execute_input":"2024-12-17T16:59:53.258324Z","iopub.status.idle":"2024-12-17T16:59:53.294713Z","shell.execute_reply.started":"2024-12-17T16:59:53.258295Z","shell.execute_reply":"2024-12-17T16:59:53.293848Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"submission['prediction'] = y_test_pred\n\nsubmission","metadata":{"id":"XctL32w6BedL","trusted":true,"execution":{"iopub.status.busy":"2024-12-17T16:59:53.295714Z","iopub.execute_input":"2024-12-17T16:59:53.295998Z","iopub.status.idle":"2024-12-17T16:59:53.306456Z","shell.execute_reply.started":"2024-12-17T16:59:53.295972Z","shell.execute_reply":"2024-12-17T16:59:53.305720Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"submission.to_csv('/kaggle/working/submission.csv', index=False)","metadata":{"id":"SMyXQD1xPv7Z","trusted":true,"execution":{"iopub.status.busy":"2024-12-17T16:59:53.312104Z","iopub.execute_input":"2024-12-17T16:59:53.312410Z","iopub.status.idle":"2024-12-17T16:59:53.647717Z","shell.execute_reply.started":"2024-12-17T16:59:53.312383Z","shell.execute_reply":"2024-12-17T16:59:53.646770Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### This work is prepared by:\n\n**Rami Abdullah**","metadata":{"id":"llkVbtZvtqga"}}]}