{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Classification of Quora Insincere Questions Using Transformers\n\n## Introduction\nThe Quora Insincere Questions Classification task revolves around identifying and categorizing inappropriate or insincere questions on the Quora platform. Quora, a popular Q&A website, hosts diverse discussions, but not all questions are genuine. Some may be disrespectful, offensive, or intended to provoke. The aim here is to create a machine learning model that can automatically label and distinguish such insincere questions.\n\n## The Dataset\nThe dataset provided for this task comprises questions sourced from Quora, labeled as either \"insincere\" or \"sincere.\" Insincere questions are those that exhibit characteristics of being disrespectful, offensive, or lacking genuine intent. On the other hand, sincere questions are those that are posed with the intention of seeking information or understanding.\n\nThis dataset includes text data representing the questions, alongside binary labels indicating whether each question is insincere or sincere. It follows the familiar pattern of supervised machine learning, where the objective is to train a model to discern patterns and features from the text that differentiate between sincere and insincere questions.\n\n## Problem Statement\nThe Quora Insincere Questions Classification problem is essentially a text classification task. Given a question, the goal is to predict whether the question is insincere or sincere. This task falls under the umbrella of natural language processing (NLP) and sentiment analysis. It involves interpreting the sentiment, tone, or intent behind text.\n\n## Steps to Solve the Problem\n\n### 1. Data Collection and Preprocessing\n   - Gather a dataset containing labeled Quora questions (sincere/insincere).\n   - Clean and tokenize the text by removing special characters, converting to lowercase, and splitting into tokens.\n\n### 2. Choose a Transformer Model\n   - Select a pre-trained transformer model suitable for text classification (e.g., BERT, RoBERTa).\n   - Utilize libraries like Hugging Face's `transformers` to load these pre-trained models.\n\n### 3. Data Formatting\n   - Convert preprocessed text data into a format compatible with the chosen transformer model.\n   - Tokenize text and transform it into numerical inputs that the model understands.\n\n### 4. Fine-Tuning the Model\n   - Load the pre-trained transformer model and add a classification layer on top.\n   - Customize the classification layer to have the desired number of output classes (2 for sincere/insincere).\n   - Define the loss function (e.g., Cross-Entropy) and optimizer (e.g., Adam) for training.\n\n### 5. Data Splitting and Training\n   - Split the dataset into training and validation sets.\n   - Train the model on the training data while fine-tuning the pre-trained transformer's weights.\n   - Monitor validation performance to avoid overfitting.\n\n### 6. Hyperparameter Tuning\n   - Experiment with hyperparameters like learning rate, batch size, and training epochs.\n   - Implement techniques such as learning rate schedules and gradient clipping for better training stability.\n\n### 7. Model Evaluation\n   - Assess the model's performance on the validation set using appropriate metrics (accuracy, precision, recall, F1-score, etc.).\n   - Adjust the model and hyperparameters based on validation results.\n\n### 8. Model Interpretation (Optional)\n   - Employ visualization techniques like attention visualization or class activation maps to understand influential parts of the input text.\n\n### 9. Inference and Deployment\n   - Use the trained model to predict whether new Quora questions are sincere or insincere.\n   - Deploy the model in suitable environments such as web applications or APIs.\n\n### 10. Continuous Improvement\n   - Regularly retrain the model with new data to enhance performance over time.","metadata":{}},{"cell_type":"code","source":"import numpy as np \nimport pandas as pd \nimport warnings\nwarnings.filterwarnings(\"ignore\")\nimport matplotlib.pyplot as plt\nimport tensorflow as tf\nimport tensorflow_hub as hub\nimport tensorflow_text as text\nfrom sklearn.model_selection import train_test_split\nfrom imblearn.under_sampling import RandomUnderSampler\nimport seaborn as sns\n%matplotlib inline","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-08-21T14:03:11.292323Z","iopub.execute_input":"2023-08-21T14:03:11.292694Z","iopub.status.idle":"2023-08-21T14:03:11.300337Z","shell.execute_reply.started":"2023-08-21T14:03:11.292663Z","shell.execute_reply":"2023-08-21T14:03:11.299235Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train = pd.read_csv('/kaggle/input/quora-insincere-questions-classification/train.csv')\ntest = pd.read_csv('/kaggle/input/quora-insincere-questions-classification/test.csv')","metadata":{"execution":{"iopub.status.busy":"2023-08-21T13:57:16.874628Z","iopub.execute_input":"2023-08-21T13:57:16.875240Z","iopub.status.idle":"2023-08-21T13:57:22.182376Z","shell.execute_reply.started":"2023-08-21T13:57:16.875209Z","shell.execute_reply":"2023-08-21T13:57:22.181388Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.head()","metadata":{"execution":{"iopub.status.busy":"2023-08-21T13:57:23.205961Z","iopub.execute_input":"2023-08-21T13:57:23.206315Z","iopub.status.idle":"2023-08-21T13:57:23.225266Z","shell.execute_reply.started":"2023-08-21T13:57:23.206285Z","shell.execute_reply":"2023-08-21T13:57:23.224230Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test.head()","metadata":{"execution":{"iopub.status.busy":"2023-08-21T13:57:24.807762Z","iopub.execute_input":"2023-08-21T13:57:24.808109Z","iopub.status.idle":"2023-08-21T13:57:24.818665Z","shell.execute_reply.started":"2023-08-21T13:57:24.808081Z","shell.execute_reply":"2023-08-21T13:57:24.817209Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Sincere Questions\ntrain[train.target==0][:5].question_text.to_list()","metadata":{"execution":{"iopub.status.busy":"2023-08-21T13:57:25.815892Z","iopub.execute_input":"2023-08-21T13:57:25.816238Z","iopub.status.idle":"2023-08-21T13:57:25.923382Z","shell.execute_reply.started":"2023-08-21T13:57:25.816210Z","shell.execute_reply":"2023-08-21T13:57:25.922431Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Insincere Questions\ntrain[train.target==1][:5].question_text.to_list()","metadata":{"execution":{"iopub.status.busy":"2023-08-21T13:57:26.530996Z","iopub.execute_input":"2023-08-21T13:57:26.531338Z","iopub.status.idle":"2023-08-21T13:57:26.559354Z","shell.execute_reply.started":"2023-08-21T13:57:26.531311Z","shell.execute_reply":"2023-08-21T13:57:26.558282Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(22,5), dpi=100)\nplt.rcParams['font.size'] = 12\n\npalette = ['#38A7D0', '#F67088']\nplt.figure(figsize=(22,5), dpi=100)\nax = sns.countplot(data =train,x=train['target'], hue=train['target'], dodge=False, palette=palette)\n\nfor bar in ax.patches:\n    ax.annotate(format(bar.get_height(), '.0f'),(bar.get_x() + bar.get_width() / 2,bar.get_height()), \n                 ha='center', va='center',size=15, xytext=(0, 8),textcoords='offset points')\nplt.gca().spines['top'].set_visible(False)\nplt.gca().spines['right'].set_visible(False)\nplt.title('Target Class Count Plot', fontsize=20)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-21T13:57:27.280325Z","iopub.execute_input":"2023-08-21T13:57:27.280902Z","iopub.status.idle":"2023-08-21T13:57:28.774520Z","shell.execute_reply.started":"2023-08-21T13:57:27.280870Z","shell.execute_reply":"2023-08-21T13:57:28.773419Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<center>\n<div class=\"alert alert-block alert-danger\">\nAs we can see there is a class imbalance in our dataset.\n</div>\n    </center>\n\n\n## What's CLASS IMBALANCE?\nClass imbalance is a common challenge encountered in many machine learning and data analysis scenarios, where the distribution of classes within a dataset is significantly skewed. In such cases, one class is often overrepresented, while the other is underrepresented, which can lead to suboptimal model performance and biased predictions.\n\nDealing with class imbalance is crucial for building accurate and fair predictive models. Ignoring class imbalance can result in models that are biased towards the majority class, leading to poor generalization and low recall for the minority class. Here are some strategies to address class imbalance:\n\n1. **Resampling Techniques**:\n   - **Oversampling**: This involves creating additional instances of the minority class by duplicating existing data points or generating synthetic examples using techniques like SMOTE (Synthetic Minority Over-sampling Technique).\n   - **Undersampling**: This entails reducing the number of instances in the majority class, which may involve randomly removing data points.\n\n2. **Cost-sensitive Learning**:\n   Introduce class-specific weights or costs during model training to penalize misclassifications of the minority class more heavily. This encourages the model to focus on improving its performance on the underrepresented class.\n\n3. **Algorithm Selection**:\n   Opt for algorithms that are less sensitive to class imbalance, such as tree-based algorithms like Random Forest or ensemble methods like AdaBoost.\n\n4. **Anomaly Detection**:\n   Treat the minority class as an anomaly detection problem and use techniques like One-Class SVM to identify instances that deviate significantly from the majority class.\n\n5. **Evaluation Metrics**:\n   Use appropriate evaluation metrics like precision, recall, F1-score, and area under the precision-recall curve (AUC-PR) instead of accuracy, which can be misleading in imbalanced settings.\n\n6. **Ensemble Methods**:\n   Combine multiple models to enhance predictive performance. Techniques like EasyEnsemble and BalancedBagging aim to balance class distribution within the ensemble.\n\n7. **Data Augmentation**:\n   Introduce variations to existing data points in the minority class, such as perturbing features, to diversify the dataset.\n\n8. **Collect More Data**:\n   If feasible, gather additional data for the minority class to balance out the distribution.\n\n9. **Threshold Adjustment**:\n   Adjust the decision threshold of the classifier to favor the minority class, depending on the specific problem's requirements.\n\nDealing with class imbalance necessitates a thoughtful and context-specific approach. Careful consideration of the problem domain, dataset characteristics, and chosen strategy is crucial to achieve effective and equitable model outcomes.","metadata":{}},{"cell_type":"code","source":"X = np.array(train.question_text).reshape(-1, 1)\ny = np.array(train.target)","metadata":{"execution":{"iopub.status.busy":"2023-08-21T13:57:28.776292Z","iopub.execute_input":"2023-08-21T13:57:28.776621Z","iopub.status.idle":"2023-08-21T13:57:28.812584Z","shell.execute_reply.started":"2023-08-21T13:57:28.776594Z","shell.execute_reply":"2023-08-21T13:57:28.811446Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's go ahead and Resample our data to get rid of the class imbalance.\n\n## **RANDOM UNDDERSAMPLING**\n\nRandom undersampling is a technique used in machine learning to address class imbalance in a dataset. Class imbalance occurs when the number of samples in one class (usually the minority class) is significantly lower than the number of samples in another class (usually the majority class). This imbalance can lead to biased model training and poor performance, especially for the minority class. Random undersampling is a way to balance the class distribution by reducing the number of samples in the majority class.\n\nHere's how random undersampling works:\n\n1. **Identify the Imbalance:** First, you need to recognize that there is a class imbalance issue in your dataset. This is usually evident when one class has substantially fewer instances compared to another class.\n\n2. **Random Selection:** In random undersampling, you randomly select a subset of samples from the majority class to match the number of samples in the minority class. The goal is to create a more balanced distribution of classes.\n\n3. **Reduction of Majority Class:** The selected subset of samples from the majority class is removed, effectively reducing the number of instances in that class.\n\n4. **Training with Balanced Data:** After random undersampling, you have a balanced dataset with equal or approximately equal representation of both classes. You can then train your machine learning model on this balanced dataset.\n\n5. **Potential Challenges:** While random undersampling can help mitigate the effects of class imbalance, it also has some drawbacks. By removing samples from the majority class, you risk losing valuable information that might be important for training a robust model. Additionally, if the dataset is extremely imbalanced, aggressive undersampling might result in a very small training set, which could negatively impact the model's ability to generalize.\n\n6. **Evaluation and Performance:** After training, it's essential to evaluate your model's performance using metrics such as precision, recall, F1-score, or area under the ROC curve. These metrics provide insights into how well the model is handling both classes, especially the minority class.\n\nIt's important to note that random undersampling is just one approach to address class imbalance. Other techniques, such as oversampling the minority class or using more advanced methods like Synthetic Minority Over-sampling Technique (SMOTE), can also be considered.","metadata":{}},{"cell_type":"markdown","source":"Once resampled we are creating a new dataframe and putting the resampled data into it since DataFrames are convenient to work with.","metadata":{}},{"cell_type":"code","source":"rus = RandomUnderSampler(random_state=0)\nX_resampled, y_resampled = rus.fit_resample(X,y)\ndata  = {\n    'question_text':np.squeeze(X_resampled),\n    'target': y_resampled\n}\ntrain_resampled = pd.DataFrame(data)\ntrain_resampled.head()","metadata":{"execution":{"iopub.status.busy":"2023-08-21T13:57:30.768656Z","iopub.execute_input":"2023-08-21T13:57:30.769010Z","iopub.status.idle":"2023-08-21T13:57:31.141988Z","shell.execute_reply.started":"2023-08-21T13:57:30.768982Z","shell.execute_reply":"2023-08-21T13:57:31.140886Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n<center>\n<div class=\"alert alert-block alert-success\">\nAnd VOILA!! The classes are now balanced as expected and visible below:-\n</div>\n    </center>\n","metadata":{}},{"cell_type":"code","source":"palette = ['#38A7D0', '#F67088']\nplt.figure(figsize=(22,5), dpi=100)\nax = sns.countplot(data =train_resampled,x=train_resampled['target'], hue=train_resampled['target'], \n                dodge=False, palette=palette)\n\nfor bar in ax.patches:\n    ax.annotate(format(bar.get_height(), '.0f'),(bar.get_x() + bar.get_width() / 2,bar.get_height()), \n                 ha='center', va='center',size=15, xytext=(0, 8),textcoords='offset points')\nplt.gca().spines['top'].set_visible(False)\nplt.gca().spines['right'].set_visible(False)\nplt.title('Resampled Target Class Count Plot', fontsize=20)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-21T13:57:32.257279Z","iopub.execute_input":"2023-08-21T13:57:32.257667Z","iopub.status.idle":"2023-08-21T13:57:32.667333Z","shell.execute_reply.started":"2023-08-21T13:57:32.257635Z","shell.execute_reply":"2023-08-21T13:57:32.666432Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now let's create our encoder and preprocesser using TensorFlow Hub to load two important components of the BERT (Bidirectional Encoder Representations from Transformers) model: the preprocessing layer and the encoder layer.\n\n**1. `bert_preprocess`:** This will load a BERT preprocessing layer using TensorFlow Hub. The preprocessing layer is responsible for preparing the input text data in a format that the BERT model can understand and process effectively. It takes care of tasks like tokenization (breaking text into individual words or subwords), padding, and creating input masks. By using this preprocessing layer, you're ensuring that the input text is properly formatted before being passed to the BERT model for further processing.\n\n**2. `bert_encoder`:** This will load the core BERT encoder layer using TensorFlow Hub. The encoder is the heart of the BERT model, responsible for understanding the context and meaning of words in the input text. It transforms the tokenized and preprocessed input text into contextualized embeddings (numerical representations) that capture the relationships between words. These embeddings are crucial for downstream tasks like text classification, sentiment analysis, and more. The specific version of BERT being loaded has 12 layers, a vocabulary size of 30,000, and embeddings of size 768.\n\nBy combining the preprocessing and encoder layers, we're setting up the infrastructure to process raw text data using BERT. The workflow would generally involve passing your input text through the `bert_preprocess` layer to get it into the right format, then forwarding the preprocessed text through the `bert_encoder` to obtain contextualized word embeddings. These embeddings can then be used for various NLP tasks or fine-tuned for specific applications.","metadata":{}},{"cell_type":"code","source":"bert_preprocess = hub.KerasLayer('https://tfhub.dev/tensorflow/bert_en_uncased_preprocess/3')\nbert_encoder = hub.KerasLayer(\"https://tfhub.dev/tensorflow/bert_en_uncased_L-12_H-768_A-12/4\")","metadata":{"execution":{"iopub.status.busy":"2023-08-21T14:02:13.453317Z","iopub.execute_input":"2023-08-21T14:02:13.453703Z","iopub.status.idle":"2023-08-21T14:02:40.936816Z","shell.execute_reply.started":"2023-08-21T14:02:13.453673Z","shell.execute_reply":"2023-08-21T14:02:40.935801Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Read More on **Tensorflow** Blog [Here](https://blog.tensorflow.org/2020/12/making-bert-easier-with-preprocessing-models-from-tensorflow-hub.html) and here is a [notebook](https://www.tensorflow.org/text/tutorials/classify_text_with_bert) you may want to look at ","metadata":{}},{"cell_type":"markdown","source":"## Embeddings\n\nBERT relies heavily on Embeddings.So, what exactly are Embeddings? \n\nEmbeddings, in the context of natural language processing (NLP), refer to numerical representations of words or other textual elements in a continuous vector space. These representations capture the semantic and contextual relationships between words, making it easier for machine learning models to process and understand text data.\n\n**Word Embeddings:**\n\nWord embeddings are a fundamental concept in NLP. They represent words as dense vectors in a high-dimensional space, where similar words are located close to each other. Each dimension in the vector space corresponds to a specific linguistic feature or context. For example, in a well-trained word embedding, words with similar meanings or roles might have similar vector coordinates.\n\n**BERT (Bidirectional Encoder Representations from Transformers) and Text Classification:**\n\nBERT takes word embeddings to the next level. Instead of just considering words in isolation, BERT generates contextualized word embeddings by considering the entire context of a word within a sentence. It does this by training on a massive amount of text data in a bidirectional manner, meaning it looks at both the left and right context of a word to generate its embedding.\n\nThis contextualization is crucial for tasks like text classification, where understanding the relationships and nuances between words in a sentence is essential. For example, in sentiment analysis (classifying whether a text expresses a positive or negative sentiment), understanding negations like \"not good\" or idiomatic expressions like \"piece of cake\" requires considering the surrounding context.\n\nHere's how BERT's embeddings play a role in text classification:\n\n1. **Contextual Understanding:** BERT generates embeddings that capture the context and meaning of words in a sentence. This means that each word's embedding takes into account the entire sentence it appears in. This is extremely beneficial for capturing the complex and nuanced meaning of text.\n\n2. **Input Representation:** When using BERT for text classification, you provide the entire sentence as input. BERT processes the sentence and generates embeddings for each word. These contextualized word embeddings form the foundation of the input representation for the downstream tasks like classification.\n\n3. **Attention Mechanism:** BERT's architecture includes an attention mechanism that helps it weigh the importance of different words in relation to each other. This mechanism ensures that the embeddings capture not only individual word meanings but also the interactions between words.\n\n4. **Transfer Learning:** BERT is pre-trained on a large corpus of text data before fine-tuning it for specific tasks like text classification. This pre-training enables the model to learn a broad understanding of language and context, making it highly effective for a wide range of NLP tasks.\n\n5. **Improved Performance:** Because BERT's embeddings capture context, nuances, and relationships, it performs exceptionally well on text classification tasks. It can understand subtle language patterns, idiomatic expressions, and the overall tone of the text, leading to improved classification accuracy.\n\nLet's go ahead and create a function that will give us the embeddings which will in turn help us convert the text into numbers which is what is required by any machine learning model to work.","metadata":{}},{"cell_type":"code","source":"def get_sentence_embeding(sentences):\n    preprocessed_text = bert_preprocess(sentences)\n    return bert_encoder(preprocessed_text)['pooled_output']\n\nget_sentence_embeding([\n    \"AI is making great strides recently especially the LLM space\", \n    \"India is a great country with so much diversity and so many cultures\"]\n)","metadata":{"execution":{"iopub.status.busy":"2023-08-20T18:26:32.537043Z","iopub.status.idle":"2023-08-20T18:26:32.537793Z","shell.execute_reply.started":"2023-08-20T18:26:32.537540Z","shell.execute_reply":"2023-08-20T18:26:32.537565Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's now split the data into train and test set. The test set that we already have will be then considered as the validation set after testing the model on the then splitted data aka `X_test`","metadata":{}},{"cell_type":"code","source":"X_train, X_test, y_train, y_test = train_test_split(train_resampled['question_text'],train_resampled['target'])","metadata":{"execution":{"iopub.status.busy":"2023-08-20T18:26:32.539875Z","iopub.status.idle":"2023-08-20T18:26:32.540931Z","shell.execute_reply.started":"2023-08-20T18:26:32.540677Z","shell.execute_reply":"2023-08-20T18:26:32.540701Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Creating the model","metadata":{}},{"cell_type":"code","source":"# Bert layers\ntext_input = tf.keras.layers.Input(shape=(), dtype=tf.string, name='text')\npreprocessed_text = bert_preprocess(text_input)\noutputs = bert_encoder(preprocessed_text)\n\n# Neural network layers\nl = tf.keras.layers.Dropout(0.1, name=\"dropout\")(outputs['pooled_output'])\nl = tf.keras.layers.Dense(1, activation='sigmoid', name=\"output\")(l)\n\n# Use inputs and outputs to construct a final model\nmodel = tf.keras.Model(inputs=[text_input], outputs = [l])","metadata":{"execution":{"iopub.status.busy":"2023-08-20T18:26:32.542124Z","iopub.status.idle":"2023-08-20T18:26:32.542956Z","shell.execute_reply.started":"2023-08-20T18:26:32.542698Z","shell.execute_reply":"2023-08-20T18:26:32.542721Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.summary()","metadata":{"execution":{"iopub.status.busy":"2023-08-20T18:26:32.544486Z","iopub.status.idle":"2023-08-20T18:26:32.545342Z","shell.execute_reply.started":"2023-08-20T18:26:32.545091Z","shell.execute_reply":"2023-08-20T18:26:32.545115Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Model Architecture\n\nLet's go over the architecture for a little bit. The model's architecture is designed to process and analyze text data using a combination of two powerful techniques: BERT (Bidirectional Encoder Representations from Transformers) and a neural network.\n\n**1. BERT Layers:**\n- BERT is a language model that's really good at understanding the context and meaning of words in a sentence. It's like a super-smart language decoder that can figure out the relationships between words.\n- In this architecture, we start by giving BERT a piece of text as input. Think of it as feeding a sentence or paragraph into a machine.\n- The `text_input` is where you provide the text that you want the model to understand. It's like the initial door through which you let your text data in.\n- The `bert_preprocess` step is where the text gets prepared for BERT to work its magic. It's like giving BERT the tools it needs to understand the language.\n- The `bert_encoder` processes the preprocessed text and generates outputs that represent the text's meaning. You can think of this step as BERT's \"thinking process\" where it figures out what the text is saying.\n\n**2. Neural Network Layers:**\n- The outputs from the BERT layers are really useful, but we want to make them even more insightful for our specific task. That's where the neural network comes in.\n- Imagine you have a bunch of puzzle pieces (the BERT outputs), and you want to arrange them in a certain way to reveal a clear picture (meaningful insights). The neural network helps with this arrangement.\n- The `Dropout` layer is like a little helper that randomly turns off some puzzle pieces. This helps prevent the model from becoming too reliant on specific pieces and encourages it to learn a more general representation.\n- The `Dense` layer is where the rearranging and final interpretation happens. It takes the processed puzzle pieces, arranges them, and reveals a single value (a number between 0 and 1) that indicates the model's confidence about a certain aspect of the text.\n\n**3. Constructing the Final Model:**\n- After all these steps, we put everything together to create the complete model that can take text as input and provide a meaningful output.\n- The `text_input` is where you give the model the text you want to analyze.\n- The `model` you get as the end result is like a trained expert. You feed it text, and it gives you back an answer based on what it has learned from similar text it has seen during its training.\n\n## **Functional API**\n\nWe are using the Functional API here as opposed to the Sequential API In tensorflow. The Functional API is a powerful tool provided by TensorFlow (a popular machine learning framework) that lets us define and connect layers in a flexible and customizable way. In the architecture you've provided, the Functional API plays a crucial role in bringing all the different components together and creating a cohesive model.\n\nHere's how the Functional API is helpful in this case:\n\n**1. Customizable Architecture:**\nWith the Functional API, you have the freedom to create complex models with multiple input and output branches, shared layers, and more. In your example, you have a specific requirement: you want to process text using BERT, then pass the output through a neural network. The Functional API lets you easily design this custom architecture by allowing you to connect layers in a way that matches your needs.\n\n**2. Multiple Inputs and Outputs:**\nIn more advanced use cases, you might have models with multiple inputs or outputs. The Functional API handles these scenarios effortlessly. For instance, if you wanted to process two different pieces of text and combine their results, you could do so by defining multiple inputs and appropriately connecting them using the API.\n\n**3. Layer Reusability:**\nIn machine learning, reusing layers is common practice to improve efficiency and maintainability. The Functional API makes it straightforward to reuse layers across different parts of your model. This is particularly useful if you want to process the same text with multiple approaches or if you want to reuse parts of your architecture in different projects.\n\n**4. Clear Visualization:**\nThe Functional API enhances the visualization of your model. You can easily create diagrams showing how different layers are connected, making it easier to understand the flow of data through your model. This is especially useful when collaborating with others or when revisiting your model after a period of time.\n\n**5. Easier Debugging and Inspection:**\nWith the Functional API, you can inspect and debug individual layers and connections more effectively. This is valuable when you're fine-tuning your architecture or diagnosing issues that might arise during training.\n\nIn the architecture, the Functional API comes into play when you define the `model` by connecting the `text_input` to the BERT layers, then passing the BERT outputs through the neural network layers. The Functional API allows you to precisely orchestrate how data flows through these different components and create a well-organized, comprehensive model.\n\nLet's go ahead and compile the model now.","metadata":{}},{"cell_type":"code","source":"model.compile(optimizer='adam',loss='binary_crossentropy',metrics=['accuracy'])","metadata":{"execution":{"iopub.status.busy":"2023-08-20T18:26:32.546700Z","iopub.status.idle":"2023-08-20T18:26:32.547546Z","shell.execute_reply.started":"2023-08-20T18:26:32.547289Z","shell.execute_reply":"2023-08-20T18:26:32.547313Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Training the Model\n\nTime for the fun stuff!!","metadata":{}},{"cell_type":"code","source":"history = model.fit(X_train, y_train, epochs=3, batch_size= 100)","metadata":{"execution":{"iopub.status.busy":"2023-08-20T18:26:32.548914Z","iopub.status.idle":"2023-08-20T18:26:32.549740Z","shell.execute_reply.started":"2023-08-20T18:26:32.549490Z","shell.execute_reply":"2023-08-20T18:26:32.549513Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(22,5), dpi=100)\nplt.plot(range(len(history.history['loss'])),history.history['loss'] , marker='o')\nplt.legend(['loss'])\nplt.xlabel('Epochs')\nplt.xticks(range(len(history.history)+1))\nplt.ylabel('Loss')\nplt.title('Loss Plot', fontsize=12)\n#Turning off the top and right spines\nplt.gca().spines['top'].set_visible(False)\nplt.gca().spines['right'].set_visible(False)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-20T18:26:32.551128Z","iopub.status.idle":"2023-08-20T18:26:32.552002Z","shell.execute_reply.started":"2023-08-20T18:26:32.551708Z","shell.execute_reply":"2023-08-20T18:26:32.551731Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(22,5), dpi=100)\nplt.plot(range(len(history.history['accuracy'])),history.history['accuracy'] , marker='o', color='orange')\nplt.legend(['accuracy'])\nplt.xlabel('Epochs')\nplt.xticks(range(len(history.history)+1))\nplt.ylabel('Accuracy')\nplt.title('Accuracy Plot', fontsize=12)\n#Turning off the top and right spines\nplt.gca().spines['top'].set_visible(False)\nplt.gca().spines['right'].set_visible(False)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-08-20T18:26:32.553559Z","iopub.status.idle":"2023-08-20T18:26:32.555423Z","shell.execute_reply.started":"2023-08-20T18:26:32.555170Z","shell.execute_reply":"2023-08-20T18:26:32.555194Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Evaluating the Model\n\nLet's evaluate how the model performs on the `X_test` data","metadata":{}},{"cell_type":"code","source":"print(\"X_test Dataset Evaluation\")\nresult = model.evaluate(X_test, y_test)\ndict(zip(model.metrics_names, result))","metadata":{"execution":{"iopub.status.busy":"2023-08-20T18:26:32.556840Z","iopub.status.idle":"2023-08-20T18:26:32.557647Z","shell.execute_reply.started":"2023-08-20T18:26:32.557392Z","shell.execute_reply":"2023-08-20T18:26:32.557414Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's go ahead and validate the model on our `test` data","metadata":{}},{"cell_type":"code","source":"questions = test.question_text\nval_preds = model.predict(questions[:20])","metadata":{"execution":{"iopub.status.busy":"2023-08-20T18:26:32.559114Z","iopub.status.idle":"2023-08-20T18:26:32.559891Z","shell.execute_reply.started":"2023-08-20T18:26:32.559641Z","shell.execute_reply":"2023-08-20T18:26:32.559664Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Making Predictions on the Validation Dataset \\n')\nfor question, prediction in zip(questions[:20], val_preds):\n    prediction_label = 1 if prediction >= 0.5 else 0\n    print(f\"Question: {question}\\nPrediction: {prediction_label}\\n{'-' * 50}\")","metadata":{"execution":{"iopub.status.busy":"2023-08-20T18:26:32.561350Z","iopub.status.idle":"2023-08-20T18:26:32.562168Z","shell.execute_reply.started":"2023-08-20T18:26:32.561887Z","shell.execute_reply":"2023-08-20T18:26:32.561921Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Resources, Acknowledgments, and Conclusion\n\nAs we conclude this journey through the Quora Insincere Questions Classification project, I'd like to express my gratitude to the data science community, educators, and resources that have paved the way for this project's success. \n\nI extend heartfelt gratitude to **codebasics** for their dedication to educating the data science community and empowering learners to build practical and impactful projects. Their contributions have been instrumental in guiding the creation of this notebook, enriching the learning experience. To access codebasics' insightful content, you can visit their [GitHub repository](https://github.com/codebasics/deep-learning-keras-tf-tutorial)\n\nThis project has delved into the world of natural language processing, exploring techniques to process text data, leverage BERT models, and construct robust machine learning systems. Throughout this journey, we covered essential steps such as data collection, preprocessing, model selection, training, evaluation, and interpretation. Each phase provided an opportunity to learn and grow in our understanding of machine learning and NLP concepts.\n\nThank you for joining me on this exploration. Here's to continued learning and many more rewarding adventures ahead! \n\n<center>\n<div class=\"alert alert-block alert-info\">\nI trust this notebook has provided a solid foundation in grasping the essentials of regex, leaving you confident in crafting your own expressions. Your support through upvoting and following for more content like this is greatly appreciated.\n</div>\n    </center>","metadata":{}}]}