{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.7.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":19018,"databundleVersionId":2703900,"sourceType":"competition"},{"sourceId":11650,"sourceType":"datasetVersion","datasetId":8327}],"dockerImageVersionId":30299,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# About this Notebook\n\nNLP is a very hot topic right now and as belived by many experts '2020 is going to be NLP's Year' ,with its ever changing dynamics it is experiencing a boom , same as computer vision once did. Owing to its popularity Kaggle launched two NLP competitions recently and me being a lover of this Hot topic prepared myself to join in my first Kaggle Competition.<br><br>\nAs I joined the competitions and since I was a complete beginner with Deep Learning Techniques for NLP, all my enthusiasm took a beating when I saw everyone Using all  kinds of BERT , everything just went over my head,I thought to quit but there is a special thing about Kaggle ,it just hooks you. I thought I have to learn someday , why not now , so I braced myself and sat on the learning curve. I wrote a kernel on the Tweet Sentiment Extraction competition that has now got a gold medal , it can be viewed here : https://www.kaggle.com/tanulsingh077/twitter-sentiment-extaction-analysis-eda-and-model <br><br>\nAfter 10 days of extensive learning(finishing all the latest NLP approaches) , I am back here to share my leaning , by writing a kernel that starts from the very Basic RNN's to built over , all the way to BERT . I invite you all to come and learn alongside with me and take a step closer towards becoming an NLP expert","metadata":{}},{"cell_type":"markdown","source":"# Contents\n\nIn this Notebook I will start with the very Basics of RNN's and Build all the way to latest deep learning architectures to solve NLP problems. It will cover the Following:\n* Simple RNN's\n* Word Embeddings : Definition and How to get them\n* LSTM's\n* GRU's\n* BI-Directional RNN's\n* Encoder-Decoder Models (Seq2Seq Models)\n* Attention Models\n* Transformers - Attention is all you need\n* BERT\n\nI will divide every Topic into four subsections:\n* Basic Overview\n* In-Depth Understanding : In this I will attach links of articles and videos to learn about the topic in depth\n* Code-Implementation\n* Code Explanation\n\nThis is a comprehensive kernel and if you follow along till the end , I promise you would learn all the techniques completely\n\nNote that the aim of this notebook is not to have a High LB score but to present a beginner guide to understand Deep Learning techniques used for NLP. Also after discussing all of these ideas , I will present a starter solution for this competiton","metadata":{}},{"cell_type":"markdown","source":"**<span style=\"color:Red\">This kernel has been a work of more than 10 days If you find my kernel useful and my efforts appreciable, Please Upvote it , it motivates me to write more Quality content**","metadata":{}},{"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nfrom tqdm import tqdm\nfrom sklearn.model_selection import train_test_split\nimport tensorflow as tf\nfrom keras.models import Sequential\nfrom keras.layers.recurrent import LSTM, GRU,SimpleRNN\nfrom keras.layers.core import Dense, Activation, Dropout\nfrom keras.layers.embeddings import Embedding\nfrom keras.layers import BatchNormalization\nfrom keras.utils import np_utils\nfrom sklearn import preprocessing, decomposition, model_selection, metrics, pipeline\nfrom keras.layers import GlobalMaxPooling1D, Conv1D, MaxPooling1D, Flatten, Bidirectional, SpatialDropout1D\nfrom keras.preprocessing import sequence, text ###keras.preprocessing是处理文本的包，sequence, text是处理具体功能的模块\nfrom keras.callbacks import EarlyStopping\nfrom sklearn.metrics import accuracy_score\n\n\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n%matplotlib inline\nfrom plotly import graph_objs as go\nimport plotly.express as px\nimport plotly.figure_factory as ff","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2025-06-06T05:38:52.989546Z","iopub.execute_input":"2025-06-06T05:38:52.990414Z","iopub.status.idle":"2025-06-06T05:38:53.000546Z","shell.execute_reply.started":"2025-06-06T05:38:52.990379Z","shell.execute_reply":"2025-06-06T05:38:52.999431Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Detect hardware, return appropriate distribution strategy\ntry:\n    # TPU detection. No parameters necessary if TPU_NAME environment variable is\n    # set: this is always the case on Kaggle.\n    tpu = tf.distribute.cluster_resolver.TPUClusterResolver()\n    print('Running on TPU ', tpu.master())\nexcept ValueError:\n    tpu = None\n\nif tpu:\n    tf.config.experimental_connect_to_cluster(tpu)\n    tf.tpu.experimental.initialize_tpu_system(tpu)\n    strategy = tf.distribute.experimental.TPUStrategy(tpu)\nelse:\n    # Default distribution strategy in Tensorflow. Works on CPU and single GPU.\n    strategy = tf.distribute.get_strategy()\n\nprint(\"REPLICAS: \", strategy.num_replicas_in_sync)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-06T06:45:17.691995Z","iopub.execute_input":"2025-06-06T06:45:17.692934Z","iopub.status.idle":"2025-06-06T06:45:17.700131Z","shell.execute_reply.started":"2025-06-06T06:45:17.692898Z","shell.execute_reply":"2025-06-06T06:45:17.699110Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Read Data ","metadata":{}},{"cell_type":"code","source":"train = pd.read_csv('/kaggle/input/jigsaw-multilingual-toxic-comment-classification/jigsaw-toxic-comment-train.csv')\nvalidation = pd.read_csv('/kaggle/input/jigsaw-multilingual-toxic-comment-classification/validation.csv')\ntest = pd.read_csv('/kaggle/input/jigsaw-multilingual-toxic-comment-classification/test.csv')","metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true,"execution":{"iopub.status.busy":"2025-06-06T05:40:17.958669Z","iopub.execute_input":"2025-06-06T05:40:17.959263Z","iopub.status.idle":"2025-06-06T05:40:21.679221Z","shell.execute_reply.started":"2025-06-06T05:40:17.959222Z","shell.execute_reply":"2025-06-06T05:40:21.678379Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"We will drop the other columns and approach this problem as a Binary Classification Problem and also we will have our exercise done on a smaller subsection of the dataset(only 12000 data points) to make it easier to train the models","metadata":{}},{"cell_type":"code","source":"train.drop(['severe_toxic','obscene','threat','insult','identity_hate'],axis=1,inplace=True)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-06T05:40:32.885000Z","iopub.execute_input":"2025-06-06T05:40:32.885402Z","iopub.status.idle":"2025-06-06T05:40:32.913251Z","shell.execute_reply.started":"2025-06-06T05:40:32.885375Z","shell.execute_reply":"2025-06-06T05:40:32.912441Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train = train.loc[:12000,:]\ntrain.shape","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-06T05:40:40.485187Z","iopub.execute_input":"2025-06-06T05:40:40.486246Z","iopub.status.idle":"2025-06-06T05:40:40.494596Z","shell.execute_reply.started":"2025-06-06T05:40:40.486211Z","shell.execute_reply":"2025-06-06T05:40:40.493426Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-06T05:40:45.113893Z","iopub.execute_input":"2025-06-06T05:40:45.114694Z","iopub.status.idle":"2025-06-06T05:40:45.126957Z","shell.execute_reply.started":"2025-06-06T05:40:45.114664Z","shell.execute_reply":"2025-06-06T05:40:45.125736Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"We will check the maximum number of words that can be present in a comment , this will help us in padding later","metadata":{}},{"cell_type":"code","source":"train['comment_text'].apply(lambda x:len(str(x).split())).max()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-06T05:42:00.136335Z","iopub.execute_input":"2025-06-06T05:42:00.137104Z","iopub.status.idle":"2025-06-06T05:42:00.200421Z","shell.execute_reply.started":"2025-06-06T05:42:00.137069Z","shell.execute_reply":"2025-06-06T05:42:00.199448Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Writing a function for getting auc score for validation","metadata":{}},{"cell_type":"code","source":"def roc_auc(predictions,target):\n    '''\n    This methods returns the AUC Score when given the Predictions\n    and Labels\n    '''\n    \n    fpr, tpr, thresholds = metrics.roc_curve(target, predictions)\n    roc_auc = metrics.auc(fpr, tpr)\n    return roc_auc","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-06T05:44:59.699216Z","iopub.execute_input":"2025-06-06T05:44:59.699573Z","iopub.status.idle":"2025-06-06T05:44:59.704807Z","shell.execute_reply.started":"2025-06-06T05:44:59.699543Z","shell.execute_reply":"2025-06-06T05:44:59.703655Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Data Preparation","metadata":{}},{"cell_type":"code","source":"xtrain, xvalid, ytrain, yvalid = train_test_split(train.comment_text.values, train.toxic.values, \n                                                  stratify=train.toxic.values, \n                                                  random_state=42, \n                                                  test_size=0.2, shuffle=True)\n\n###stratify=train.toxic.values表示分层抽样\n###数据集中标签分布比例是3比7则分割得到的训练集和测试集中标签比例也是3比7","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-06T05:45:02.589509Z","iopub.execute_input":"2025-06-06T05:45:02.590412Z","iopub.status.idle":"2025-06-06T05:45:02.602325Z","shell.execute_reply.started":"2025-06-06T05:45:02.590369Z","shell.execute_reply":"2025-06-06T05:45:02.601375Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Before We Begin\n\nBefore we Begin If you are a complete starter with NLP and never worked with text data, I am attaching a few kernels that will serve as a starting point of your journey\n* https://www.kaggle.com/arthurtok/spooky-nlp-and-topic-modelling-tutorial\n* https://www.kaggle.com/abhishek/approaching-almost-any-nlp-problem-on-kaggle\n\nIf you want a more basic dataset to practice with here is another kernel which I wrote:\n* https://www.kaggle.com/tanulsingh077/what-s-cooking\n\nBelow are some Resources to get started with basic level Neural Networks, It will help us to easily understand the upcoming parts\n* https://www.youtube.com/watch?v=aircAruvnKk&list=PL_h2yd2CGtBHEKwEH5iqTZH85wLS-eUzv\n* https://www.youtube.com/watch?v=IHZwWFHWa-w&list=PL_h2yd2CGtBHEKwEH5iqTZH85wLS-eUzv&index=2\n* https://www.youtube.com/watch?v=Ilg3gGewQ5U&list=PL_h2yd2CGtBHEKwEH5iqTZH85wLS-eUzv&index=3\n* https://www.youtube.com/watch?v=tIeHLnjs5U8&list=PL_h2yd2CGtBHEKwEH5iqTZH85wLS-eUzv&index=4\n\nFor Learning how to visualize test data and what to use view:\n* https://www.kaggle.com/tanulsingh077/twitter-sentiment-extaction-analysis-eda-and-model\n* https://www.kaggle.com/jagangupta/stop-the-s-toxic-comments-eda","metadata":{}},{"cell_type":"markdown","source":"# Simple RNN\n\n## Basic Overview\n\nWhat is a RNN?\n\nRecurrent Neural Network(RNN) are a type of Neural Network where the output from previous step are fed as input to the current step. In traditional neural networks, all the inputs and outputs are independent of each other, but in cases like when it is required to predict the next word of a sentence, the previous words are required and hence there is a need to remember the previous words. Thus RNN came into existence, which solved this issue with the help of a Hidden Layer.\n\nWhy RNN's?\n\nhttps://www.quora.com/Why-do-we-use-an-RNN-instead-of-a-simple-neural-network\n\n## In-Depth Understanding\n\n* https://medium.com/mindorks/understanding-the-recurrent-neural-network-44d593f112a2\n* https://www.youtube.com/watch?v=2E65LDnM2cA&list=PL1F3ABbhcqa3BBWo170U4Ev2wfsF7FN8l\n* https://www.d2l.ai/chapter_recurrent-neural-networks/rnn.html\n\n## Code Implementation\n\nSo first I will implement the and then I will explain the code step by step","metadata":{}},{"cell_type":"code","source":"# using keras tokenizer here\ntoken = text.Tokenizer(num_words=None) ###使用词嵌入模块text的Tokenizer方法，用于向量化单词\nmax_len = 1500\n\ntoken.fit_on_texts(list(xtrain) + list(xvalid)) ###将训练集 xtrain 和验证集 xvalid 的文本合并，统计所有单词的频率并构建词典\n###{\"the\": 1, \"cat\": 2, \"sat\": 3, ...}\nxtrain_seq = token.texts_to_sequences(xtrain) ###将每条文本转换为对应的整数序列（根据 word_index 映射）。若 xtrain = [\"cat sat\", \"the cat\"]，word_index = {\"cat\":1, \"sat\":2, \"the\":3}，则输出：\n###xtrain_seq = [[1, 2], [3, 1]]\nxvalid_seq = token.texts_to_sequences(xvalid)\n\n#zero pad the sequences\nxtrain_pad = sequence.pad_sequences(xtrain_seq, maxlen=max_len)###将所有序列填充/截断为固定长度 max_len（此处为1500），确保输入维度一致。\nxvalid_pad = sequence.pad_sequences(xvalid_seq, maxlen=max_len)\n\nword_index = token.word_index","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-06T05:45:07.768141Z","iopub.execute_input":"2025-06-06T05:45:07.768504Z","iopub.status.idle":"2025-06-06T05:45:09.049626Z","shell.execute_reply.started":"2025-06-06T05:45:07.768475Z","shell.execute_reply":"2025-06-06T05:45:09.048722Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"len(word_index)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-06T06:09:46.490602Z","iopub.execute_input":"2025-06-06T06:09:46.491621Z","iopub.status.idle":"2025-06-06T06:09:46.497406Z","shell.execute_reply.started":"2025-06-06T06:09:46.491579Z","shell.execute_reply":"2025-06-06T06:09:46.496257Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# A simpleRNN without any pretrained embeddings and one dense layer\nmodel = Sequential()\nmodel.add(Embedding(len(word_index) + 1,300,input_length=max_len))\nmodel.add(SimpleRNN(100))\nmodel.add(Dense(1, activation='sigmoid'))\nmodel.compile(loss='binary_crossentropy', optimizer='adam', metrics=['accuracy'])\nmodel.summary()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-06T05:45:52.060546Z","iopub.execute_input":"2025-06-06T05:45:52.061447Z","iopub.status.idle":"2025-06-06T05:45:55.183727Z","shell.execute_reply.started":"2025-06-06T05:45:52.061412Z","shell.execute_reply":"2025-06-06T05:45:55.182837Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"词嵌入层，这个层构建了一个二维表，行表示单词，列表示单词的嵌入。比如本例子中，行数是43497，列数是300，则是一个43497*300的二维表\n该表中参数是学习出来的，初始化为随机值，通过交叉熵损失，优化参数，逐步降低该损失值。\nsimpleRNN是我们课堂上所讲的RNN的具体实现，参数100表示该层RNN中有100个RNN神经元，注意的是，每个RNN神经元在t时刻的隐藏状态会传递给\nt+1时刻所有的RNN神经元，所以参数量计算：300*100（输入到神经元）+100*100（隐状态到隐状态）+100*1（偏置）=0100","metadata":{}},{"cell_type":"code","source":"model.fit(xtrain_pad, ytrain, epochs=5, batch_size=64) #Multiplying by Strategy to run on TPU's","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-06T05:45:59.060924Z","iopub.execute_input":"2025-06-06T05:45:59.061607Z","iopub.status.idle":"2025-06-06T05:59:11.470414Z","shell.execute_reply.started":"2025-06-06T05:45:59.061575Z","shell.execute_reply":"2025-06-06T05:59:11.469245Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"###在验证集上的效果\nscores = model.predict(xvalid_pad)\npredicted_labels = (scores > 0.5).astype(int)  # 转为0/1标签\nprint(predicted_labels[:10])  # 输出前10个预测标签","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-06T06:31:30.560138Z","iopub.execute_input":"2025-06-06T06:31:30.560991Z","iopub.status.idle":"2025-06-06T06:31:37.491999Z","shell.execute_reply.started":"2025-06-06T06:31:30.560959Z","shell.execute_reply":"2025-06-06T06:31:37.491028Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"accuracy = accuracy_score(yvalid, predicted_labels)\nprint(f\"Accuracy: {accuracy:.4f}\")  # 输出格式化为4位小数\n#print(\"Auc: %.2f%%\" % (roc_auc(scores,yvalid)))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-06T06:32:28.039441Z","iopub.execute_input":"2025-06-06T06:32:28.040457Z","iopub.status.idle":"2025-06-06T06:32:28.046430Z","shell.execute_reply.started":"2025-06-06T06:32:28.040420Z","shell.execute_reply":"2025-06-06T06:32:28.045472Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"scores_model = []\nscores_model.append({'Model': 'SimpleRNN','AUC_Score': accuracy})","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-06T06:33:28.174278Z","iopub.execute_input":"2025-06-06T06:33:28.175236Z","iopub.status.idle":"2025-06-06T06:33:28.179420Z","shell.execute_reply.started":"2025-06-06T06:33:28.175199Z","shell.execute_reply":"2025-06-06T06:33:28.178408Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Code Explanantion\n* Tokenization<br><br>\n So if you have watched the videos and referred to the links, you would know that in an RNN we input a sentence word by word. We represent every word as one hot vectors of dimensions : Numbers of words in Vocab +1. <br>\n  What keras Tokenizer does is , it takes all the unique words in the corpus,forms a dictionary with words as keys and their number of occurences as values,it then sorts the dictionary in descending order of counts. It then assigns the first value 1 , second value 2 and so on. So let's suppose word 'the' occured the most in the corpus then it will assigned index 1 and vector representing 'the' would be a one-hot vector with value 1 at position 1 and rest zereos.<br>\n  Try printing first 2 elements of xtrain_seq you will see every word is represented as a digit now","metadata":{}},{"cell_type":"code","source":"xtrain_seq[:1]","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<b>Now you might be wondering What is padding? Why its done</b><br><br>\n\nHere is the answer :\n* https://www.quora.com/Which-effect-does-sequence-padding-have-on-the-training-of-a-neural-network\n* https://machinelearningmastery.com/data-preparation-variable-length-input-sequences-sequence-prediction/\n* https://www.coursera.org/lecture/natural-language-processing-tensorflow/padding-2Cyzs\n\nAlso sometimes people might use special tokens while tokenizing like EOS(end of string) and BOS(Begining of string). Here is the reason why it's done\n* https://stackoverflow.com/questions/44579161/why-do-we-do-padding-in-nlp-tasks\n\n\nThe code token.word_index simply gives the dictionary of vocab that keras created for us","metadata":{}},{"cell_type":"markdown","source":"* Building the Neural Network\n\nTo understand the Dimensions of input and output given to RNN in keras her is a beautiful article : https://medium.com/@shivajbd/understanding-input-and-output-shape-in-lstm-keras-c501ee95c65e\n\nThe first line model.Sequential() tells keras that we will be building our network sequentially . Then we first add the Embedding layer.\nEmbedding layer is also a layer of neurons which takes in as input the nth dimensional one hot vector of every word and converts it into 300 dimensional vector , it gives us word embeddings similar to word2vec. We could have used word2vec but the embeddings layer learns during training to enhance the embeddings.\nNext we add an 100 LSTM units without any dropout or regularization\nAt last we add a single neuron with sigmoid function which takes output from 100 LSTM cells (Please note we have 100 LSTM cells not layers) to predict the results and then we compile the model using adam optimizer \n\n* Comments on the model<br><br>\nWe can see our model achieves an accuracy of 1 which is just insane , we are clearly overfitting I know , but this was the simplest model of all ,we can tune a lot of hyperparameters like RNN units, we can do batch normalization , dropouts etc to get better result. The point is we got an AUC score of 0.82 without much efforts and we know have learnt about RNN's .Deep learning is really revolutionary","metadata":{}},{"cell_type":"markdown","source":"# Word Embeddings\n\nWhile building our simple RNN models we talked about using word-embeddings , So what is word-embeddings and how do we get word-embeddings?\nHere is the answer :\n* https://www.coursera.org/learn/nlp-sequence-models/lecture/6Oq70/word-representation\n* https://machinelearningmastery.com/what-are-word-embeddings/\n<br> <br>\nThe latest approach to getting word Embeddings is using pretained GLoVe or using Fasttext. Without going into too much details, I would explain how to create sentence vectors and how can we use them to create a machine learning model on top of it and since I am a fan of GloVe vectors, word2vec and fasttext. In this Notebook, I'll be using the GloVe vectors. You can download the GloVe vectors from here http://www-nlp.stanford.edu/data/glove.840B.300d.zip or you can search for GloVe in datasets on Kaggle and add the file","metadata":{}},{"cell_type":"code","source":"# load the GloVe vectors in a dictionary:\n\nembeddings_index = {}\nf = open('/kaggle/input/glove840b300dtxt/glove.840B.300d.txt','r',encoding='utf-8')\nfor line in tqdm(f):\n    values = line.split(' ')\n    word = values[0]\n    coefs = np.asarray([float(val) for val in values[1:]])\n    embeddings_index[word] = coefs\nf.close()\n\nprint('Found %s word vectors.' % len(embeddings_index))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-06T06:33:44.118389Z","iopub.execute_input":"2025-06-06T06:33:44.118742Z","iopub.status.idle":"2025-06-06T06:37:27.486382Z","shell.execute_reply.started":"2025-06-06T06:33:44.118713Z","shell.execute_reply":"2025-06-06T06:37:27.485387Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# LSTM's\n\n## Basic Overview\n\nSimple RNN's were certainly better than classical ML algorithms and gave state of the art results, but it failed to capture long term dependencies that is present in sentences . So in 1998-99 LSTM's were introduced to counter to these drawbacks.\n\n## In Depth Understanding\n\nWhy LSTM's?\n* https://www.coursera.org/learn/nlp-sequence-models/lecture/PKMRR/vanishing-gradients-with-rnns\n* https://www.analyticsvidhya.com/blog/2017/12/fundamentals-of-deep-learning-introduction-to-lstm/\n\nWhat are LSTM's?\n* https://www.coursera.org/learn/nlp-sequence-models/lecture/KXoay/long-short-term-memory-lstm\n* https://distill.pub/2019/memorization-in-rnns/\n* https://towardsdatascience.com/illustrated-guide-to-lstms-and-gru-s-a-step-by-step-explanation-44e9eb85bf21\n\n# Code Implementation\n\nWe have already tokenized and paded our text for input to LSTM's","metadata":{}},{"cell_type":"code","source":"# create an embedding matrix for the words we have in the dataset\nembedding_matrix = np.zeros((len(word_index) + 1, 300))\nfor word, i in tqdm(word_index.items()):\n    embedding_vector = embeddings_index.get(word)\n    if embedding_vector is not None:\n        embedding_matrix[i] = embedding_vector","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-06T06:45:34.637050Z","iopub.execute_input":"2025-06-06T06:45:34.637436Z","iopub.status.idle":"2025-06-06T06:45:34.805306Z","shell.execute_reply.started":"2025-06-06T06:45:34.637405Z","shell.execute_reply":"2025-06-06T06:45:34.804475Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"%%time\nwith strategy.scope():\n    \n    # A simple LSTM with glove embeddings and one dense layer\n    model = Sequential()\n    model.add(Embedding(len(word_index) + 1,\n                     300,\n                     weights=[embedding_matrix],\n                     input_length=max_len,\n                     trainable=False))\n\n    model.add(LSTM(100, dropout=0.3, recurrent_dropout=0.3))\n    model.add(Dense(1, activation='sigmoid'))\n    model.compile(loss='binary_crossentropy', optimizer='adam',metrics=['accuracy'])\n    \nmodel.summary()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-06T06:45:38.532946Z","iopub.execute_input":"2025-06-06T06:45:38.533286Z","iopub.status.idle":"2025-06-06T06:45:38.761600Z","shell.execute_reply.started":"2025-06-06T06:45:38.533260Z","shell.execute_reply":"2025-06-06T06:45:38.760657Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"model.fit(xtrain_pad, ytrain, epochs=5, batch_size=64*strategy.num_replicas_in_sync)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-06-06T06:45:55.596594Z","iopub.execute_input":"2025-06-06T06:45:55.597481Z","iopub.status.idle":"2025-06-06T07:20:16.061136Z","shell.execute_reply.started":"2025-06-06T06:45:55.597444Z","shell.execute_reply":"2025-06-06T07:20:16.060305Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"scores = model.predict(xvalid_pad)\npredicted_labels = (scores > 0.5).astype(int)  # 转为0/1标签\naccuracy = accuracy_score(yvalid, predicted_labels)\nprint(f\"Accuracy: {accuracy:.4f}\")  # 输出格式化为4位小数","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"scores_model.append({'Model': 'LSTM','AUC_Score': accuracy})","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Code Explanation\n\nAs a first step we calculate embedding matrix for our vocabulary from the pretrained GLoVe vectors . Then while building the embedding layer we pass Embedding Matrix as weights to the layer instead of training it over Vocabulary and thus we pass trainable = False.\nRest of the model is same as before except we have replaced the SimpleRNN By LSTM Units\n\n* Comments on the Model\n\nWe now see that the model is not overfitting and achieves an auc score of 0.96 which is quite commendable , also we close in on the gap between accuracy and auc .\nWe see that in this case we used dropout and prevented overfitting the data","metadata":{}},{"cell_type":"markdown","source":"# GRU's\n\n## Basic  Overview\n\nIntroduced by Cho, et al. in 2014, GRU (Gated Recurrent Unit) aims to solve the vanishing gradient problem which comes with a standard recurrent neural network. GRU's are a variation on the LSTM because both are designed similarly and, in some cases, produce equally excellent results . GRU's were designed to be simpler and faster than LSTM's and in most cases produce equally good results and thus there is no clear winner.\n\n## In Depth Explanation\n\n* https://towardsdatascience.com/understanding-gru-networks-2ef37df6c9be\n* https://www.coursera.org/learn/nlp-sequence-models/lecture/agZiL/gated-recurrent-unit-gru\n* https://www.geeksforgeeks.org/gated-recurrent-unit-networks/\n\n## Code Implementation","metadata":{}},{"cell_type":"code","source":"%%time\nwith strategy.scope():\n    # GRU with glove embeddings and two dense layers\n     model = Sequential()\n     model.add(Embedding(len(word_index) + 1,\n                     300,\n                     weights=[embedding_matrix],\n                     input_length=max_len,\n                     trainable=False))\n     model.add(SpatialDropout1D(0.3))\n     model.add(GRU(300))\n     model.add(Dense(1, activation='sigmoid'))\n\n     model.compile(loss='binary_crossentropy', optimizer='adam',metrics=['accuracy'])   \n    \nmodel.summary()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"model.fit(xtrain_pad, ytrain, epochs=5, batch_size=64*strategy.num_replicas_in_sync)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"scores = model.predict(xvalid_pad)\npredicted_labels = (scores > 0.5).astype(int)  # 转为0/1标签\naccuracy = accuracy_score(yvalid, predicted_labels)\nprint(f\"Accuracy: {accuracy:.4f}\")  # 输出格式化为4位小数","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"scores_model.append({'Model': 'GRU','AUC_Score': accuracy})","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Bi-Directional RNN's\n\n## In Depth Explanation\n\n* https://www.coursera.org/learn/nlp-sequence-models/lecture/fyXnn/bidirectional-rnn\n* https://towardsdatascience.com/understanding-bidirectional-rnn-in-pytorch-5bd25a5dd66\n* https://d2l.ai/chapter_recurrent-modern/bi-rnn.html\n\n## Code Implementation","metadata":{}},{"cell_type":"code","source":"%%time\nwith strategy.scope():\n    # A simple bidirectional LSTM with glove embeddings and one dense layer\n    model = Sequential()\n    model.add(Embedding(len(word_index) + 1,\n                     300,\n                     weights=[embedding_matrix],\n                     input_length=max_len,\n                     trainable=False))\n    model.add(Bidirectional(LSTM(300, dropout=0.3, recurrent_dropout=0.3)))\n\n    model.add(Dense(1,activation='sigmoid'))\n    model.compile(loss='binary_crossentropy', optimizer='adam',metrics=['accuracy'])\n    \n    \nmodel.summary()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"model.fit(xtrain_pad, ytrain, nb_epoch=5, batch_size=64*strategy.num_replicas_in_sync)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"scores = model.predict(xvalid_pad)\npredicted_labels = (scores > 0.5).astype(int)  # 转为0/1标签\naccuracy = accuracy_score(yvalid, predicted_labels)\nprint(f\"Accuracy: {accuracy:.4f}\")  # 输出格式化为4位小数","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"scores_model.append({'Model': 'Bi-directional LSTM','AUC_Score': accuracy})","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"scores_model","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Code Explanation\n\nCode is same as before,only we have added bidirectional nature to the LSTM cells we used before and is self explanatory. We have achieve similar accuracy and auc score as before and now we have learned all the types of typical RNN architectures","metadata":{}},{"cell_type":"markdown","source":"**We are now at the end of part 1 of this notebook and things are about to go wild now as we Enter more complex and State of the art models .If you have followed along from the starting and read all the articles and understood everything , these complex models would be fairly easy to understand.I recommend Finishing Part 1 before continuing as the upcoming techniques can be quite overwhelming**","metadata":{}},{"cell_type":"markdown","source":"# Seq2Seq Model Architecture\n\n## Overview\n\nRNN's are of many types  and different architectures are used for different purposes. Here is a nice video explanining different types of model architectures : https://www.coursera.org/learn/nlp-sequence-models/lecture/BO8PS/different-types-of-rnns.\nSeq2Seq is a many to many RNN architecture where the input is a sequence and the output is also a sequence (where input and output sequences can be or cannot be of different lengths). This architecture is used in a lot of applications like Machine Translation, text summarization, question answering etc\n\n## In Depth Understanding\n\nI will not write the code implementation for this,but rather I will provide the resources where code has already been implemented and explained in a much better way than I could have ever explained.\n\n* https://www.coursera.org/learn/nlp-sequence-models/lecture/HyEui/basic-models ---> A basic idea of different Seq2Seq Models\n\n* https://blog.keras.io/a-ten-minute-introduction-to-sequence-to-sequence-learning-in-keras.html , https://machinelearningmastery.com/define-encoder-decoder-sequence-sequence-model-neural-machine-translation-keras/ ---> Basic Encoder-Decoder Model and its explanation respectively\n\n* https://towardsdatascience.com/how-to-implement-seq2seq-lstm-model-in-keras-shortcutnlp-6f355f3e5639 ---> A More advanced Seq2seq Model and its explanation\n\n* https://d2l.ai/chapter_recurrent-modern/machine-translation-and-dataset.html , https://d2l.ai/chapter_recurrent-modern/encoder-decoder.html ---> Implementation of Encoder-Decoder Model from scratch\n\n* https://www.youtube.com/watch?v=IfsjMg4fLWQ&list=PLtmWHNX-gukKocXQOkQjuVxglSDYWsSh9&index=8&t=0s ---> Introduction to Seq2seq By fast.ai","metadata":{}},{"cell_type":"code","source":"# Visualization of Results obtained from various Deep learning models\nresults = pd.DataFrame(scores_model).sort_values(by='AUC_Score',ascending=False)\nresults.style.background_gradient(cmap='Blues')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"fig = go.Figure(go.Funnelarea(\n    text =results.Model,\n    values = results.AUC_Score,\n    title = {\"position\": \"top center\", \"text\": \"Funnel-Chart of Sentiment Distribution\"}\n    ))\nfig.show()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Attention Models\n\nThis is the toughest and most tricky part. If you are able to understand the intiuition and working of attention block , understanding transformers and transformer based architectures like BERT will be a piece of cake. This is the part where I spent the most time on and I suggest you do the same . Please read and view the following resources in the order I am providing to ignore getting confused, also at the end of this try to write and draw an attention block in your own way :-\n\n* https://www.coursera.org/learn/nlp-sequence-models/lecture/RDXpX/attention-model-intuition --> Only watch this video and not the next one\n* https://towardsdatascience.com/sequence-2-sequence-model-with-attention-mechanism-9e9ca2a613a\n* https://towardsdatascience.com/attention-and-its-different-forms-7fc3674d14dc\n* https://distill.pub/2016/augmented-rnns/ \n\n## Code Implementation\n\n* https://www.analyticsvidhya.com/blog/2019/11/comprehensive-guide-attention-mechanism-deep-learning/ --> Basic Level\n* https://pytorch.org/tutorials/intermediate/seq2seq_translation_tutorial.html ---> Implementation from Scratch in Pytorch","metadata":{}},{"cell_type":"markdown","source":"# Transformers : Attention is all you need\n\nSo finally we have reached the end of the learning curve and are about to start learning the technology that changed NLP completely and are the reasons for the state of the art NLP techniques .Transformers were introduced in the paper Attention is all you need by Google. If you have understood the Attention models,this will be very easy , Here is transformers fully explained:\n\n* http://jalammar.github.io/illustrated-transformer/\n\n## Code Implementation\n\n* http://nlp.seas.harvard.edu/2018/04/03/attention.html ---> This presents the code implementation of the architecture presented in the paper by Google","metadata":{}},{"cell_type":"markdown","source":"# BERT and Its Implementation on this Competition\n\nAs Promised I am back with Resiurces , to understand about BERT architecture , please follow the contents in the given order :-\n\n* http://jalammar.github.io/illustrated-bert/ ---> In Depth Understanding of BERT\n\nAfter going through the post Above , I guess you must have understood how transformer architecture have been utilized by the current SOTA models . Now these architectures can be used in two ways :<br><br>\n1) We can use the model for prediction on our problems using the pretrained weights without fine-tuning or training the model for our sepcific tasks\n* EG: http://jalammar.github.io/a-visual-guide-to-using-bert-for-the-first-time/ ---> Using Pre-trained BERT without Tuning\n\n2) We can fine-tune or train these transformer models for our task by tweaking the already pre-trained weights and training on a much smaller dataset\n* EG:* https://www.youtube.com/watch?v=hinZO--TEk4&t=2933s ---> Tuning BERT For your TASK\n\nWe will be using the first example as a base for our implementation of BERT model using Hugging Face and KERAS , but contrary to first example we will also Fine-Tune our model for our task\n\nAcknowledgements : https://www.kaggle.com/xhlulu/jigsaw-tpu-distilbert-with-huggingface-and-keras\n\n\nSteps Involved :\n* Data Preparation : Tokenization and encoding of data\n* Configuring TPU's \n* Building a Function for Model Training and adding an output layer for classification\n* Train the model and get the results","metadata":{}},{"cell_type":"code","source":"# Loading Dependencies\nimport os\nimport tensorflow as tf\nfrom tensorflow.keras.layers import Dense, Input\nfrom tensorflow.keras.optimizers import Adam\nfrom tensorflow.keras.models import Model\nfrom tensorflow.keras.callbacks import ModelCheckpoint\nfrom kaggle_datasets import KaggleDatasets\nimport transformers\n\nfrom tokenizers import BertWordPieceTokenizer","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# LOADING THE DATA\n\ntrain1 = pd.read_csv(\"/kaggle/input/jigsaw-multilingual-toxic-comment-classification/jigsaw-toxic-comment-train.csv\")\nvalid = pd.read_csv('/kaggle/input/jigsaw-multilingual-toxic-comment-classification/validation.csv')\ntest = pd.read_csv('/kaggle/input/jigsaw-multilingual-toxic-comment-classification/test.csv')\nsub = pd.read_csv('/kaggle/input/jigsaw-multilingual-toxic-comment-classification/sample_submission.csv')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Encoder FOr DATA for understanding waht encode batch does read documentation of hugging face tokenizer :\nhttps://huggingface.co/transformers/main_classes/tokenizer.html here","metadata":{}},{"cell_type":"markdown","source":"## Tokenization\n\nFor understanding please refer to hugging face documentation again","metadata":{}},{"cell_type":"code","source":"# First load the real tokenizer\ntokenizer = transformers.DistilBertTokenizer.from_pretrained('distilbert-base-multilingual-cased')\n# Save the loaded tokenizer locally\ntokenizer.save_pretrained('.')\n# Reload it with the huggingface tokenizers library\nfast_tokenizer = BertWordPieceTokenizer('vocab.txt', lowercase=False)\nfast_tokenizer","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def fast_encode(texts, tokenizer, chunk_size=256, maxlen=512):\n    \"\"\"\n    Encoder for encoding the text into sequence of integers for BERT Input\n    将输入的文本列表（texts）通过指定的 tokenizer 批量编码为固定长度的整数序列（Token IDs），\n    并支持分块处理以优化内存使用。\n    \"\"\"\n    tokenizer.enable_truncation(max_length=maxlen)\n    tokenizer.enable_padding(max_length=maxlen)\n    all_ids = []\n    \n    for i in tqdm(range(0, len(texts), chunk_size)):\n        text_chunk = texts[i:i+chunk_size].tolist()\n        encs = tokenizer.encode_batch(text_chunk)\n        all_ids.extend([enc.ids for enc in encs])\n    \n    return np.array(all_ids)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"x_train = fast_encode(train1.comment_text.astype(str), fast_tokenizer, maxlen=MAX_LEN)\nx_valid = fast_encode(valid.comment_text.astype(str), fast_tokenizer, maxlen=MAX_LEN)\nx_test = fast_encode(test.content.astype(str), fast_tokenizer, maxlen=MAX_LEN)\n\ny_train = train1.toxic.values\ny_valid = valid.toxic.values","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#IMP DATA FOR CONFIG\n\nAUTO = tf.data.experimental.AUTOTUNE\n\n\n# Configuration\nEPOCHS = 3\nBATCH_SIZE = 16 * strategy.num_replicas_in_sync\nMAX_LEN = 192","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_dataset = (\n    tf.data.Dataset\n    .from_tensor_slices((x_train, y_train))\n    .repeat()\n    .shuffle(2048)\n    .batch(BATCH_SIZE)\n    .prefetch(AUTO)\n)\n\nvalid_dataset = (\n    tf.data.Dataset\n    .from_tensor_slices((x_valid, y_valid))\n    .batch(BATCH_SIZE)\n    .cache()\n    .prefetch(AUTO)\n)\n\ntest_dataset = (\n    tf.data.Dataset\n    .from_tensor_slices(x_test)\n    .batch(BATCH_SIZE)\n)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def build_model(transformer, max_len=512):\n    \"\"\"\n    function for training the BERT model\n    \"\"\"\n    input_word_ids = Input(shape=(max_len,), dtype=tf.int32, name=\"input_word_ids\") \n    ###输入 (input_word_ids)为 (batch_size, max_len) 的整数张量，表示批量文本的Token ID序列。\n    #### 假设 max_len=5, batch_size=2\n    ###input_word_ids = [\n    ### [101, 7592, 2088, 102, 0],   # 样本1: [CLS], \"Hello\", \"world\", [SEP], [PAD]\n    ###[101, 2023, 2003, 102, 0]     # 样本2: [CLS], \"This\", \"is\", [SEP], [PAD]\n    ###]\n    sequence_output = transformer(input_word_ids)[0] ###对于大多数HuggingFace的Transformer模型（如BertModel、TFDistilBertModel），\n    ###输出是一个元组：outputs[0]: 所有Token的上下文向量（形状 (batch_size, max_len, hidden_dim)）\n\n    cls_token = sequence_output[:, 0, :] ### # 取每个样本的第0个Token（[CLS]）\n    ###在BERT预训练中，[CLS] Token的向量被设计为聚合整个序列的信息。分类任务中，通常仅用 [CLS] 向量作为分类头的输入。\n    out = Dense(1, activation='sigmoid')(cls_token)\n    \n    model = Model(inputs=input_word_ids, outputs=out)\n    model.compile(Adam(lr=1e-5), loss='binary_crossentropy', metrics=['accuracy'])\n    \n    return model","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Starting Training\n\nIf you want to use any another model just replace the model name in transformers._____ and use accordingly","metadata":{}},{"cell_type":"code","source":"%%time\nwith strategy.scope():\n    transformer_layer = (\n        transformers.TFDistilBertModel\n        .from_pretrained('distilbert-base-multilingual-cased')\n    )\n    model = build_model(transformer_layer, max_len=MAX_LEN)\nmodel.summary()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"n_steps = x_train.shape[0] // BATCH_SIZE\ntrain_history = model.fit(\n    train_dataset,\n    steps_per_epoch=n_steps,\n    validation_data=valid_dataset,\n    epochs=EPOCHS\n)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"n_steps = x_valid.shape[0] // BATCH_SIZE\ntrain_history_2 = model.fit(\n    valid_dataset.repeat(),\n    steps_per_epoch=n_steps,\n    epochs=EPOCHS*2\n)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"sub['toxic'] = model.predict(test_dataset, verbose=1)\nsub.to_csv('submission.csv', index=False)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# End Notes\n\nThis was my effort to share my learnings so that everyone can benifit from it.As this community has been very kind to me and helped me in learning all of this , I want to take this forward. I have shared all the resources I used to learn all the stuff .Join me and make these NLP competitions your first ,without being overwhelmed by the shear number of techniques used . It took me 10 days to learn all of this , you can learn it at your pace and dont give in , at the end of all this you will be a different person and it will all be worth it.\n\n\n### I am attaching more resources if you want NLP end to end:\n\n1) Books\n\n* https://d2l.ai/\n* Jason Brownlee's Books\n\n2) Courses\n\n* https://www.coursera.org/learn/nlp-sequence-models/home/welcome\n* Fast.ai NLP Course\n\n3) Blogs and websites\n\n* Machine Learning Mastery\n* https://distill.pub/\n* http://jalammar.github.io/\n\n**<span style=\"color:Red\">This is subtle effort of contributing towards the community, if it helped you in any way please show a token of love by upvoting**","metadata":{}}]}