{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.7.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceId":19018,"databundleVersionId":2703900,"sourceType":"competition"},{"sourceId":11650,"sourceType":"datasetVersion","datasetId":8327}],"dockerImageVersionId":30299,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"목표  \n:다국어 댓글들 중에서 부정 댓글 식별하기  \n  \n데이터셋 [링크](https://www.kaggle.com/code/tanulsingh077/deep-learning-for-nlp-zero-to-transformers-bert)\n- jigsaw-toxic-comment-train.csv: 위키피디아의 talk 페이지에서 추출된 댓글 데이터 (eng)   \n- test.csv와 validation.csv는 다국어 댓글 포함  \n- comment_text: 댓글  \n- toxic: 악성댓글(1) 일반댓글(0)  \n  \n\n\n","metadata":{}},{"cell_type":"markdown","source":"# Contents\n\nIn this Notebook I will start with the very Basics of RNN's and Build all the way to latest deep learning architectures to solve NLP problems. It will cover the Following:\n* Simple RNN's\n* Word Embeddings : Definition and How to get them\n* LSTM's\n* GRU's\n* BI-Directional RNN's\n* Encoder-Decoder Models (Seq2Seq Models)\n* Attention Models\n* Transformers - Attention is all you need\n* BERT\n\nI will divide every Topic into four subsections:\n* Basic Overview\n* In-Depth Understanding : In this I will attach links of articles and videos to learn about the topic in depth\n* Code-Implementation\n* Code Explanation\n\nThis is a comprehensive kernel and if you follow along till the end , I promise you would learn all the techniques completely\n\nNote that the aim of this notebook is not to have a High LB score but to present a beginner guide to understand Deep Learning techniques used for NLP. Also after discussing all of these ideas , I will present a starter solution for this competiton","metadata":{}},{"cell_type":"code","source":"# Shapely 와 PyGEOS 버전 맞춰주기\n!pip install --upgrade shapely pygeos","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-31T09:29:27.619019Z","iopub.execute_input":"2025-03-31T09:29:27.619626Z","iopub.status.idle":"2025-03-31T09:29:38.088704Z","shell.execute_reply.started":"2025-03-31T09:29:27.619596Z","shell.execute_reply":"2025-03-31T09:29:38.087732Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nfrom tqdm import tqdm\nfrom sklearn.model_selection import train_test_split\nfrom tensorflow import keras\nimport tensorflow as tf\nfrom sklearn import preprocessing, decomposition, model_selection, metrics, pipeline\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n%matplotlib inline\nfrom plotly import graph_objs as go\nimport plotly.express as px\nimport plotly.figure_factory as ff","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2025-03-31T09:29:45.825462Z","iopub.execute_input":"2025-03-31T09:29:45.826565Z","iopub.status.idle":"2025-03-31T09:29:45.833832Z","shell.execute_reply.started":"2025-03-31T09:29:45.826525Z","shell.execute_reply":"2025-03-31T09:29:45.832917Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(tf.__version__)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-31T09:29:47.612942Z","iopub.execute_input":"2025-03-31T09:29:47.613310Z","iopub.status.idle":"2025-03-31T09:29:47.618291Z","shell.execute_reply.started":"2025-03-31T09:29:47.613280Z","shell.execute_reply":"2025-03-31T09:29:47.617393Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Configuring TPU's\n\nFor this version of Notebook we will be using TPU's as we have to built a BERT Model","metadata":{}},{"cell_type":"code","source":"# Detect hardware, return appropriate distribution strategy\ntry:\n    # TPU detection. No parameters necessary if TPU_NAME environment variable is\n    # set: this is always the case on Kaggle.\n    tpu = tf.distribute.cluster_resolver.TPUClusterResolver()\n    print('Running on TPU ', tpu.master())\nexcept ValueError:\n    tpu = None\n\nif tpu:\n    tf.config.experimental_connect_to_cluster(tpu)\n    tf.tpu.experimental.initialize_tpu_system(tpu)\n    strategy = tf.distribute.experimental.TPUStrategy(tpu)\nelse:\n    # Default distribution strategy in Tensorflow. Works on CPU and single GPU.\n    strategy = tf.distribute.get_strategy()\n\nprint(\"REPLICAS(device): \", strategy.num_replicas_in_sync)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-31T09:29:49.575528Z","iopub.execute_input":"2025-03-31T09:29:49.575885Z","iopub.status.idle":"2025-03-31T09:29:49.588999Z","shell.execute_reply.started":"2025-03-31T09:29:49.575848Z","shell.execute_reply":"2025-03-31T09:29:49.588083Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train = pd.read_csv('/kaggle/input/jigsaw-multilingual-toxic-comment-classification/jigsaw-toxic-comment-train.csv')\nvalidation = pd.read_csv('/kaggle/input/jigsaw-multilingual-toxic-comment-classification/validation.csv')\ntest = pd.read_csv('/kaggle/input/jigsaw-multilingual-toxic-comment-classification/test.csv')","metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true,"execution":{"iopub.status.busy":"2025-03-31T09:29:51.452041Z","iopub.execute_input":"2025-03-31T09:29:51.452923Z","iopub.status.idle":"2025-03-31T09:29:54.807819Z","shell.execute_reply.started":"2025-03-31T09:29:51.452886Z","shell.execute_reply":"2025-03-31T09:29:54.807061Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 기본 정보 확인\nprint(f\"훈련데이터\\n\")\nprint(train.info())\nprint('---'*20)\nprint(f\"검증 데이터\\n\")\nprint(validation.info())\nprint('---'*20)\nprint(f\"테스트 데이터\\n\")\nprint(test.info())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-31T09:29:55.858416Z","iopub.execute_input":"2025-03-31T09:29:55.858802Z","iopub.status.idle":"2025-03-31T09:29:55.943471Z","shell.execute_reply.started":"2025-03-31T09:29:55.858770Z","shell.execute_reply":"2025-03-31T09:29:55.942644Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 문장 길이 계산\ntrain['comment_length'] = train['comment_text'].apply(len)\nprint(f\"문장 길이:\\n{train['comment_length'].describe()}\")\n# 결측값 확인\nprint()\nprint(f\"결측값:\\n{train.isnull().sum()}\")\nprint()\n# 'toxic' 또는 목표 변수 분포 확인\nprint(f\"target 분포:\\n{train['toxic'].value_counts()}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-31T09:29:58.477867Z","iopub.execute_input":"2025-03-31T09:29:58.478221Z","iopub.status.idle":"2025-03-31T09:29:58.612801Z","shell.execute_reply.started":"2025-03-31T09:29:58.478191Z","shell.execute_reply":"2025-03-31T09:29:58.611747Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"평균 문장 길이가 391자로 상당히 긴 문장의 댓글이 많다.  \n특히 최소길이 1자, 최대길이 5000자의 댓글 등은 이상치 처리를 하거나 길이를 자르는 등의 특별한 접근이 필요해 보인다.  \n그 외 대부분은 중간 길이 범위에 집중되어있음 (93 ~ 431자)  \n표준편차도 592.86로 큰 편으로 보인다, 문장 길이에 변동성이 큰 것을 알 수 있다.  \n\n악성댓글수(1)이 일반댓글수(0) 보다 분포가 적어서 클래스 불균형이 있음","metadata":{}},{"cell_type":"code","source":"# 컬럼 이름 확인\nprint(f\"column names:\\n{train.columns}\")\nprint('---'*40)\n# 데이터의 첫 5줄 확인\nprint(f\"train head:\\n{train.head(30)}\")\nprint(f\"vallidation head:\\n{validation.head()}\")\nprint(f\"test head:\\n{test.head()}\")\nprint('---'*40)\n# 수치형 데이터의 기술 통계 확인\nprint(f\"describe:\\n{train.describe()}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-31T09:30:00.815202Z","iopub.execute_input":"2025-03-31T09:30:00.816060Z","iopub.status.idle":"2025-03-31T09:30:00.887935Z","shell.execute_reply.started":"2025-03-31T09:30:00.816011Z","shell.execute_reply":"2025-03-31T09:30:00.886889Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"We will drop the other columns and approach this problem as a Binary Classification Problem and also we will have our exercise done on a smaller subsection of the dataset(only 12000 data points) to make it easier to train the models","metadata":{}},{"cell_type":"code","source":"# 불필요 컬럼 삭제: 이진 레이블 예측 문제이므로 toxic에 집중하기 위해 유해 댓글의 세부 유형을 나타내는 컬럼 삭제 \ntrain.drop(['severe_toxic','obscene','threat','insult','identity_hate'],axis=1,inplace=True)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-31T09:30:04.382336Z","iopub.execute_input":"2025-03-31T09:30:04.382707Z","iopub.status.idle":"2025-03-31T09:30:04.399622Z","shell.execute_reply.started":"2025-03-31T09:30:04.382675Z","shell.execute_reply":"2025-03-31T09:30:04.398810Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 슬라이싱\ntrain = train.loc[:12000,:]\ntrain.shape","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-31T09:30:07.870761Z","iopub.execute_input":"2025-03-31T09:30:07.871616Z","iopub.status.idle":"2025-03-31T09:30:07.878961Z","shell.execute_reply.started":"2025-03-31T09:30:07.871584Z","shell.execute_reply":"2025-03-31T09:30:07.878038Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"We will check the maximum number of words that can be present in a comment , this will help us in padding later","metadata":{}},{"cell_type":"code","source":"# 각 댓글의 단어 수 계산 후 단어 수가 가장 많은 댓글 찾기\ntrain['comment_text'].apply(lambda x:len(str(x).split())).max()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-31T09:30:09.860534Z","iopub.execute_input":"2025-03-31T09:30:09.861408Z","iopub.status.idle":"2025-03-31T09:30:09.920259Z","shell.execute_reply.started":"2025-03-31T09:30:09.861373Z","shell.execute_reply":"2025-03-31T09:30:09.919378Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Writing a function for getting auc score for validation\n이진 분류 모델의 성능 평가 지표 함수 준비\n\nmetrics.roc_curve()로 ROC 곡선 계산 \n","metadata":{}},{"cell_type":"code","source":"def roc_auc(predictions,target):\n    '''\n    This methods returns the AUC Score when given the Predictions\n    and Labels\n\n    Returns\n    -------\n    roc_auc:\n        모델의 분류 성능을 평가하는 AUC 점수 반환\n    '''\n    # thresholds: 예측 확률에 대한 임계값. 이를 기준으로 예측이 양성 또는 음성으로 분류된다.\n    fpr, tpr, thresholds = metrics.roc_curve(target, predictions)\n    roc_auc = metrics.auc(fpr, tpr)\n    return roc_auc","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-31T10:15:17.949108Z","iopub.execute_input":"2025-03-31T10:15:17.949431Z","iopub.status.idle":"2025-03-31T10:15:17.954845Z","shell.execute_reply.started":"2025-03-31T10:15:17.949406Z","shell.execute_reply":"2025-03-31T10:15:17.953728Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Data Preparation","metadata":{}},{"cell_type":"code","source":"xtrain, xvalid, ytrain, yvalid = train_test_split(train.comment_text.values, train.toxic.values, #.values가 넘파이배열 형태로 반환\n                                                  stratify=train.toxic.values, # 훈련/검증 데이터 간 클래스 비율 유지\n                                                  random_state=42, \n                                                  test_size=0.2, shuffle=True)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-31T09:30:17.236713Z","iopub.execute_input":"2025-03-31T09:30:17.237632Z","iopub.status.idle":"2025-03-31T09:30:17.249493Z","shell.execute_reply.started":"2025-03-31T09:30:17.237592Z","shell.execute_reply":"2025-03-31T09:30:17.248584Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Before We Begin\n\nBefore we Begin If you are a complete starter with NLP and never worked with text data, I am attaching a few kernels that will serve as a starting point of your journey\n* https://www.kaggle.com/arthurtok/spooky-nlp-and-topic-modelling-tutorial\n* https://www.kaggle.com/abhishek/approaching-almost-any-nlp-problem-on-kaggle\n\nIf you want a more basic dataset to practice with here is another kernel which I wrote:\n* https://www.kaggle.com/tanulsingh077/what-s-cooking\n\nBelow are some Resources to get started with basic level Neural Networks, It will help us to easily understand the upcoming parts\n* https://www.youtube.com/watch?v=aircAruvnKk&list=PL_h2yd2CGtBHEKwEH5iqTZH85wLS-eUzv\n* https://www.youtube.com/watch?v=IHZwWFHWa-w&list=PL_h2yd2CGtBHEKwEH5iqTZH85wLS-eUzv&index=2\n* https://www.youtube.com/watch?v=Ilg3gGewQ5U&list=PL_h2yd2CGtBHEKwEH5iqTZH85wLS-eUzv&index=3\n* https://www.youtube.com/watch?v=tIeHLnjs5U8&list=PL_h2yd2CGtBHEKwEH5iqTZH85wLS-eUzv&index=4\n\nFor Learning how to visualize test data and what to use view:\n* https://www.kaggle.com/tanulsingh077/twitter-sentiment-extaction-analysis-eda-and-model\n* https://www.kaggle.com/jagangupta/stop-the-s-toxic-comments-eda","metadata":{}},{"cell_type":"markdown","source":"# Simple RNN\n\n## Basic Overview\n\nWhat is a RNN?\n\nRecurrent Neural Network(RNN) are a type of Neural Network where the output from previous step are fed as input to the current step. In traditional neural networks, all the inputs and outputs are independent of each other, but in cases like when it is required to predict the next word of a sentence, the previous words are required and hence there is a need to remember the previous words. Thus RNN came into existence, which solved this issue with the help of a Hidden Layer.\n\nWhy RNN's?\n\nhttps://www.quora.com/Why-do-we-use-an-RNN-instead-of-a-simple-neural-network\n\n## In-Depth Understanding\n\n* https://medium.com/mindorks/understanding-the-recurrent-neural-network-44d593f112a2\n* https://www.youtube.com/watch?v=2E65LDnM2cA&list=PL1F3ABbhcqa3BBWo170U4Ev2wfsF7FN8l\n* https://www.d2l.ai/chapter_recurrent-neural-networks/rnn.html\n\n## Code Implementation\n\nSo first I will implement the and then I will explain the code step by step","metadata":{}},{"cell_type":"code","source":"# 각 문장의 단어 수(토큰 개수) 계산\ncomment_lengths = [len(comment.split()) for comment in xtrain]\n\n# 히스토그램 시각화: max_length를 정하기 위해 문장 길이 분포 확인\nplt.figure(figsize=(10,5))\nplt.hist(comment_lengths, bins=50, edgecolor='black')\nplt.xlabel('comment length(number of words)')\nplt.ylabel('number of comments')\nplt.title('comment length distribution')\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-31T09:30:21.169382Z","iopub.execute_input":"2025-03-31T09:30:21.169778Z","iopub.status.idle":"2025-03-31T09:30:21.676897Z","shell.execute_reply.started":"2025-03-31T09:30:21.169750Z","shell.execute_reply":"2025-03-31T09:30:21.676076Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**토큰화(Tokenization)**\n\nRNN은 문장을 단어별로 입력합니다. 우리는 각 단어를 원-핫 벡터로 표현하는데, 이 벡터의 차원은 어휘에 있는 단어 수 + 1입니다.\n\nKeras의 Tokenizer는 다음과 같은 작업을 수행합니다:\n\n1. 말뭉치에서 모든 고유한 단어들을 추출하여, 단어들을 **키(key)**로 하고 그 단어의 등장 횟수를 **값(value)**으로 하는 사전을 만듭니다.\n\n2. 그런 다음 등장 횟수를 기준으로 사전을 내림차순으로 정렬합니다.\n\n3. 이 후, 첫 번째 단어에 1, 두 번째 단어에 2, 이렇게 차례대로 인덱스를 부여합니다. 예를 들어, 만약 'the'라는 단어가 말뭉치에서 가장 많이 등장했다면, 그 단어는 1번 인덱스를 할당받고, 'the'를 나타내는 벡터는 **첫 번째 위치에 1을, 나머지 위치는 모두 0**이 되는 원-핫 벡터로 표현됩니다.\n\nxtrain_seq의 첫 번째 두 요소를 출력해 보세요. 그러면 각 단어가 숫자로 표현된 것을 확인할 수 있을 것입니다.","metadata":{}},{"cell_type":"code","source":"# Keras Tokenizer 로 텍스트 데이터를 단어 인덱스로 변환 \n# !TensorFlow 2.x 부터 최신권장문법은 tensorflow.keras (라이브러리 변경해줘야함)\ntoken = keras.preprocessing.text.Tokenizer(num_words=None)\nmax_len = 1500 # 패딩할 시퀀스 길이의 기준 정의\n\n# xtrain과 xvalid 데이터에서 나오는 모든 단어들을 기반으로 만들어진 단어 사전 학습\ntoken.fit_on_texts(list(xtrain) + list(xvalid))\n# 각 문장의 각 단어를 정수 시퀀스(인덱스 리스트)로 변환\nxtrain_seq = token.texts_to_sequences(xtrain) \nxvalid_seq = token.texts_to_sequences(xvalid)\n\n# max_len에 맞게 정수 시퀀스에 제로 패딩 추가. 기본 시퀀스의 앞에서부터 0이 붙는다.\n# padding='pre' 또는 padding='post' 파라미터를 통해 시퀀스의 앞과 뒤 중 어느 위치에 패딩을 줄지 결정할 수 있다.\nxtrain_pad = keras.preprocessing.sequence.pad_sequences(xtrain_seq, maxlen=max_len)\nxvalid_pad = keras.preprocessing.sequence.pad_sequences(xvalid_seq, maxlen=max_len)\n\n# 모델의 인풋 데이터 형식으로 변환\n# word_index 는 Tokenizer 객체에서 각 단어를 고유한 정수 인덱스에 매핑하는 딕셔너리\nword_index = token.word_index","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-31T09:30:24.288080Z","iopub.execute_input":"2025-03-31T09:30:24.288421Z","iopub.status.idle":"2025-03-31T09:30:26.248599Z","shell.execute_reply.started":"2025-03-31T09:30:24.288392Z","shell.execute_reply":"2025-03-31T09:30:26.247579Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"xtrain_seq[:1]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-31T09:30:28.787133Z","iopub.execute_input":"2025-03-31T09:30:28.787968Z","iopub.status.idle":"2025-03-31T09:30:28.793672Z","shell.execute_reply.started":"2025-03-31T09:30:28.787933Z","shell.execute_reply":"2025-03-31T09:30:28.792752Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"word_index는 단어 -> 인덱스 형태의 딕셔너리를 반환한다. 이 정보를 활용해서 텍스트를 정수로 변환하고 모델에 입력할 수 있다. ","metadata":{}},{"cell_type":"code","source":"# 특정 단어의 인덱스 확인\nprint(f\"단어'love'의 인덱스:\\n{word_index['love']}\")\nprint(f\"word_index의 길이:\\n{len(word_index.items())}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-31T09:26:02.095598Z","iopub.execute_input":"2025-03-31T09:26:02.096015Z","iopub.status.idle":"2025-03-31T09:26:02.102019Z","shell.execute_reply.started":"2025-03-31T09:26:02.095979Z","shell.execute_reply":"2025-03-31T09:26:02.100803Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"RNN 모델을 정의하고 컴파일하기\n- %%time: 셀의 실행 시간을 측정해 출력하는 Jypyter 매직 명령어.  \n- strategy.scope(): TensorFlow의 분산 학습 전략(여러 GPU에서 학습시)을 사용할 때 모델 생성에 필요하다.\n- Keras의 Embedding()층은 정수 인덱스로 표현된 단어 시퀀스를 밀집 벡터로 변환한다.  \n    - 입력 파라미터: \n        - input_dim: 단어 사전 크기(word_index 크기 + 1)\n        - output_dim: 각 단어를 임베딩할 벡터 차원\n        - input_length: 입력 문장의 최대 길이\n    - 출력 차원:\n        - 각 단어를 N임베딩 차원 벡터로 매핑한 행렬이 출력된다. (단어를 의미 공간에서 벡터로 변환한 형태)  \n- word_index는 정수 1부터 인덱스가 시작하지만, Embedding 층에서는 시퀀스 데이터의 패딩을 위해서 인덱스 0의 공간을 확보한다. 그래서 단어의 인덱스 범위를 정할 때 +1을 추가한다.  \n- SimpleRNN(출력차원수): RNN층은 순차 데이터(텍스트, 시간 시퀀스 등)의 특징을 학습한다.\n    - 출력차원(RNN의 Hidden Units) 선택 기준:  \n        - 적당한 표현력을 가지면서 과적합을 방지하고, 연산 비용을 최소화하는 균형을 잡는 것이 좋다.  \n        - RNN의 연산량은 O(n x d x h)(n: 시퀀스 길이, d: 입력 차원, h: hidden size)로 증가한다. 따라서 너무 큰 차원은 계산이 느려질 수 있다. \n    - NLP의 일반적인 경험값:\n        - 간단한 모델: 50 ~ 100  \n        - 중간 복잡도: 128 ~ 256  \n        - 복잡한 모델(LSTM, GRU, Transformer): 256 ~ 1024     \n\n","metadata":{}},{"cell_type":"code","source":"%%time\nwith strategy.scope():\n    # A simpleRNN without any pretrained embeddings and one dense layer\n    model = keras.models.Sequential() # sequential 모델 객체 생성\n    model.add(keras.layers.Embedding(len(word_index) + 1,    # 단어 인덱스의 크기(범위)\n                     300,  # 임베딩 차원:각 단어를 300차원의 실수 벡터로 변환\n                     input_length=max_len)) # 입력 시퀀스 길이\n    model.add(keras.layers.SimpleRNN(100))   # RNN의 출력 차원으로, 100개의 뉴런 사용\n    model.add(keras.layers.Dense(1, activation='sigmoid'))\n    model.compile(loss='binary_crossentropy', optimizer='adam', metrics=['accuracy'])\n    \nmodel.summary()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-31T09:30:35.957213Z","iopub.execute_input":"2025-03-31T09:30:35.957673Z","iopub.status.idle":"2025-03-31T09:30:39.738062Z","shell.execute_reply.started":"2025-03-31T09:30:35.957638Z","shell.execute_reply":"2025-03-31T09:30:39.737084Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Embedding 층의 313,049,100 파라미터(가중치)가 단어 벡터를 학습하는 데 사용된다.   \nsimple_rnn 층의 40,100 파라미터(가중치)가 RNN이 문장 패턴을 학습하는 데 사용된다.","metadata":{}},{"cell_type":"code","source":"model.fit(xtrain_pad, ytrain, epochs=5, batch_size=64*strategy.num_replicas_in_sync) #Multiplying by Strategy to run on TPU's","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-31T09:30:45.302978Z","iopub.execute_input":"2025-03-31T09:30:45.303597Z","iopub.status.idle":"2025-03-31T09:43:46.795875Z","shell.execute_reply.started":"2025-03-31T09:30:45.303561Z","shell.execute_reply":"2025-03-31T09:43:46.794953Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from sklearn.metrics import roc_auc_score\n\nscores = model.predict(xvalid_pad)\nprint(\"Auc: %.2f%%\" % (roc_auc(scores,yvalid))) # AUC가 높을 수록 모델의 분류 성능이 좋다. ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-31T10:15:24.634189Z","iopub.execute_input":"2025-03-31T10:15:24.635071Z","iopub.status.idle":"2025-03-31T10:15:31.956548Z","shell.execute_reply.started":"2025-03-31T10:15:24.635037Z","shell.execute_reply":"2025-03-31T10:15:31.955498Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 여러 모델의 AUC 점수를 비교할 수 있도록 리스트에 저장\nscores_model = []\nscores_model.append({'Model': 'SimpleRNN','AUC_Score': roc_auc(scores,yvalid)})","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-31T10:15:36.954832Z","iopub.execute_input":"2025-03-31T10:15:36.955255Z","iopub.status.idle":"2025-03-31T10:15:36.964092Z","shell.execute_reply.started":"2025-03-31T10:15:36.955218Z","shell.execute_reply":"2025-03-31T10:15:36.962816Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<b>Now you might be wondering What is padding? Why its done</b><br><br>\n\nHere is the answer :\n* https://www.quora.com/Which-effect-does-sequence-padding-have-on-the-training-of-a-neural-network\n* https://machinelearningmastery.com/data-preparation-variable-length-input-sequences-sequence-prediction/\n* https://www.coursera.org/lecture/natural-language-processing-tensorflow/padding-2Cyzs\n\nAlso sometimes people might use special tokens while tokenizing like EOS(end of string) and BOS(Begining of string). Here is the reason why it's done\n* https://stackoverflow.com/questions/44579161/why-do-we-do-padding-in-nlp-tasks\n\n\nThe code token.word_index simply gives the dictionary of vocab that keras created for us","metadata":{}},{"cell_type":"markdown","source":"* Building the Neural Network\n\nTo understand the Dimensions of input and output given to RNN in keras her is a beautiful article : https://medium.com/@shivajbd/understanding-input-and-output-shape-in-lstm-keras-c501ee95c65e\n\nThe first line model.Sequential() tells keras that we will be building our network sequentially . Then we first add the Embedding layer.\nEmbedding layer is also a layer of neurons which takes in as input the nth dimensional one hot vector of every word and converts it into 300 dimensional vector , it gives us word embeddings similar to word2vec. We could have used word2vec but the embeddings layer learns during training to enhance the embeddings.\nNext we add an 100 LSTM units without any dropout or regularization\nAt last we add a single neuron with sigmoid function which takes output from 100 LSTM cells (Please note we have 100 LSTM cells not layers) to predict the results and then we compile the model using adam optimizer \n\n* Comments on the model<br><br>\nWe can see our model achieves an accuracy of 1 which is just insane , we are clearly overfitting I know , but this was the simplest model of all ,we can tune a lot of hyperparameters like RNN units, we can do batch normalization , dropouts etc to get better result. The point is we got an AUC score of 0.82 without much efforts and we know have learnt about RNN's .Deep learning is really revolutionary","metadata":{}},{"cell_type":"markdown","source":"# Word Embeddings\n\nWhile building our simple RNN models we talked about using word-embeddings , So what is word-embeddings and how do we get word-embeddings?\nHere is the answer :\n* https://www.coursera.org/learn/nlp-sequence-models/lecture/6Oq70/word-representation\n* https://machinelearningmastery.com/what-are-word-embeddings/\n<br> <br>\nThe latest approach to getting word Embeddings is using pretained GLoVe or using Fasttext. Without going into too much details, I would explain how to create sentence vectors and how can we use them to create a machine learning model on top of it and since I am a fan of GloVe vectors, word2vec and fasttext. In this Notebook, I'll be using the GloVe vectors. You can download the GloVe vectors from here http://www-nlp.stanford.edu/data/glove.840B.300d.zip or you can search for GloVe in datasets on Kaggle and add the file","metadata":{}},{"cell_type":"code","source":"# load the GloVe vectors in a dictionary:\n\nembeddings_index = {} # {단어(키) : 해당 단어의 300차원 벡터(numpy arrays)}\nf = open('/kaggle/input/glove840b300dtxt/glove.840B.300d.txt','r',encoding='utf-8')\nfor line in tqdm(f):\n    values = line.split(' ')\n    word = values[0] # 첫번째 값을 word에 저장\n    coefs = np.asarray([float(val) for val in values[1:]]) # 나머지 300개의 숫자를 실수로 변환 -> numpy array 생성\n    embeddings_index[word] = coefs # 딕셔너리에 단어와 벡터를 저장\nf.close() # 파일 닫기 (리소스 절약)\n\n# 임베딩된 단어 개수 출력\nprint('Found %s word vectors.' % len(embeddings_index))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-31T10:16:14.294316Z","iopub.execute_input":"2025-03-31T10:16:14.295027Z","iopub.status.idle":"2025-03-31T10:19:49.820893Z","shell.execute_reply.started":"2025-03-31T10:16:14.294991Z","shell.execute_reply":"2025-03-31T10:19:49.819916Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(embeddings_index.get(\"love\"))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-31T10:20:50.786576Z","iopub.execute_input":"2025-03-31T10:20:50.787461Z","iopub.status.idle":"2025-03-31T10:20:50.794874Z","shell.execute_reply.started":"2025-03-31T10:20:50.787410Z","shell.execute_reply":"2025-03-31T10:20:50.793640Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# LSTM's\n\n## Basic Overview\n\nSimple RNN's were certainly better than classical ML algorithms and gave state of the art results, but it failed to capture long term dependencies that is present in sentences . So in 1998-99 LSTM's were introduced to counter to these drawbacks.\n\n## In Depth Understanding\n\nWhy LSTM's?\n* https://www.coursera.org/learn/nlp-sequence-models/lecture/PKMRR/vanishing-gradients-with-rnns\n* https://www.analyticsvidhya.com/blog/2017/12/fundamentals-of-deep-learning-introduction-to-lstm/\n\nWhat are LSTM's?\n* https://www.coursera.org/learn/nlp-sequence-models/lecture/KXoay/long-short-term-memory-lstm\n* https://distill.pub/2019/memorization-in-rnns/\n* https://towardsdatascience.com/illustrated-guide-to-lstms-and-gru-s-a-step-by-step-explanation-44e9eb85bf21\n\n# Code Implementation\n\nWe have already tokenized and paded our text for input to LSTM's","metadata":{}},{"cell_type":"code","source":"# create an embedding matrix for the words we have in the dataset\nembedding_matrix = np.zeros((len(word_index) + 1, 300))\nfor word, i in tqdm(word_index.items()):\n    embedding_vector = embeddings_index.get(word)\n    if embedding_vector is not None:\n        embedding_matrix[i] = embedding_vector","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-31T10:21:02.914137Z","iopub.execute_input":"2025-03-31T10:21:02.914924Z","iopub.status.idle":"2025-03-31T10:21:03.077115Z","shell.execute_reply.started":"2025-03-31T10:21:02.914874Z","shell.execute_reply":"2025-03-31T10:21:03.076149Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"embedding_matrix","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-31T10:21:10.288288Z","iopub.execute_input":"2025-03-31T10:21:10.289037Z","iopub.status.idle":"2025-03-31T10:21:10.295352Z","shell.execute_reply.started":"2025-03-31T10:21:10.289003Z","shell.execute_reply":"2025-03-31T10:21:10.294299Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"\nGloVe 사전 학습된 임베딩을 활용한 LSTM 모델 정의\n\n1. 입력 단어를 GloVe 임베딩 벡터로 변환\n2. LSTM을 사용하여 문맥 학습\n3. 최종적으로 이진 분류(sigmoid) 수행\n\n- strategy.scope() 안에서 실행되므로 TPU/GPU 병렬 처리 지원\n- Embedding 레이어에 사전 학습된 GloVe 임베딩(embedding_matrix)을 적용\n- LSTM을 사용하여 시퀀스 데이터를 학습 후, 마지막으로 이진 분류(sigmoid)를 수행\n    -  LSTM(100) → 출력 뉴런 수 = 100개\n    -  dropout=0.3 → 입력 뉴런의 30%를 랜덤으로 제외 (과적합 방지)\n    -  recurrent_dropout=0.3 → LSTM의 내부 순환 상태에도 드롭아웃 적용","metadata":{}},{"cell_type":"code","source":"%%time\nwith strategy.scope():\n    \n    # A simple LSTM with glove embeddings and one dense layer\n    model = keras.models.Sequential()\n    model.add(keras.layers.Embedding(len(word_index) + 1,\n                     300,\n                     weights=[embedding_matrix],  # GloVe 임베딩 로드\n                     input_length=max_len,\n                     trainable=False))            # 파인튜닝 미사용\n\n    model.add(keras.layers.LSTM(100, dropout=0.3, recurrent_dropout=0.3)) \n    model.add(keras.layers.Dense(1, activation='sigmoid'))\n    model.compile(loss='binary_crossentropy', optimizer='adam',metrics=['accuracy'])\n    \nmodel.summary()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-31T10:22:17.947034Z","iopub.execute_input":"2025-03-31T10:22:17.947663Z","iopub.status.idle":"2025-03-31T10:22:18.160548Z","shell.execute_reply.started":"2025-03-31T10:22:17.947631Z","shell.execute_reply":"2025-03-31T10:22:18.159558Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"model.fit(xtrain_pad, ytrain, epochs=5, batch_size=64*strategy.num_replicas_in_sync)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-31T10:22:22.003686Z","iopub.execute_input":"2025-03-31T10:22:22.004032Z","iopub.status.idle":"2025-03-31T10:50:16.948739Z","shell.execute_reply.started":"2025-03-31T10:22:22.004004Z","shell.execute_reply":"2025-03-31T10:50:16.947911Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"scores = model.predict(xvalid_pad)\nprint(\"Auc: %.2f%%\" % (roc_auc(scores,yvalid)))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-31T11:04:31.373709Z","iopub.execute_input":"2025-03-31T11:04:31.374087Z","iopub.status.idle":"2025-03-31T11:04:57.229839Z","shell.execute_reply.started":"2025-03-31T11:04:31.374058Z","shell.execute_reply":"2025-03-31T11:04:57.228884Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"scores_model.append({'Model': 'LSTM','AUC_Score': roc_auc(scores,yvalid)})","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-31T11:05:04.645960Z","iopub.execute_input":"2025-03-31T11:05:04.646669Z","iopub.status.idle":"2025-03-31T11:05:04.652339Z","shell.execute_reply.started":"2025-03-31T11:05:04.646632Z","shell.execute_reply":"2025-03-31T11:05:04.651499Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Code Explanation\n\nAs a first step we calculate embedding matrix for our vocabulary from the pretrained GLoVe vectors . Then while building the embedding layer we pass Embedding Matrix as weights to the layer instead of training it over Vocabulary and thus we pass trainable = False.\nRest of the model is same as before except we have replaced the SimpleRNN By LSTM Units\n\n* Comments on the Model\n\nWe now see that the model is not overfitting and achieves an auc score of 0.96 which is quite commendable , also we close in on the gap between accuracy and auc .\nWe see that in this case we used dropout and prevented overfitting the data","metadata":{}},{"cell_type":"markdown","source":"# GRU's\n\n## Basic  Overview\n\nIntroduced by Cho, et al. in 2014, GRU (Gated Recurrent Unit) aims to solve the vanishing gradient problem which comes with a standard recurrent neural network. GRU's are a variation on the LSTM because both are designed similarly and, in some cases, produce equally excellent results . GRU's were designed to be simpler and faster than LSTM's and in most cases produce equally good results and thus there is no clear winner.\n\n## In Depth Explanation\n\n* https://towardsdatascience.com/understanding-gru-networks-2ef37df6c9be\n* https://www.coursera.org/learn/nlp-sequence-models/lecture/agZiL/gated-recurrent-unit-gru\n* https://www.geeksforgeeks.org/gated-recurrent-unit-networks/\n\n## Code Implementation","metadata":{}},{"cell_type":"code","source":"%%time\nwith strategy.scope():\n    # GRU with glove embeddings and two dense layers\n     model = keras.models.Sequential()\n     model.add(keras.layers.Embedding(len(word_index) + 1,\n                     300,\n                     weights=[embedding_matrix],\n                     input_length=max_len,\n                     trainable=False))\n     model.add(keras.layers.SpatialDropout1D(0.3))\n     model.add(keras.layers.GRU(300))\n     model.add(keras.layers.Dense(1, activation='sigmoid'))\n\n     model.compile(loss='binary_crossentropy', optimizer='adam',metrics=['accuracy'])   \n    \nmodel.summary()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-31T11:06:27.390357Z","iopub.execute_input":"2025-03-31T11:06:27.391124Z","iopub.status.idle":"2025-03-31T11:06:27.707716Z","shell.execute_reply.started":"2025-03-31T11:06:27.391087Z","shell.execute_reply":"2025-03-31T11:06:27.706457Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"model.fit(xtrain_pad, ytrain, epochs=5, batch_size=64*strategy.num_replicas_in_sync)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-31T11:06:31.795256Z","iopub.execute_input":"2025-03-31T11:06:31.796180Z","iopub.status.idle":"2025-03-31T11:08:51.439133Z","shell.execute_reply.started":"2025-03-31T11:06:31.796143Z","shell.execute_reply":"2025-03-31T11:08:51.438295Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"scores = model.predict(xvalid_pad)\nprint(\"Auc: %.2f%%\" % (roc_auc(scores,yvalid)))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-31T11:11:11.831762Z","iopub.execute_input":"2025-03-31T11:11:11.832125Z","iopub.status.idle":"2025-03-31T11:11:14.808031Z","shell.execute_reply.started":"2025-03-31T11:11:11.832093Z","shell.execute_reply":"2025-03-31T11:11:14.807093Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 점수 비교를 위해 저장\nscores_model.append({'Model': 'GRU','AUC_Score': roc_auc(scores,yvalid)})","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-31T11:11:24.013179Z","iopub.execute_input":"2025-03-31T11:11:24.014004Z","iopub.status.idle":"2025-03-31T11:11:24.019396Z","shell.execute_reply.started":"2025-03-31T11:11:24.013969Z","shell.execute_reply":"2025-03-31T11:11:24.018691Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"scores_model","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-31T11:11:29.188275Z","iopub.execute_input":"2025-03-31T11:11:29.189183Z","iopub.status.idle":"2025-03-31T11:11:29.195139Z","shell.execute_reply.started":"2025-03-31T11:11:29.189144Z","shell.execute_reply":"2025-03-31T11:11:29.194177Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# scores_kmodel에서 모델 이름과 AUC 점수 추출\nmodels = [model['Model'] for model in scores_model]\nauc_scores = [model['AUC_Score'] for model in scores_model]\n\n# 데이터프레임으로 변환\ndf = pd.DataFrame(scores_model)\n\n# 시각화 (Seaborn을 사용한 Barplot)\nplt.figure(figsize=(8, 6))\nsns.barplot(x='Model', y='AUC_Score', data=df, palette='Pastel1')\n\n# 그래프 꾸미기\nplt.xlabel('Model')\nplt.ylabel('AUC Score')\nplt.title('Comparison of Model AUC Scores')\nplt.ylim(0, 1)  # AUC는 0과 1 사이의 값\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-31T11:11:33.568898Z","iopub.execute_input":"2025-03-31T11:11:33.569253Z","iopub.status.idle":"2025-03-31T11:11:33.677618Z","shell.execute_reply.started":"2025-03-31T11:11:33.569223Z","shell.execute_reply":"2025-03-31T11:11:33.676386Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Bi-Directional RNN's\n\n## In Depth Explanation\n\n* https://www.coursera.org/learn/nlp-sequence-models/lecture/fyXnn/bidirectional-rnn\n* https://towardsdatascience.com/understanding-bidirectional-rnn-in-pytorch-5bd25a5dd66\n* https://d2l.ai/chapter_recurrent-modern/bi-rnn.html\n\n## Code Implementation","metadata":{}},{"cell_type":"code","source":"%%time\nwith strategy.scope():\n    # A simple bidirectional LSTM with glove embeddings and one dense layer\n    model = keras.models.Sequential()\n    model.add(keras.layers.Embedding(len(word_index) + 1,\n                     300,\n                     weights=[embedding_matrix],\n                     input_length=max_len,\n                     trainable=False))\n    model.add(keras.layers.Bidirectional(keras.layers.LSTM(300, dropout=0.3, recurrent_dropout=0.3)))\n\n    model.add(keras.layers.Dense(1,activation='sigmoid'))\n    model.compile(loss='binary_crossentropy', optimizer='adam',metrics=['accuracy'])\n    \n    \nmodel.summary()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-31T11:29:01.059325Z","iopub.execute_input":"2025-03-31T11:29:01.060241Z","iopub.status.idle":"2025-03-31T11:29:01.469046Z","shell.execute_reply.started":"2025-03-31T11:29:01.060203Z","shell.execute_reply":"2025-03-31T11:29:01.468119Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Bidirectional(LSTM())\n- 양방향 LSTM 레이어. LSTM은 순차 데이터를 처리하는 RNN계열 모델인데, 양방향 LSTM은 데이터를 양방향으로 처리하여 과거와 미래의 정보를 모두 사용할 수 있게 합니다. ","metadata":{}},{"cell_type":"code","source":"model.fit(xtrain_pad, ytrain, epochs=5, batch_size=32*strategy.num_replicas_in_sync)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-31T11:29:05.106621Z","iopub.execute_input":"2025-03-31T11:29:05.106978Z","iopub.status.idle":"2025-03-31T11:45:49.357483Z","shell.execute_reply.started":"2025-03-31T11:29:05.106947Z","shell.execute_reply":"2025-03-31T11:45:49.356231Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"양방향 LSTM 레이어를 추가한 모델의 훈련이 끝나기를 기다리다가 해가 지는구나~ ","metadata":{}},{"cell_type":"code","source":"model.save('/kaggle/working/my_model.h5')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-29T09:37:19.363853Z","iopub.execute_input":"2025-03-29T09:37:19.364605Z","iopub.status.idle":"2025-03-29T09:37:19.535232Z","shell.execute_reply.started":"2025-03-29T09:37:19.364575Z","shell.execute_reply":"2025-03-29T09:37:19.534225Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"scores = model.predict(xvalid_pad)\nprint(\"Auc: %.2f%%\" % (roc_auc(scores,yvalid)))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-29T09:37:29.750001Z","iopub.execute_input":"2025-03-29T09:37:29.750335Z","iopub.status.idle":"2025-03-29T09:39:12.472953Z","shell.execute_reply.started":"2025-03-29T09:37:29.750307Z","shell.execute_reply":"2025-03-29T09:39:12.472020Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"scores_model.append({'Model': 'Bi-directional LSTM','AUC_Score': roc_auc(scores,yvalid)})","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-29T09:39:25.833219Z","iopub.execute_input":"2025-03-29T09:39:25.833560Z","iopub.status.idle":"2025-03-29T09:39:25.839918Z","shell.execute_reply.started":"2025-03-29T09:39:25.833533Z","shell.execute_reply":"2025-03-29T09:39:25.839092Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Code Explanation\n\nCode is same as before,only we have added bidirectional nature to the LSTM cells we used before and is self explanatory. We have achieve similar accuracy and auc score as before and now we have learned all the types of typical RNN architectures","metadata":{}},{"cell_type":"markdown","source":"**We are now at the end of part 1 of this notebook and things are about to go wild now as we Enter more complex and State of the art models .If you have followed along from the starting and read all the articles and understood everything , these complex models would be fairly easy to understand.I recommend Finishing Part 1 before continuing as the upcoming techniques can be quite overwhelming**","metadata":{}},{"cell_type":"markdown","source":"# Seq2Seq Model Architecture\n\n## Overview\n\nRNN's are of many types  and different architectures are used for different purposes. Here is a nice video explanining different types of model architectures : https://www.coursera.org/learn/nlp-sequence-models/lecture/BO8PS/different-types-of-rnns.\nSeq2Seq is a many to many RNN architecture where the input is a sequence and the output is also a sequence (where input and output sequences can be or cannot be of different lengths). This architecture is used in a lot of applications like Machine Translation, text summarization, question answering etc\n\n## In Depth Understanding\n\nI will not write the code implementation for this,but rather I will provide the resources where code has already been implemented and explained in a much better way than I could have ever explained.\n\n* https://www.coursera.org/learn/nlp-sequence-models/lecture/HyEui/basic-models ---> A basic idea of different Seq2Seq Models\n\n* https://blog.keras.io/a-ten-minute-introduction-to-sequence-to-sequence-learning-in-keras.html , https://machinelearningmastery.com/define-encoder-decoder-sequence-sequence-model-neural-machine-translation-keras/ ---> Basic Encoder-Decoder Model and its explanation respectively\n\n* https://towardsdatascience.com/how-to-implement-seq2seq-lstm-model-in-keras-shortcutnlp-6f355f3e5639 ---> A More advanced Seq2seq Model and its explanation\n\n* https://d2l.ai/chapter_recurrent-modern/machine-translation-and-dataset.html , https://d2l.ai/chapter_recurrent-modern/encoder-decoder.html ---> Implementation of Encoder-Decoder Model from scratch\n\n* https://www.youtube.com/watch?v=IfsjMg4fLWQ&list=PLtmWHNX-gukKocXQOkQjuVxglSDYWsSh9&index=8&t=0s ---> Introduction to Seq2seq By fast.ai","metadata":{}},{"cell_type":"markdown","source":"아래 스코어 비교에서 Bi-directionalRNN 모델은 훈련 시간이 오래걸리는 관계로 생략했다. ","metadata":{}},{"cell_type":"code","source":"# Visualization of Results obtained from various Deep learning models\nresults = pd.DataFrame(scores_model).sort_values(by='AUC_Score',ascending=False)\nresults.style.background_gradient(cmap='Blues')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-31T11:11:53.842957Z","iopub.execute_input":"2025-03-31T11:11:53.843665Z","iopub.status.idle":"2025-03-31T11:11:53.900371Z","shell.execute_reply.started":"2025-03-31T11:11:53.843628Z","shell.execute_reply":"2025-03-31T11:11:53.899467Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import plotly\nprint(plotly.__version__)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-31T11:16:35.044077Z","iopub.execute_input":"2025-03-31T11:16:35.044781Z","iopub.status.idle":"2025-03-31T11:16:35.049495Z","shell.execute_reply.started":"2025-03-31T11:16:35.044746Z","shell.execute_reply":"2025-03-31T11:16:35.048465Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import plotly.express as px\nfig = px.funnel_area(names= results.Model,\n                     values= results.AUC_Score,\n                     title= {\"position\": \"top center\", \"text\": \"Funnel-Chart of Sentiment Distribution\"})\nfig.show(renderer='iframe')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-31T11:45:49.358585Z","iopub.status.idle":"2025-03-31T11:45:49.359048Z","shell.execute_reply.started":"2025-03-31T11:45:49.358813Z","shell.execute_reply":"2025-03-31T11:45:49.358835Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"fig = go.Figure(go.Funnelarea(\n    text =results.Model,\n    values = results.AUC_Score,\n    title = {\"position\": \"top center\", \"text\": \"Funnel-Chart of Sentiment Distribution\"}\n    ))\nfig.show(renderer='iframe') # or fig.show(renderer='iframe_connected')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-31T11:18:03.456208Z","iopub.execute_input":"2025-03-31T11:18:03.456948Z","iopub.status.idle":"2025-03-31T11:18:03.487328Z","shell.execute_reply.started":"2025-03-31T11:18:03.456912Z","shell.execute_reply":"2025-03-31T11:18:03.486551Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"퍼널 영역 차트(Funnel Area Chart)\n- 데이터의 상대적 비율이나 분포를 시각화하는 도구입니다.\n- 일반적인 퍼널 차트가 단계별 감소를 표현하는 반면, 퍼널 영역 차트는 각 항목의 상대적 크기를 영역으로 보여줍니다.\n- 퍼널 영역 차트는 비율만 보여주기 때문에 실제 값의 절대적인 크기는 알 수 없습니다.\n- 작은 데이터셋에서는 미미한 AUC 차이가 과장되어 보일 수 있습니다.\n\n퍼널 영역 차트를 통해 최고 성능 모델과 다른 모델들 간의 성능 차이가 얼마나 큰 지 확인할 수 있다.  \nglove 워드 임베딩을 사용한 LSTM과 GRU가 simpleRNN 보다 좋은 성능을 보여주는 반면, LSTM과 GRU의 차이는 미미하다. ","metadata":{}},{"cell_type":"markdown","source":"# Attention Models\n\nThis is the toughest and most tricky part. If you are able to understand the intiuition and working of attention block , understanding transformers and transformer based architectures like BERT will be a piece of cake. This is the part where I spent the most time on and I suggest you do the same . Please read and view the following resources in the order I am providing to ignore getting confused, also at the end of this try to write and draw an attention block in your own way :-\n\n* https://www.coursera.org/learn/nlp-sequence-models/lecture/RDXpX/attention-model-intuition --> Only watch this video and not the next one\n* https://towardsdatascience.com/sequence-2-sequence-model-with-attention-mechanism-9e9ca2a613a\n* https://towardsdatascience.com/attention-and-its-different-forms-7fc3674d14dc\n* https://distill.pub/2016/augmented-rnns/ \n\n## Code Implementation\n\n* https://www.analyticsvidhya.com/blog/2019/11/comprehensive-guide-attention-mechanism-deep-learning/ --> Basic Level\n* https://pytorch.org/tutorials/intermediate/seq2seq_translation_tutorial.html ---> Implementation from Scratch in Pytorch","metadata":{}},{"cell_type":"markdown","source":"# Transformers : Attention is all you need\n\nSo finally we have reached the end of the learning curve and are about to start learning the technology that changed NLP completely and are the reasons for the state of the art NLP techniques .Transformers were introduced in the paper Attention is all you need by Google. If you have understood the Attention models,this will be very easy , Here is transformers fully explained:\n\n* http://jalammar.github.io/illustrated-transformer/\n\n## Code Implementation\n\n* http://nlp.seas.harvard.edu/2018/04/03/attention.html ---> This presents the code implementation of the architecture presented in the paper by Google","metadata":{}},{"cell_type":"markdown","source":"# BERT and Its Implementation on this Competition\n\nAs Promised I am back with Resiurces , to understand about BERT architecture , please follow the contents in the given order :-\n\n* http://jalammar.github.io/illustrated-bert/ ---> In Depth Understanding of BERT\n\nAfter going through the post Above , I guess you must have understood how transformer architecture have been utilized by the current SOTA models . Now these architectures can be used in two ways :<br><br>\n1) We can use the model for prediction on our problems using the pretrained weights without fine-tuning or training the model for our sepcific tasks\n* EG: http://jalammar.github.io/a-visual-guide-to-using-bert-for-the-first-time/ ---> Using Pre-trained BERT without Tuning\n\n2) We can fine-tune or train these transformer models for our task by tweaking the already pre-trained weights and training on a much smaller dataset\n* EG:* https://www.youtube.com/watch?v=hinZO--TEk4&t=2933s ---> Tuning BERT For your TASK\n\nWe will be using the first example as a base for our implementation of BERT model using Hugging Face and KERAS , but contrary to first example we will also Fine-Tune our model for our task\n\nAcknowledgements : https://www.kaggle.com/xhlulu/jigsaw-tpu-distilbert-with-huggingface-and-keras\n\n\nSteps Involved :\n* Data Preparation : Tokenization and encoding of data\n* Configuring TPU's \n* Building a Function for Model Training and adding an output layer for classification\n* Train the model and get the results","metadata":{}},{"cell_type":"code","source":"# Loading Dependencies\nimport os\nimport tensorflow as tf\nfrom tensorflow.keras.layers import Dense, Input\nfrom tensorflow.keras.optimizers import Adam\nfrom tensorflow.keras.models import Model\nfrom tensorflow.keras.callbacks import ModelCheckpoint\nfrom kaggle_datasets import KaggleDatasets\nimport transformers\n\nfrom tokenizers import BertWordPieceTokenizer","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-31T11:46:51.746405Z","iopub.execute_input":"2025-03-31T11:46:51.747094Z","iopub.status.idle":"2025-03-31T11:46:51.974148Z","shell.execute_reply.started":"2025-03-31T11:46:51.747057Z","shell.execute_reply":"2025-03-31T11:46:51.973477Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# LOADING THE DATA\n\ntrain1 = pd.read_csv(\"/kaggle/input/jigsaw-multilingual-toxic-comment-classification/jigsaw-toxic-comment-train.csv\")\nvalid = pd.read_csv('/kaggle/input/jigsaw-multilingual-toxic-comment-classification/validation.csv')\ntest = pd.read_csv('/kaggle/input/jigsaw-multilingual-toxic-comment-classification/test.csv')\nsub = pd.read_csv('/kaggle/input/jigsaw-multilingual-toxic-comment-classification/sample_submission.csv')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-31T11:46:55.178464Z","iopub.execute_input":"2025-03-31T11:46:55.179300Z","iopub.status.idle":"2025-03-31T11:46:57.038747Z","shell.execute_reply.started":"2025-03-31T11:46:55.179266Z","shell.execute_reply":"2025-03-31T11:46:57.037968Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Encoder FOr DATA for understanding waht encode batch does read documentation of hugging face tokenizer :\nhttps://huggingface.co/transformers/main_classes/tokenizer.html here","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"code","source":"# 텍스트 데이터를 Bert 모델 입력에 적합한 정수 시퀀스로 인코딩하는 함수 \ndef fast_encode(texts, tokenizer, chunk_size=256, maxlen=512):\n    \"\"\"\n    Encoder for encoding the text into sequence of integers for BERT Input\n\n    parameters\n    ----------\n    texts:\n        인코딩할 텍스트 리스트나 배열\n    tokenizer:\n        토크나이저 객체\n        (라이브러리 사용)\n    chunk_size:\n        한 번에 처리할 텍스트 청크 크기\n        (기본값: 256)\n    maxlen:\n        최대 시퀀스 길이\n        (기본값: 512)\n    \"\"\"\n    tokenizer.enable_truncation(max_length=maxlen) # maxlen보다 길면 잘라내기\n    tokenizer.enable_padding(length=maxlen)    # maxlen보다 짧으면 패딩\n    all_ids = []\n    \n    for i in tqdm(range(0, len(texts), chunk_size)):\n        text_chunk = texts[i:i+chunk_size].tolist()  # 메모리 효율성을 위해 텍스트를 작은 청크로 나눠 리스트 변환 \n        encs = tokenizer.encode_batch(text_chunk)    # 청크 내 모든 텍스트를 한 번에 인코딩\n        all_ids.extend([enc.ids for enc in encs])    # 인코딩된 각 객체에서 ID 목록을 추출해서 리스트에 추가\n    \n    return np.array(all_ids) # 인코딩된 ID를 NumPy 배열로 반환 [텍스트 수, maxlen] -> BERT모델의 입력 데이터로 사용","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-31T12:06:15.672809Z","iopub.execute_input":"2025-03-31T12:06:15.673727Z","iopub.status.idle":"2025-03-31T12:06:15.680082Z","shell.execute_reply.started":"2025-03-31T12:06:15.673691Z","shell.execute_reply":"2025-03-31T12:06:15.679183Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# 모델의 설정 매개변수 정의\n#IMP DATA FOR CONFIG\n\nAUTO = tf.data.experimental.AUTOTUNE\n\n\n# Configuration\nEPOCHS = 3\nBATCH_SIZE = 16 * strategy.num_replicas_in_sync\nMAX_LEN = 192","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-31T11:55:05.606908Z","iopub.execute_input":"2025-03-31T11:55:05.607685Z","iopub.status.idle":"2025-03-31T11:55:05.612263Z","shell.execute_reply.started":"2025-03-31T11:55:05.607650Z","shell.execute_reply":"2025-03-31T11:55:05.611251Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Tokenization\n\nFor understanding please refer to hugging face documentation again","metadata":{}},{"cell_type":"code","source":"# ipywidgets 패키지 업데이트\n!pip install --upgrade ipywidgets","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-31T11:56:37.908672Z","iopub.execute_input":"2025-03-31T11:56:37.909045Z","iopub.status.idle":"2025-03-31T11:56:46.988700Z","shell.execute_reply.started":"2025-03-31T11:56:37.909014Z","shell.execute_reply":"2025-03-31T11:56:46.987501Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# jupyter 위젯 버전 충돌 오류 무시\nimport warnings\nwarnings.filterwarnings('ignore', category=UserWarning, module='ipywidgets')\n\n# 커널 재시작시 위젯 버전 다운그레이드 및 재설치\n!pip uninstall -y ipywidgets\n!pip install ipywidgets==7.6.5\n!pip install jupyterlab-widgets==1.0.0","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-31T11:59:53.129054Z","iopub.execute_input":"2025-03-31T11:59:53.129799Z","iopub.status.idle":"2025-03-31T11:59:53.134304Z","shell.execute_reply.started":"2025-03-31T11:59:53.129764Z","shell.execute_reply":"2025-03-31T11:59:53.133367Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# First load the real tokenizer\ntokenizer = transformers.DistilBertTokenizer.from_pretrained('distilbert-base-multilingual-cased')\n# Save the loaded tokenizer locally\ntokenizer.save_pretrained('.')\n# Reload it with the huggingface tokenizers library\nfast_tokenizer = BertWordPieceTokenizer('vocab.txt', lowercase=False)\nfast_tokenizer","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-31T12:00:23.755413Z","iopub.execute_input":"2025-03-31T12:00:23.756279Z","iopub.status.idle":"2025-03-31T12:00:24.733811Z","shell.execute_reply.started":"2025-03-31T12:00:23.756243Z","shell.execute_reply":"2025-03-31T12:00:24.732779Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Bert 모델에서 사용하는 WordPiece 토크나이저의 구성정보가 출력된다.  \n토큰화를 수행할 때, 단어를 서브워드 단위로 분할하는 방식을 사용한다.  \n위에서 정의한 fast_encode() 함수와 함께 사용","metadata":{}},{"cell_type":"code","source":"x_train = fast_encode(train1.comment_text.astype(str), fast_tokenizer, maxlen=MAX_LEN)\nx_valid = fast_encode(valid.comment_text.astype(str), fast_tokenizer, maxlen=MAX_LEN)\nx_test = fast_encode(test.content.astype(str), fast_tokenizer, maxlen=MAX_LEN)\n\ny_train = train1.toxic.values\ny_valid = valid.toxic.values","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-31T12:06:29.321046Z","iopub.execute_input":"2025-03-31T12:06:29.321405Z","iopub.status.idle":"2025-03-31T12:07:08.775708Z","shell.execute_reply.started":"2025-03-31T12:06:29.321374Z","shell.execute_reply":"2025-03-31T12:07:08.774944Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_dataset = (\n    tf.data.Dataset\n    .from_tensor_slices((x_train, y_train))\n    .repeat()\n    .shuffle(2048)\n    .batch(BATCH_SIZE)\n    .prefetch(AUTO)\n)\n\nvalid_dataset = (\n    tf.data.Dataset\n    .from_tensor_slices((x_valid, y_valid))\n    .batch(BATCH_SIZE)\n    .cache()\n    .prefetch(AUTO)\n)\n\ntest_dataset = (\n    tf.data.Dataset\n    .from_tensor_slices(x_test)\n    .batch(BATCH_SIZE)\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-31T12:07:59.542591Z","iopub.execute_input":"2025-03-31T12:07:59.543615Z","iopub.status.idle":"2025-03-31T12:08:00.544878Z","shell.execute_reply.started":"2025-03-31T12:07:59.543571Z","shell.execute_reply":"2025-03-31T12:08:00.543898Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def build_model(transformer, max_len=512):\n    \"\"\"\n    function for training the BERT model\n    \"\"\"\n    # 최대 길이가 512인 단어 ID 입력 레이어 생성\n    input_word_ids = Input(shape=(max_len,), dtype=tf.int32, name=\"input_word_ids\")\n    # 트랜스포머(BERT) 모델의 첫번째 시퀀스 출력 \n    sequence_output = transformer(input_word_ids)[0]\n    # 전체 문장의 정보를 담고 있는 cls_token\n    cls_token = sequence_output[:, 0, :]\n\n    out = Dense(1, activation='sigmoid')(cls_token)\n    \n    # 입력과 출력을 연결해 모델 생성\n    model = Model(inputs=input_word_ids, outputs=out)\n    model.compile(Adam(learning_rate=1e-5), loss='binary_crossentropy', metrics=['accuracy'])\n    \n    return model","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-31T12:14:27.526674Z","iopub.execute_input":"2025-03-31T12:14:27.527048Z","iopub.status.idle":"2025-03-31T12:14:27.534104Z","shell.execute_reply.started":"2025-03-31T12:14:27.527017Z","shell.execute_reply":"2025-03-31T12:14:27.532783Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Starting Training\n\nIf you want to use any another model just replace the model name in transformers._____ and use accordingly","metadata":{}},{"cell_type":"code","source":"# 다국어 DistillBERT 모델 기반 이진 분류 모델\n%%time\nwith strategy.scope():\n    transformer_layer = (\n        transformers.TFDistilBertModel\n        .from_pretrained('distilbert-base-multilingual-cased')\n    )\n    model = build_model(transformer_layer, max_len=MAX_LEN)\nmodel.summary()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-31T12:14:46.359560Z","iopub.execute_input":"2025-03-31T12:14:46.360253Z","iopub.status.idle":"2025-03-31T12:15:10.633830Z","shell.execute_reply.started":"2025-03-31T12:14:46.360224Z","shell.execute_reply":"2025-03-31T12:15:10.632931Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"n_steps = x_train.shape[0] // BATCH_SIZE\ntrain_history = model.fit(\n    train_dataset,\n    steps_per_epoch=n_steps,\n    validation_data=valid_dataset,\n    epochs=EPOCHS\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-31T12:17:40.596171Z","iopub.execute_input":"2025-03-31T12:17:40.597405Z","iopub.status.idle":"2025-03-31T17:17:50.797412Z","shell.execute_reply.started":"2025-03-31T12:17:40.597369Z","shell.execute_reply":"2025-03-31T17:17:50.796477Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"n_steps = x_valid.shape[0] // BATCH_SIZE\ntrain_history_2 = model.fit(\n    valid_dataset.repeat(),\n    steps_per_epoch=n_steps,\n    epochs=EPOCHS*2\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-31T17:18:02.057582Z","iopub.execute_input":"2025-03-31T17:18:02.058324Z","iopub.status.idle":"2025-03-31T17:39:17.898403Z","shell.execute_reply.started":"2025-03-31T17:18:02.058295Z","shell.execute_reply":"2025-03-31T17:39:17.897492Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"sub['toxic'] = model.predict(test_dataset, verbose=1)\nsub.to_csv('submission.csv', index=False)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-31T17:39:35.281471Z","iopub.execute_input":"2025-03-31T17:39:35.282288Z","iopub.status.idle":"2025-03-31T17:48:15.518267Z","shell.execute_reply.started":"2025-03-31T17:39:35.282255Z","shell.execute_reply":"2025-03-31T17:48:15.517263Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# End Notes\n\nThis was my effort to share my learnings so that everyone can benifit from it.As this community has been very kind to me and helped me in learning all of this , I want to take this forward. I have shared all the resources I used to learn all the stuff .Join me and make these NLP competitions your first ,without being overwhelmed by the shear number of techniques used . It took me 10 days to learn all of this , you can learn it at your pace and dont give in , at the end of all this you will be a different person and it will all be worth it.\n\n\n### I am attaching more resources if you want NLP end to end:\n\n1) Books\n\n* https://d2l.ai/\n* Jason Brownlee's Books\n\n2) Courses\n\n* https://www.coursera.org/learn/nlp-sequence-models/home/welcome\n* Fast.ai NLP Course\n\n3) Blogs and websites\n\n* Machine Learning Mastery\n* https://distill.pub/\n* http://jalammar.github.io/\n\n**<span style=\"color:Red\">This is subtle effort of contributing towards the community, if it helped you in any way please show a token of love by upvoting**","metadata":{}}]}