{"cells":[{"metadata":{"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5"},"cell_type":"markdown","source":"# CLASSIFIY TOXIC COMMENT \n### by XLM-R \n#### Sarah , Abdulelah and Norah","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"XLM-R It is based on Facebook’s RoBERTa model released in 2019. It is a large multi-lingual language model, trained on 2.5TB of filtered CommonCrawl data. It is optimized version of  RoBERTa  the only noteworthy difference  between XLM-R and RoBERTa is the vocabulary size: 250 thousand tokens in XLM-R compared to RoBERTa’s 50,000 tokens. Note that this makes the model significantly larger, 550 million parameters compared to the 355 million of RoBERTa.  significantly outperforms multilingual BERT (mBERT) on a variety of cross-lingual benchmarks, including +13.8% average accuracy on XNLI, +12.3% average F1 score on MLQA, and +2.1%","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"Import Libraries\nRun Bert Model on TPU\nFunctions and Variables\n3.1 Function for Encoding the comment\n3.2 Function for Neural Network model\nPreprocessing \n4.1 confegration\n4.2 Import Datasets\n4.3 tokenaizer\n4.4 Encode The Comments\n4.5 Prepare tensorflow dataset for modeling\nMachine Learning\n5.1 Build the model\n5.2 Training The Model, Tuning Hyper-Parameters\n5.3 Testing The Model","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"### Importing libraries ","execution_count":null},{"metadata":{"trusted":false},"cell_type":"code","source":"import os\nimport numpy as np\nimport pandas as pd\nimport tensorflow as tf\nfrom tensorflow.keras.layers import Dense, Input\nfrom tensorflow.keras.optimizers import Adam\nfrom tensorflow.keras.models import Model\nfrom tensorflow.keras.callbacks import ModelCheckpoint\nfrom kaggle_datasets import KaggleDatasets\nimport transformers\nfrom transformers import TFAutoModel, AutoTokenizer\nfrom tqdm.notebook import tqdm\nfrom tokenizers import Tokenizer, models, pre_tokenizers, decoders, processors","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Encoding the comment\n#### encoder job to take word and embedding it  to be in  “numeric representation”  \nnote:we just need decoder in translation task","execution_count":null},{"metadata":{"trusted":false},"cell_type":"code","source":"def regular_encode(texts, tokenizer, maxlen=512):\n    \"\"\"\n    Function to encode the word\n    \"\"\"\n    # encode the word to vector of integer\n    enc_di = tokenizer.batch_encode_plus(\n        texts, \n        return_attention_masks=False, \n        return_token_type_ids=False,\n        pad_to_max_length=True,\n        max_length=maxlen\n    )\n    \n    return np.array(enc_di['input_ids'])","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Building The model and compile it ","execution_count":null},{"metadata":{"trusted":false},"cell_type":"code","source":"def build_model(transformer, max_len=512):\n    \"\"\"\n    This function to build and compile Keras model\n    \n    \"\"\"\n    #Input: for define input layer\n    #shape is vector with 512-dimensional vectors\n    input_word_ids = Input(shape=(max_len,), dtype=tf.int32, name=\"input_word_ids\") # name is optional \n    sequence_output = transformer(input_word_ids)[0]\n    # to get the vector\n    cls_token = sequence_output[:, 0, :]\n    # define output layer\n    out = Dense(1, activation='sigmoid')(cls_token)\n    \n    # initiate the model with inputs and outputs\n    model = Model(inputs=input_word_ids, outputs=out)\n    model.compile(Adam(lr=1e-5), loss='binary_crossentropy',metrics=[tf.keras.metrics.AUC()])\n    \n    return model","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"#### Run Model in TPU","execution_count":null},{"metadata":{"trusted":false},"cell_type":"code","source":"# Detect hardware, return appropriate distribution strategy\ntry:\n    # TPU detection. No parameters necessary if TPU_NAME environment variable is\n    # set: this is always the case on Kaggle.\n    tpu = tf.distribute.cluster_resolver.TPUClusterResolver()\n    print('Running on TPU ', tpu.master())\nexcept ValueError:\n    tpu = None\n\nif tpu:\n    tf.config.experimental_connect_to_cluster(tpu)\n    tf.tpu.experimental.initialize_tpu_system(tpu)\n    strategy = tf.distribute.experimental.TPUStrategy(tpu)\nelse:\n    # Default distribution strategy in Tensorflow. Works on CPU and single GPU.\n    strategy = tf.distribute.get_strategy()\n\nprint(\"REPLICAS: \", strategy.num_replicas_in_sync)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"#### Some configration variables","execution_count":null},{"metadata":{"trusted":false},"cell_type":"code","source":"\n# input pipeline that delivers data for the next step before the current step has finished.\n# The tf.data API helps to build flexible and efficient input pipelines.\n# This document demonstrates how to use the tf.data \n# API to build highly performant TensorFlow input pipelines.\nAUTO = tf.data.experimental.AUTOTUNE\n\n# upload data into google cloud storage\nGCS_DS_PATH = KaggleDatasets().get_gcs_path()\n\n# Configuration\nEPOCHS = 2\nBATCH_SIZE = 16 * strategy.num_replicas_in_sync\nMAX_LEN = 192\nMODEL = 'jplu/tf-xlm-roberta-large'","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### download the xlm model Tokinazer","execution_count":null},{"metadata":{"trusted":false},"cell_type":"code","source":"# First load the real tokenizer\ntokenizer = AutoTokenizer.from_pretrained(MODEL)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Load the data and apply encoding to the data","execution_count":null},{"metadata":{"trusted":false},"cell_type":"code","source":"\ntrain1 = pd.read_csv(\"/kaggle/input/jigsaw-multilingual-toxic-comment-classification/jigsaw-toxic-comment-train.csv\")\nvalid = pd.read_csv('/kaggle/input/jigsaw-multilingual-toxic-comment-classification/validation.csv')\ntest = pd.read_csv('/kaggle/input/jigsaw-multilingual-toxic-comment-classification/test.csv')\nsub = pd.read_csv('/kaggle/input/jigsaw-multilingual-toxic-comment-classification/sample_submission.csv')","execution_count":null,"outputs":[]},{"metadata":{"trusted":false},"cell_type":"code","source":"#applied regular_encode() function to all data set\n%%time \n\nx_train = regular_encode(train1.comment_text.values, tokenizer, maxlen=MAX_LEN)\nx_valid = regular_encode(valid.comment_text.values, tokenizer, maxlen=MAX_LEN)\nx_test = regular_encode(test.content.values, tokenizer, maxlen=MAX_LEN)\n\ny_train = train1.toxic.values\ny_valid = valid.toxic.values","execution_count":null,"outputs":[]},{"metadata":{"trusted":false},"cell_type":"code","source":"\n# Create a source dataset from your input data.\n# Apply dataset transformations to preprocess the data.\n# Iterate over the dataset and process the elements.\n\ntrain_dataset = (\n    tf.data.Dataset # create dataset\n    .from_tensor_slices((x_train, y_train)) # Once you have a dataset, you can apply transformations \n    .repeat()\n    .shuffle(2048)\n    .batch(BATCH_SIZE)# Combines consecutive elements of this dataset into batches.\n    .prefetch(AUTO) #This allows later elements to be prepared while the current element is being processed.\n)\n\nvalid_dataset = (\n    tf.data.Dataset # create dataset\n    .from_tensor_slices((x_valid, y_valid)) # Once you have a dataset, you can apply transformations \n    .batch(BATCH_SIZE) #Combines consecutive elements of this dataset into batches.\n    .cache()\n    .prefetch(AUTO)#This allows later elements to be prepared while the current element is being processed.\n)\n\ntest_dataset = (\n    tf.data.Dataset# create dataset\n    .from_tensor_slices(x_test) # Once you have a dataset, you can apply transformations \n    .batch(BATCH_SIZE)\n)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Dwonload the model and create Keras model","execution_count":null},{"metadata":{"trusted":false},"cell_type":"code","source":"%%time\nwith strategy.scope():\n    # load model and build it\n    transformer_layer = TFAutoModel.from_pretrained(MODEL)\n    model = build_model(transformer_layer, max_len=MAX_LEN)\nmodel.summary()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Modil fit on train and validate data","execution_count":null},{"metadata":{"trusted":false},"cell_type":"code","source":"# fit the model on train data\nn_steps = x_train.shape[0] // BATCH_SIZE\ntrain_history = model.fit(\n    train_dataset,\n    steps_per_epoch=n_steps,\n    validation_data=valid_dataset,\n    epochs=EPOCHS +3\n)","execution_count":null,"outputs":[]},{"metadata":{"trusted":false},"cell_type":"code","source":"# fit model on validation\nn_steps = x_valid.shape[0] // BATCH_SIZE\ntrain_history_2 = model.fit(\n    valid_dataset.repeat(),\n    steps_per_epoch=n_steps,\n    epochs=EPOCHS * 3\n)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Predict ","execution_count":null},{"metadata":{"trusted":false},"cell_type":"code","source":"sub['toxic'] = model.predict(test_dataset, verbose=1)\nsub.to_csv('submission.csv', index=False)","execution_count":null,"outputs":[]},{"metadata":{"trusted":false},"cell_type":"code","source":"# save model weights \nmodel.save_weights('xlm_weights.h5')","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}