{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.7.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceId":19018,"databundleVersionId":2703900,"sourceType":"competition"}],"dockerImageVersionId":30299,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## About this notebook\n\n*[Jigsaw Multilingual Toxic Comment Classification](https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification)* is the 3rd annual competition organized by the Jigsaw team. It follows *[Toxic Comment Classification Challenge](https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge)*, the original 2018 competition, and *[Jigsaw Unintended Bias in Toxicity Classification](https://www.kaggle.com/c/jigsaw-unintended-bias-in-toxicity-classification)*, which required the competitors to consider biased ML predictions in their new models. This year, the goal is to use english only training data to run toxicity predictions on many different languages, which can be done using multilingual models, and speed up using TPUs.\n\nMany awesome notebooks has already been made so far. Many of them used really cool technologies like [Pytorch XLA](https://www.kaggle.com/theoviel/bert-pytorch-huggingface-starter). This notebook instead aims at constructing a **fast, concise, reusable, and beginner-friendly model scaffold**. It will focus on the following points:\n* **Using Tensorflow and Keras**: Tensorflow is a powerful framework, and Keras makes the training process extremely easy to understand. This is especially good for beginners to learn how to use TPUs, and for experts to focus on the modelling aspect.\n* **Huggingface's transformers library**: [This library](https://huggingface.co/transformers/) is extremely popular, so using this let you easily integrate the end result into your ML pipelines, and can be easily reused for your other projects.\n* **Multilingual DistilBERT**: DistilBERT is **2 times faster and 25% lighter** than multilingual BERT base, all while retaining **92% of its performance**. This model let you quickly experiments with different ideas, and when you are ready for the real thing, just change two lines of code to use `bert-base-multilingual-cased`.\n* **Blazing fast tokenization**: [Huggingface's `tokenizers`](https://github.com/huggingface/tokenizers/tree/master/bindings/python) is order of magnitude faster than the default BERT tokenizer, since it is written in Rust, and uses a Python interface.\n* **Native TPU usage**: The TPU usage is abstracted using the native `strategy` that was created using Tensorflow's `tf.distribute.experimental.TPUStrategy`. This avoids getting too much into the lower-level aspect of TPU management.\n* **Subset of the data**: Instead of using the entire dataset, we will only use the 2018 subset of the data available, which makes this much faster, all while achieving a respectable accuracy.\n\n\n### References\n* Original Author: [@xhlulu](https://www.kaggle.com/xhlulu/)\n* Original notebook: [Link](https://www.kaggle.com/xhlulu/jigsaw-tpu-distilbert-with-huggingface-and-keras)","metadata":{}},{"cell_type":"code","source":"import os\n\nimport numpy as np\nimport pandas as pd\nimport tensorflow as tf\nfrom tensorflow.keras.layers import Dense, Input\nfrom tensorflow.keras.optimizers import Adam\nfrom tensorflow.keras.models import Model\nfrom tensorflow.keras.callbacks import ModelCheckpoint\nfrom kaggle_datasets import KaggleDatasets\nimport transformers\nfrom tqdm.notebook import tqdm\nfrom tokenizers import BertWordPieceTokenizer","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2024-10-13T22:04:26.785339Z","iopub.execute_input":"2024-10-13T22:04:26.785978Z","iopub.status.idle":"2024-10-13T22:04:26.792712Z","shell.execute_reply.started":"2024-10-13T22:04:26.785941Z","shell.execute_reply":"2024-10-13T22:04:26.791731Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Helper Functions","metadata":{}},{"cell_type":"code","source":"def fast_encode(texts, tokenizer, chunk_size=256, maxlen=512):\n    \"\"\"\n    https://www.kaggle.com/xhlulu/jigsaw-tpu-distilbert-with-huggingface-and-keras\n    \"\"\"\n    tokenizer.enable_truncation(max_length=maxlen)\n    tokenizer.enable_padding()\n    all_ids = []\n    \n    for i in tqdm(range(0, len(texts), chunk_size)):\n        text_chunk = texts[i:i+chunk_size].tolist()\n        encs = tokenizer.encode_batch(text_chunk)\n        all_ids.extend([enc.ids for enc in encs])\n    \n    return np.array(all_ids)","metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","execution":{"iopub.status.busy":"2024-10-13T22:04:27.606530Z","iopub.execute_input":"2024-10-13T22:04:27.607197Z","iopub.status.idle":"2024-10-13T22:04:27.614437Z","shell.execute_reply.started":"2024-10-13T22:04:27.607163Z","shell.execute_reply":"2024-10-13T22:04:27.613471Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def build_model(transformer, max_len=512):\n    \"\"\"\n    https://www.kaggle.com/xhlulu/jigsaw-tpu-distilbert-with-huggingface-and-keras\n    \"\"\"\n    input_word_ids = Input(shape=(max_len,), dtype=tf.int32, name=\"input_word_ids\")\n    sequence_output = transformer(input_word_ids)[0]\n    cls_token = sequence_output[:, 0, :]\n    out = Dense(1, activation='sigmoid')(cls_token)\n    \n    model = Model(inputs=input_word_ids, outputs=out)\n    model.compile(Adam(lr=1e-5), loss='binary_crossentropy', metrics=['accuracy'])\n    \n    return model","metadata":{"execution":{"iopub.status.busy":"2024-10-13T22:04:32.227588Z","iopub.execute_input":"2024-10-13T22:04:32.227944Z","iopub.status.idle":"2024-10-13T22:04:32.234922Z","shell.execute_reply.started":"2024-10-13T22:04:32.227916Z","shell.execute_reply":"2024-10-13T22:04:32.233986Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## TPU Configs","metadata":{}},{"cell_type":"code","source":"# Detect hardware, return appropriate distribution strategy\ntry:\n    # TPU detection. No parameters necessary if TPU_NAME environment variable is\n    # set: this is always the case on Kaggle.\n    tpu = tf.distribute.cluster_resolver.TPUClusterResolver()\n    print('Running on TPU ', tpu.master())\nexcept ValueError:\n    tpu = None\n\nif tpu:\n    tf.config.experimental_connect_to_cluster(tpu)\n    tf.tpu.experimental.initialize_tpu_system(tpu)\n    strategy = tf.distribute.experimental.TPUStrategy(tpu)\nelse:\n    # Default distribution strategy in Tensorflow. Works on CPU and single GPU.\n    strategy = tf.distribute.get_strategy()\n\nprint(\"REPLICAS: \", strategy.num_replicas_in_sync)","metadata":{"execution":{"iopub.status.busy":"2024-10-13T21:45:52.474906Z","iopub.execute_input":"2024-10-13T21:45:52.475825Z","iopub.status.idle":"2024-10-13T21:45:52.492321Z","shell.execute_reply.started":"2024-10-13T21:45:52.475787Z","shell.execute_reply":"2024-10-13T21:45:52.491299Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"AUTO = tf.data.experimental.AUTOTUNE\n\n# Data access\nGCS_DS_PATH = KaggleDatasets().get_gcs_path()\n\n# Configuration\nEPOCHS = 3\nBATCH_SIZE = 16 * strategy.num_replicas_in_sync\nMAX_LEN = 192","metadata":{"execution":{"iopub.status.busy":"2024-10-13T21:45:59.930553Z","iopub.execute_input":"2024-10-13T21:45:59.931543Z","iopub.status.idle":"2024-10-13T21:48:29.666931Z","shell.execute_reply.started":"2024-10-13T21:45:59.931497Z","shell.execute_reply":"2024-10-13T21:48:29.665996Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Create fast tokenizer","metadata":{}},{"cell_type":"code","source":"# First load the real tokenizer\ntokenizer = transformers.DistilBertTokenizer.from_pretrained('distilbert-base-multilingual-cased')\n# Save the loaded tokenizer locally\ntokenizer.save_pretrained('.')\n# Reload it with the huggingface tokenizers library\nfast_tokenizer = BertWordPieceTokenizer('vocab.txt', lowercase=False)\nfast_tokenizer","metadata":{"execution":{"iopub.status.busy":"2024-10-13T21:48:29.668897Z","iopub.execute_input":"2024-10-13T21:48:29.669606Z","iopub.status.idle":"2024-10-13T21:48:31.646482Z","shell.execute_reply.started":"2024-10-13T21:48:29.669563Z","shell.execute_reply":"2024-10-13T21:48:31.645488Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Load text data into memory","metadata":{}},{"cell_type":"code","source":"train1 = pd.read_csv(\"/kaggle/input/jigsaw-multilingual-toxic-comment-classification/jigsaw-toxic-comment-train.csv\")\ntrain2 = pd.read_csv(\"/kaggle/input/jigsaw-multilingual-toxic-comment-classification/jigsaw-unintended-bias-train.csv\")\ntrain2.toxic = train2.toxic.round().astype(int)\n\nvalid = pd.read_csv('/kaggle/input/jigsaw-multilingual-toxic-comment-classification/validation.csv')\ntest = pd.read_csv('/kaggle/input/jigsaw-multilingual-toxic-comment-classification/test.csv')\nsub = pd.read_csv('/kaggle/input/jigsaw-multilingual-toxic-comment-classification/sample_submission.csv')","metadata":{"execution":{"iopub.status.busy":"2024-10-13T21:48:35.181099Z","iopub.execute_input":"2024-10-13T21:48:35.181463Z","iopub.status.idle":"2024-10-13T21:49:01.320574Z","shell.execute_reply.started":"2024-10-13T21:48:35.181432Z","shell.execute_reply":"2024-10-13T21:49:01.319548Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Combine train1 with a subset of train2\ntrain = pd.concat([\n    train1[['comment_text', 'toxic']],\n    train2[['comment_text', 'toxic']].query('toxic==1'),\n    train2[['comment_text', 'toxic']].query('toxic==0').sample(n=150000, random_state=0)\n])","metadata":{"execution":{"iopub.status.busy":"2024-10-13T21:49:06.992752Z","iopub.execute_input":"2024-10-13T21:49:06.993629Z","iopub.status.idle":"2024-10-13T21:49:07.524460Z","shell.execute_reply.started":"2024-10-13T21:49:06.993590Z","shell.execute_reply":"2024-10-13T21:49:07.523307Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"x_train = fast_encode(train.comment_text.astype(str), fast_tokenizer, maxlen=MAX_LEN)\nx_valid = fast_encode(valid.comment_text.astype(str), fast_tokenizer, maxlen=MAX_LEN)\nx_test = fast_encode(test.content.astype(str), fast_tokenizer, maxlen=MAX_LEN)\n\ny_train = train.toxic.values\ny_valid = valid.toxic.values","metadata":{"execution":{"iopub.status.busy":"2024-10-13T21:50:41.020868Z","iopub.execute_input":"2024-10-13T21:50:41.021904Z","iopub.status.idle":"2024-10-13T21:52:00.854754Z","shell.execute_reply.started":"2024-10-13T21:50:41.021864Z","shell.execute_reply":"2024-10-13T21:52:00.853776Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Build datasets objects","metadata":{}},{"cell_type":"code","source":"train_dataset = (\n    tf.data.Dataset\n    .from_tensor_slices((x_train, y_train))\n    .repeat()\n    .shuffle(2048)\n    .batch(BATCH_SIZE)\n    .prefetch(AUTO)\n)\n\nvalid_dataset = (\n    tf.data.Dataset\n    .from_tensor_slices((x_valid, y_valid))\n    .batch(BATCH_SIZE)\n    .cache()\n    .prefetch(AUTO)\n)\n\ntest_dataset = (\n    tf.data.Dataset\n    .from_tensor_slices(x_test)\n    .batch(BATCH_SIZE)\n)","metadata":{"execution":{"iopub.status.busy":"2024-10-13T21:52:00.856360Z","iopub.execute_input":"2024-10-13T21:52:00.856627Z","iopub.status.idle":"2024-10-13T21:52:06.150917Z","shell.execute_reply.started":"2024-10-13T21:52:00.856601Z","shell.execute_reply":"2024-10-13T21:52:06.150073Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Load model into the TPU","metadata":{}},{"cell_type":"code","source":"%%time\nwith strategy.scope():\n    transformer_layer = (\n        transformers.TFDistilBertModel\n        .from_pretrained('distilbert-base-multilingual-cased')\n    )\n    model = build_model(transformer_layer, max_len=MAX_LEN)\nmodel.summary()","metadata":{"execution":{"iopub.status.busy":"2024-10-13T21:52:41.683947Z","iopub.execute_input":"2024-10-13T21:52:41.684792Z","iopub.status.idle":"2024-10-13T21:53:17.685224Z","shell.execute_reply.started":"2024-10-13T21:52:41.684756Z","shell.execute_reply":"2024-10-13T21:53:17.684243Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Train Model","metadata":{}},{"cell_type":"markdown","source":"First, we train on the subset of the training set, which is completely in English.","metadata":{}},{"cell_type":"code","source":"n_steps = x_train.shape[0] // BATCH_SIZE\ntrain_history = model.fit(\n    train_dataset,\n    steps_per_epoch=n_steps,\n    validation_data=valid_dataset,\n    epochs=1\n)","metadata":{"execution":{"iopub.status.busy":"2024-10-13T21:54:58.995767Z","iopub.execute_input":"2024-10-13T21:54:58.996514Z","iopub.status.idle":"2024-10-13T21:57:57.446259Z","shell.execute_reply.started":"2024-10-13T21:54:58.996476Z","shell.execute_reply":"2024-10-13T21:57:57.444517Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Now that we have pretty much saturated the learning potential of the model on english only data, we train it for one more epoch on the `validation` set, which is significantly smaller but contains a mixture of different languages.","metadata":{}},{"cell_type":"code","source":"n_steps = x_valid.shape[0] // BATCH_SIZE\ntrain_history_2 = model.fit(\n    valid_dataset.repeat(),\n    steps_per_epoch=n_steps,\n    epochs=EPOCHS*2\n)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Submission","metadata":{}},{"cell_type":"code","source":"sub['toxic'] = model.predict(test_dataset, verbose=1)\nsub.to_csv('submission.csv', index=False)","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}