{"cells":[{"metadata":{"_uuid":"ebd8119d-fb40-454f-a46f-940e80e7f8ce","_cell_guid":"f2ef93cd-2bc2-4681-bc6e-9a9395b25691","trusted":true},"cell_type":"markdown","source":"# Overview\n\nIt only takes one toxic comment to sour an online discussion. The Conversation AI team, a research initiative founded by [Jigsaw](https://jigsaw.google.com/) and Google, builds technology to protect voices in conversation. A main area of focus is machine learning models that can identify toxicity in online conversations, where toxicity is defined as anything *rude, disrespectful or otherwise likely to make someone leave a discussion*. Our API, [Perspective](http://perspectiveapi.com/), serves these models and others in a growing set of languages (see our [documentation](https://github.com/conversationai/perspectiveapi/blob/master/2-api/models.md#all-attribute-types) for the full list). If these toxic contributions can be identified, we could have a safer, more collaborative internet.\n\nIn this competition, we'll explore how models for recognizing toxicity in online conversations might generalize across different languages. Specifically, in this notebook, we'll demonstrate this with a multilingual BERT (m-BERT) model. Multilingual BERT is pretrained on monolingual data in a variety of languages, and through this learns multilingual representations of text. These multilingual representations enable *zero-shot cross-lingual transfer*, that is, by fine-tuning on a task in one language, m-BERT can learn to perform that same task in another language (for some examples, see e.g. [How multilingual is Multilingual BERT?](https://arxiv.org/abs/1906.01502)).\n\nWe'll study this zero-shot transfer in the context of toxicity in online conversations, similar to past competitions we've hosted ([[1]](https://www.kaggle.com/c/jigsaw-unintended-bias-in-toxicity-classification), [[2]](https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge)). But rather than analyzing toxicity in English as in those competitions, here we'll ask you to do it in several different languages. For training, we're including the (English) datasets from our earlier competitions, as well as a small amount of new toxicity data in other languages."},{"metadata":{"_uuid":"cc13af9e-5629-4bcc-8906-5875bcc3b62a","_cell_guid":"3bb3b650-48ba-4dbe-a23d-55740b04b6f2","trusted":true},"cell_type":"markdown","source":"# Benchmark notebook\n\nWe'll begin by importing TensorFlow, our datasets, and TensorFlow Hub, which has a pretrained multilingual model we'll use."},{"metadata":{"_uuid":"cba509cb-708d-4fd0-8e2a-e5c1c7a0982d","_cell_guid":"593619be-8090-4dad-a462-e883e560ec1c","trusted":true},"cell_type":"code","source":"import os, time\nimport pandas\nimport tensorflow as tf\nimport tensorflow_hub as hub\nfrom kaggle_datasets import KaggleDatasets\nprint(tf.version.VERSION)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"add2f478-e634-42db-b3c6-732ebd484ce3","_cell_guid":"ae9fd8d4-bb00-4b3d-8452-0ec605f6ebba","trusted":true},"cell_type":"markdown","source":"Detect TPUs or GPUs:"},{"metadata":{"_uuid":"bc97f110-17eb-44dd-a792-d66c27a0b3a6","_cell_guid":"f5ccaf08-c532-4fde-9306-b897c890d0f8","trusted":true},"cell_type":"code","source":"# Detect hardware, return appropriate distribution strategy\ntry:\n    # TPU detection. No parameters necessary if TPU_NAME environment variable is\n    # set: this is always the case on Kaggle.\n    tpu = tf.distribute.cluster_resolver.TPUClusterResolver()\n    print('Running on TPU ', tpu.master())\nexcept ValueError:\n    tpu = None\n\nif tpu:\n    tf.config.experimental_connect_to_cluster(tpu)\n    tf.tpu.experimental.initialize_tpu_system(tpu)\n    strategy = tf.distribute.experimental.TPUStrategy(tpu)\nelse:\n    # Default distribution strategy in Tensorflow. Works on CPU and single GPU.\n    strategy = tf.distribute.get_strategy()\n\nprint(\"REPLICAS: \", strategy.num_replicas_in_sync)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"b1188d07-d23d-4d04-b9f5-125d16b982ed","_cell_guid":"c55830a8-39e7-4607-9691-cd0848a092f7","trusted":true},"cell_type":"markdown","source":"Set maximum sequence length and path variables."},{"metadata":{"_uuid":"7d8d39c7-1fd2-4af7-a88a-9ca1242a6277","_cell_guid":"7acc953d-b1f6-4d29-a006-69af6839f7a9","trusted":true},"cell_type":"code","source":"SEQUENCE_LENGTH = 128\n\n# Note that private datasets cannot be copied - you'll have to share any pretrained models \n# you want to use with other competitors!\nGCS_PATH = KaggleDatasets().get_gcs_path('jigsaw-multilingual-toxic-comment-classification')\nBERT_GCS_PATH = KaggleDatasets().get_gcs_path('bert-multi')\nBERT_GCS_PATH_SAVEDMODEL = BERT_GCS_PATH + \"/bert_multi_from_tfhub\"","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"85bb06d9-cd83-41a8-a72f-e25c0b336d82","_cell_guid":"b4c5c163-88f3-4021-9478-4005a360e7ce","trusted":true},"cell_type":"markdown","source":"Define the model. We convert m-BERT's output to a final probabilty estimate. We're using an [m-BERT model from TensorFlow Hub](https://tfhub.dev/tensorflow/bert_multi_cased_L-12_H-768_A-12/1)."},{"metadata":{"_uuid":"1392a5e0-c8e4-46ea-b45d-0d9289682e09","_cell_guid":"fcd2093b-0774-4adf-9e4c-e9096f156d32","trusted":true},"cell_type":"code","source":"def multilingual_bert_model(max_seq_length=SEQUENCE_LENGTH, trainable_bert=True):\n    \"\"\"Build and return a multilingual BERT model and tokenizer.\"\"\"\n    input_word_ids = tf.keras.layers.Input(\n        shape=(max_seq_length,), dtype=tf.int32, name=\"input_word_ids\")\n    input_mask = tf.keras.layers.Input(\n        shape=(max_seq_length,), dtype=tf.int32, name=\"input_mask\")\n    segment_ids = tf.keras.layers.Input(\n        shape=(max_seq_length,), dtype=tf.int32, name=\"all_segment_id\")\n    \n    # Load a SavedModel on TPU from GCS. This model is available online at \n    # https://tfhub.dev/tensorflow/bert_multi_cased_L-12_H-768_A-12/1. You can use your own \n    # pretrained models, but will need to add them as a Kaggle dataset.\n    bert_layer = tf.saved_model.load(BERT_GCS_PATH_SAVEDMODEL)\n    # Cast the loaded model to a TFHub KerasLayer.\n    bert_layer = hub.KerasLayer(bert_layer, trainable=trainable_bert)\n\n    pooled_output, _ = bert_layer([input_word_ids, input_mask, segment_ids])\n    output = tf.keras.layers.Dense(32, activation='relu')(pooled_output)\n    output = tf.keras.layers.Dense(1, activation='sigmoid', name='labels')(output)\n\n    return tf.keras.Model(inputs={'input_word_ids': input_word_ids,\n                                  'input_mask': input_mask,\n                                  'all_segment_id': segment_ids},\n                          outputs=output)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"b24c47ad-a156-41f2-97b8-f5618181382c","_cell_guid":"48479a32-25c1-40c9-bd1c-d076eb39d86e","trusted":true},"cell_type":"markdown","source":"Load the preprocessed dataset. See the demo notebook for sample code for performing this preprocessing."},{"metadata":{"trusted":true},"cell_type":"code","source":"def parse_string_list_into_ints(strlist):\n    s = tf.strings.strip(strlist)\n    s = tf.strings.substr(\n        strlist, 1, tf.strings.length(s) - 2)  # Remove parentheses around list\n    s = tf.strings.split(s, ',', maxsplit=SEQUENCE_LENGTH)\n    s = tf.strings.to_number(s, tf.int32)\n    s = tf.reshape(s, [SEQUENCE_LENGTH])  # Force shape here needed for XLA compilation (TPU)\n    return s\n\ndef format_sentences(data, label='toxic', remove_language=False):\n    labels = {'labels': data.pop(label)}\n    if remove_language:\n        languages = {'language': data.pop('lang')}\n    # The remaining three items in the dict parsed from the CSV are lists of integers\n    for k,v in data.items():  # \"input_word_ids\", \"input_mask\", \"all_segment_id\"\n        data[k] = parse_string_list_into_ints(v)\n    return data, labels\n\ndef make_sentence_dataset_from_csv(filename, label='toxic', language_to_filter=None):\n    # This assumes the column order label, input_word_ids, input_mask, segment_ids\n    SELECTED_COLUMNS = [label, \"input_word_ids\", \"input_mask\", \"all_segment_id\"]\n    label_default = tf.int32 if label == 'id' else tf.float32\n    COLUMN_DEFAULTS  = [label_default, tf.string, tf.string, tf.string]\n\n    if language_to_filter:\n        insert_pos = 0 if label != 'id' else 1\n        SELECTED_COLUMNS.insert(insert_pos, 'lang')\n        COLUMN_DEFAULTS.insert(insert_pos, tf.string)\n\n    preprocessed_sentences_dataset = tf.data.experimental.make_csv_dataset(\n        filename, column_defaults=COLUMN_DEFAULTS, select_columns=SELECTED_COLUMNS,\n        batch_size=1, num_epochs=1, shuffle=False)  # We'll do repeating and shuffling ourselves\n    # make_csv_dataset required a batch size, but we want to batch later\n    preprocessed_sentences_dataset = preprocessed_sentences_dataset.unbatch()\n    \n    if language_to_filter:\n        preprocessed_sentences_dataset = preprocessed_sentences_dataset.filter(\n            lambda data: tf.math.equal(data['lang'], tf.constant(language_to_filter)))\n        #preprocessed_sentences.pop('lang')\n    preprocessed_sentences_dataset = preprocessed_sentences_dataset.map(\n        lambda data: format_sentences(data, label=label,\n                                      remove_language=language_to_filter))\n\n    return preprocessed_sentences_dataset","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"10ff216a-2248-4104-858c-2b2461b42fba","_cell_guid":"e006f30f-6350-45d5-b452-338c4bc78cc5","trusted":true},"cell_type":"markdown","source":"Set up our data pipelines for training and evaluation."},{"metadata":{"_uuid":"864f31ac-8285-442b-93cf-dcae11d7fe62","_cell_guid":"b17b24d7-1b0f-442b-bfd2-c42748bcd067","trusted":true},"cell_type":"code","source":"def make_dataset_pipeline(dataset, repeat_and_shuffle=True):\n    \"\"\"Set up the pipeline for the given dataset.\n    \n    Caches, repeats, shuffles, and sets the pipeline up to prefetch batches.\"\"\"\n    cached_dataset = dataset.cache()\n    if repeat_and_shuffle:\n        cached_dataset = cached_dataset.repeat().shuffle(2048)\n    cached_dataset = cached_dataset.batch(32 * strategy.num_replicas_in_sync)\n    cached_dataset = cached_dataset.prefetch(tf.data.experimental.AUTOTUNE)\n    return cached_dataset\n\n# Load the preprocessed English dataframe.\npreprocessed_en_filename = (\n    GCS_PATH + \"/jigsaw-toxic-comment-train-processed-seqlen{}.csv\".format(\n        SEQUENCE_LENGTH))\n\n# Set up the dataset and pipeline.\nenglish_train_dataset = make_dataset_pipeline(\n    make_sentence_dataset_from_csv(preprocessed_en_filename))\n\n# Process the new datasets by language.\npreprocessed_val_filename = (\n    GCS_PATH + \"/validation-processed-seqlen{}.csv\".format(SEQUENCE_LENGTH))\n\nnonenglish_val_datasets = {}\nfor language_name, language_label in [('Spanish', 'es'), ('Italian', 'it'),\n                                      ('Turkish', 'tr')]:\n    nonenglish_val_datasets[language_name] = make_sentence_dataset_from_csv(\n        preprocessed_val_filename, language_to_filter=language_label)\n    nonenglish_val_datasets[language_name] = make_dataset_pipeline(\n        nonenglish_val_datasets[language_name])\n\nnonenglish_val_datasets['Combined'] = tf.data.experimental.sample_from_datasets(\n        (nonenglish_val_datasets['Spanish'], nonenglish_val_datasets['Italian'],\n         nonenglish_val_datasets['Turkish']))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"0779be08-0502-47c0-b284-5f3d851de2e1","_cell_guid":"ffd5c9ef-a806-4ae6-a1c5-8723ed822232","trusted":true},"cell_type":"markdown","source":"Compile our model. We'll first evaluate it on our new toxicity dataset in the \ndifferent languages to see its performance. After that, we'll train it on one of our English datasets, and then again evaluate its performance on the new multilingual toxicity data. As our metric, we'll use the [AUC](https://www.tensorflow.org/api_docs/python/tf/keras/metrics/AUC)."},{"metadata":{"_uuid":"e3d569ca-0bf4-4bde-aa69-95a46908f65a","_cell_guid":"422a984e-e571-4898-9667-b95d38416ddd","trusted":true},"cell_type":"code","source":"with strategy.scope():\n    multilingual_bert = multilingual_bert_model()\n\n    # Compile the model. Optimize using stochastic gradient descent.\n    multilingual_bert.compile(\n        loss=tf.keras.losses.BinaryCrossentropy(),\n        optimizer=tf.keras.optimizers.SGD(learning_rate=0.001),\n        metrics=[tf.keras.metrics.AUC()])\n\nmultilingual_bert.summary()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"730e2b4b-6e7c-43f3-912e-3cbea3db62c2","_cell_guid":"f477d7f4-20ec-4858-87d6-c041313ad276","trusted":true},"cell_type":"code","source":"# Test the model's performance on non-English comments before training.\nfor language in nonenglish_val_datasets:\n    results = multilingual_bert.evaluate(nonenglish_val_datasets[language],\n                                         steps=100, verbose=0)\n    print('{} loss, AUC before training:'.format(language), results)\n\nresults = multilingual_bert.evaluate(english_train_dataset,\n                                     steps=100, verbose=0)\nprint('\\nEnglish loss, AUC before training:', results)\n\nprint()\n# Train on English Wikipedia comment data.\nhistory = multilingual_bert.fit(\n    # Set steps such that the number of examples per epoch is fixed.\n    # This makes training on different accelerators more comparable.\n    english_train_dataset, steps_per_epoch=4000/strategy.num_replicas_in_sync,\n    epochs=5, verbose=1, validation_data=nonenglish_val_datasets['Combined'],\n    validation_steps=100)\nprint()\n\n# Re-evaluate the model's performance on non-English comments after training.\nfor language in nonenglish_val_datasets:\n    results = multilingual_bert.evaluate(nonenglish_val_datasets[language],\n                                         steps=100, verbose=0)\n    print('{} loss, AUC after training:'.format(language), results)\n\nresults = multilingual_bert.evaluate(english_train_dataset,\n                                     steps=100, verbose=0)\nprint('\\nEnglish loss, AUC after training:', results)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Generate predictions\n\nFinally, we'll use our trained multilingual model to generate predictions for the test data."},{"metadata":{"trusted":true},"cell_type":"code","source":"import numpy as np\n\nTEST_DATASET_SIZE = 63812\n\nprint('Making dataset...')\npreprocessed_test_filename = (\n    GCS_PATH + \"/test-processed-seqlen{}.csv\".format(SEQUENCE_LENGTH))\ntest_dataset = make_sentence_dataset_from_csv(preprocessed_test_filename, label='id')\ntest_dataset = make_dataset_pipeline(test_dataset, repeat_and_shuffle=False)\n\nprint('Computing predictions...')\ntest_sentences_dataset = test_dataset.map(lambda sentence, idnum: sentence)\nprobabilities = np.squeeze(multilingual_bert.predict(test_sentences_dataset))\nprint(probabilities)\n\nprint('Generating submission file...')\ntest_ids_dataset = test_dataset.map(lambda sentence, idnum: idnum).unbatch()\ntest_ids = next(iter(test_ids_dataset.batch(TEST_DATASET_SIZE)))[\n    'labels'].numpy().astype('U')  # All in one batch\n\nnp.savetxt('submission.csv', np.rec.fromarrays([test_ids, probabilities]),\n           fmt=['%s', '%f'], delimiter=',', header='id,toxic', comments='')\n!head submission.csv","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"1b4ae146-474c-4abb-a0a3-1ab08f7c6fe4","_cell_guid":"7f245cf7-e58e-437d-bcfc-49be54cce416","trusted":true},"cell_type":"markdown","source":"# Future thoughts\n\nOur results show improvements in recognizing toxicity in every language, even though we only trained in English. But this notebook is only as a basic demonstration of cross-lingual model transfer, and how it applies to toxicity in online comments — we've left a lot for you to explore in the competition! Here are a few starter thoughts for how to improve upon what we've done:\n1.   How might toxicity vary by language? The demo illustrates that some properties are shared, i.e. we do learn toxicity in other languages from toxicity in English. But our final AUC values vary significantly by language. What factors might contribute to this? Are there ways you could use this?\n2.   The scores in the `jigsaw-unintended-bias-train` dataset reflect the *probability* that an individual human rater will find a comment toxic (see [our blog post](https://medium.com/the-false-positive/creating-labeled-datasets-and-exploring-the-role-of-human-raters-56367b6db298) for more). However, for the dataset `jigsaw-toxic-comment-train` and for the non-English data, we only have binary labels. What information is there in the non-binary scores, i.e. in terms of rater agreement and disagreement?\n3.   There are a lot of things you could try to improve upon the demo (different models, maximum sequence lengths, optimizers, learning rates and learning rate schedules, etc). You could also include the dataset from [our first competition](https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/data) (which we've included, and preprocessed, in this one) as part of your training and see what improvements that might lead to, or consider using information from some of the other labels in the datasets.\n\nWe're excited to see the ways you can implement or expand upon these ideas, as well as the new ones you'll come up with!"}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}