{
  "id": 159616,
  "title": "[SOLVED] Reusing downloaded encodings",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/159616",
  "author_name": "",
  "post_date": "2020-06-18T03:43:51.901710200Z",
  "votes": null,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hi everyone, I just started working on this competition. As a starter, I am using <a href=\"https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta\">this notebook</a>. What I observed is that the notebook is downloading encoder and model for reach run. So I tried the following code to save model in *MODEL_LOC* directory.</p>\n\n<p>```python\nfrom pathlib import Path\nfrom transformers import TFAutoModel, AutoTokenizer\n...\ndef download_and_save_tokenizer(MODEL, MODEL_LOC):\n    tokenizer = AutoTokenizer.from_pretrained(MODEL)\n    Path(MODEL_LOC).mkdir(exist_ok=True, parents=True)\n    tokenizer.save_pretrained(MODEL_LOC)\n    return tokenizer</p>\n\n<p>MODEL = \"jplu/tf-xlm-roberta-large\"\nMODEL_LOC = \"../input/tf-xlm-roberta-large/\"\nif Path(MODEL_LOC).is_dir():\n    tokenizer = AutoTokenizer.from_pretrained(MODEL_LOC)\nelse:\n    tokenizer = download_and_save_tokenizer(MODEL, MODEL_LOC)\n...\n```</p>\n\n<p>But when I am trying this code, I am getting the following error:\n```python\nOSError: Can't load config for '../input/tf-xlm-roberta-large/'. Make sure that:</p>\n\n<ul>\n<li><p>'../input/tf-xlm-roberta-large/' is a correct model identifier listed on '<a href=\"https://huggingface.co/models\">https://huggingface.co/models</a>'</p></li>\n<li><p>or '../input/tf-xlm-roberta-large/' is the correct path to a directory containing a config.json file\n```</p></li>\n</ul>\n\n<p>The following files have been downloaded:\n1. sentencepiece.bpe.model\n2. special_tokens_map.json\n3. tokenizer_config.json </p>\n\n<p>but there is no config.json file. </p>\n\n<p>What am I doing wrong. Is it possible to load tokenizer from local saved directory? Any comments is appreciated. </p>\n\n<p>Thanks to <a href=\"/xhlulu\">@xhlulu</a> for <a href=\"https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta\">this notebook with good explanation</a>. </p>",
  "messages": [
    {
      "id": "891252",
      "postDate": "06/18/2020 03:43:51",
      "content": "<p>Hi everyone, I just started working on this competition. As a starter, I am using <a href=\"https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta\">this notebook</a>. What I observed is that the notebook is downloading encoder and model for reach run. So I tried the following code to save model in *MODEL_LOC* directory.</p>\n\n<p>```python\nfrom pathlib import Path\nfrom transformers import TFAutoModel, AutoTokenizer\n...\ndef download_and_save_tokenizer(MODEL, MODEL_LOC):\n    tokenizer = AutoTokenizer.from_pretrained(MODEL)\n    Path(MODEL_LOC).mkdir(exist_ok=True, parents=True)\n    tokenizer.save_pretrained(MODEL_LOC)\n    return tokenizer</p>\n\n<p>MODEL = \"jplu/tf-xlm-roberta-large\"\nMODEL_LOC = \"../input/tf-xlm-roberta-large/\"\nif Path(MODEL_LOC).is_dir():\n    tokenizer = AutoTokenizer.from_pretrained(MODEL_LOC)\nelse:\n    tokenizer = download_and_save_tokenizer(MODEL, MODEL_LOC)\n...\n```</p>\n\n<p>But when I am trying this code, I am getting the following error:\n```python\nOSError: Can't load config for '../input/tf-xlm-roberta-large/'. Make sure that:</p>\n\n<ul>\n<li><p>'../input/tf-xlm-roberta-large/' is a correct model identifier listed on '<a href=\"https://huggingface.co/models\">https://huggingface.co/models</a>'</p></li>\n<li><p>or '../input/tf-xlm-roberta-large/' is the correct path to a directory containing a config.json file\n```</p></li>\n</ul>\n\n<p>The following files have been downloaded:\n1. sentencepiece.bpe.model\n2. special_tokens_map.json\n3. tokenizer_config.json </p>\n\n<p>but there is no config.json file. </p>\n\n<p>What am I doing wrong. Is it possible to load tokenizer from local saved directory? Any comments is appreciated. </p>\n\n<p>Thanks to <a href=\"/xhlulu\">@xhlulu</a> for <a href=\"https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta\">this notebook with good explanation</a>. </p>",
      "rawMarkdown": "Hi everyone, I just started working on this competition. As a starter, I am using [this notebook](https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta). What I observed is that the notebook is downloading encoder and model for reach run. So I tried the following code to save model in *MODEL_LOC* directory.\n\n```python\nfrom pathlib import Path\nfrom transformers import TFAutoModel, AutoTokenizer\n...\ndef download_and_save_tokenizer(MODEL, MODEL_LOC):\n    tokenizer = AutoTokenizer.from_pretrained(MODEL)\n    Path(MODEL_LOC).mkdir(exist_ok=True, parents=True)\n    tokenizer.save_pretrained(MODEL_LOC)\n    return tokenizer\n\nMODEL = \"jplu/tf-xlm-roberta-large\"\nMODEL_LOC = \"../input/tf-xlm-roberta-large/\"\nif Path(MODEL_LOC).is_dir():\n    tokenizer = AutoTokenizer.from_pretrained(MODEL_LOC)\nelse:\n    tokenizer = download_and_save_tokenizer(MODEL, MODEL_LOC)\n...\n```\n\nBut when I am trying this code, I am getting the following error:\n```python\nOSError: Can't load config for '../input/tf-xlm-roberta-large/'. Make sure that:\n\n- '../input/tf-xlm-roberta-large/' is a correct model identifier listed on 'https://huggingface.co/models'\n\n- or '../input/tf-xlm-roberta-large/' is the correct path to a directory containing a config.json file\n```\n\nThe following files have been downloaded:\n1. sentencepiece.bpe.model\n2. special_tokens_map.json\n3. tokenizer_config.json \n\nbut there is no config.json file. \n\nWhat am I doing wrong. Is it possible to load tokenizer from local saved directory? Any comments is appreciated. \n\nThanks to @xhlulu for [this notebook with good explanation](https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta).",
      "votes": null
    },
    {
      "id": "891327",
      "postDate": "06/18/2020 05:17:21",
      "content": "<p>I was able to solve the problem. In order to use a downloaded tokenizer, we need to have config.json file. This file will only be created once we have downloaded the model. </p>\n\n<p>For those interested in using the downloaded tokenizer, you can use the following code. </p>\n\n<p>```python\nfrom pathlib import Path\nfrom transformers import TFAutoModel, AutoTokenizer</p>\n\n<p>def download_and_save_tokenizer(MODEL, MODEL_LOC):\n    tokenizer = AutoTokenizer.from_pretrained(MODEL)\n    Path(MODEL_LOC).mkdir(exist_ok=True, parents=True)\n    tokenizer.save_pretrained(MODEL_LOC)\n    return tokenizer</p>\n\n<p>def download_and_save_model(MODEL, MODEL_LOC):\n    transformer_layer = TFAutoModel.from_pretrained(MODEL)\n    Path(MODEL_LOC).mkdir(exist_ok=True, parents=True)\n    transformer_layer.save_pretrained(MODEL_LOC)\n    return transformer_layer</p>\n\n<p>MODEL = \"jplu/tf-xlm-roberta-large\"\nMODEL_LOC = \"../input/tfxlmrobertalarge/tf-xlm-roberta-large\"  # location of saved dir</p>\n\n<p>if Path(MODEL_LOC).is_dir():\n    try:\n        tokenizer = AutoTokenizer.from_pretrained(MODEL_LOC)\n        print(\"Tokenizer loaded from local dir\")\n    except:\n        tokenizer = download_and_save_tokenizer(MODEL, MODEL_LOC)\n        print(\"Tokenizer downloaded\")\nelse:\n    tokenizer = download_and_save_tokenizer(MODEL, MODEL_LOC)\n    print(\"Tokenizer downloaded\")\n...\nwith strategy.scope():\n    if Path(MODEL_LOC).is_dir():\n        try:\n            transformer_layer = TFAutoModel.from_pretrained(MODEL_LOC)\n            print(\"Model imported from local dir!\")\n        except:\n            transformer_layer = download_and_save_model(MODEL, MODEL_LOC)\n            print(\"Model downloaded!\")\n    else:\n        transformer_layer = download_and_save_model(MODEL, MODEL_LOC)\n        print(\"Model downloaded!\")\n    model = build_model(transformer_layer, maxlen=MAXLEN)\nmodel.summary()\n```</p>\n\n<p>This code worked for me. </p>",
      "rawMarkdown": "I was able to solve the problem. In order to use a downloaded tokenizer, we need to have config.json file. This file will only be created once we have downloaded the model. \n\nFor those interested in using the downloaded tokenizer, you can use the following code. \n\n```python\nfrom pathlib import Path\nfrom transformers import TFAutoModel, AutoTokenizer\n\ndef download_and_save_tokenizer(MODEL, MODEL_LOC):\n    tokenizer = AutoTokenizer.from_pretrained(MODEL)\n    Path(MODEL_LOC).mkdir(exist_ok=True, parents=True)\n    tokenizer.save_pretrained(MODEL_LOC)\n    return tokenizer\n\ndef download_and_save_model(MODEL, MODEL_LOC):\n    transformer_layer = TFAutoModel.from_pretrained(MODEL)\n    Path(MODEL_LOC).mkdir(exist_ok=True, parents=True)\n    transformer_layer.save_pretrained(MODEL_LOC)\n    return transformer_layer\n\nMODEL = \"jplu/tf-xlm-roberta-large\"\nMODEL_LOC = \"../input/tfxlmrobertalarge/tf-xlm-roberta-large\"  # location of saved dir\n\nif Path(MODEL_LOC).is_dir():\n    try:\n        tokenizer = AutoTokenizer.from_pretrained(MODEL_LOC)\n        print(\"Tokenizer loaded from local dir\")\n    except:\n        tokenizer = download_and_save_tokenizer(MODEL, MODEL_LOC)\n        print(\"Tokenizer downloaded\")\nelse:\n    tokenizer = download_and_save_tokenizer(MODEL, MODEL_LOC)\n    print(\"Tokenizer downloaded\")\n...\nwith strategy.scope():\n    if Path(MODEL_LOC).is_dir():\n        try:\n            transformer_layer = TFAutoModel.from_pretrained(MODEL_LOC)\n            print(\"Model imported from local dir!\")\n        except:\n            transformer_layer = download_and_save_model(MODEL, MODEL_LOC)\n            print(\"Model downloaded!\")\n    else:\n        transformer_layer = download_and_save_model(MODEL, MODEL_LOC)\n        print(\"Model downloaded!\")\n    model = build_model(transformer_layer, maxlen=MAXLEN)\nmodel.summary()\n```\n\nThis code worked for me.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 891327,
      "author_name": "manikanthr5",
      "author_url": "",
      "post_date": "06/18/2020 05:17:21",
      "content": "<p>I was able to solve the problem. In order to use a downloaded tokenizer, we need to have config.json file. This file will only be created once we have downloaded the model. </p>\n\n<p>For those interested in using the downloaded tokenizer, you can use the following code. </p>\n\n<p>```python\nfrom pathlib import Path\nfrom transformers import TFAutoModel, AutoTokenizer</p>\n\n<p>def download_and_save_tokenizer(MODEL, MODEL_LOC):\n    tokenizer = AutoTokenizer.from_pretrained(MODEL)\n    Path(MODEL_LOC).mkdir(exist_ok=True, parents=True)\n    tokenizer.save_pretrained(MODEL_LOC)\n    return tokenizer</p>\n\n<p>def download_and_save_model(MODEL, MODEL_LOC):\n    transformer_layer = TFAutoModel.from_pretrained(MODEL)\n    Path(MODEL_LOC).mkdir(exist_ok=True, parents=True)\n    transformer_layer.save_pretrained(MODEL_LOC)\n    return transformer_layer</p>\n\n<p>MODEL = \"jplu/tf-xlm-roberta-large\"\nMODEL_LOC = \"../input/tfxlmrobertalarge/tf-xlm-roberta-large\"  # location of saved dir</p>\n\n<p>if Path(MODEL_LOC).is_dir():\n    try:\n        tokenizer = AutoTokenizer.from_pretrained(MODEL_LOC)\n        print(\"Tokenizer loaded from local dir\")\n    except:\n        tokenizer = download_and_save_tokenizer(MODEL, MODEL_LOC)\n        print(\"Tokenizer downloaded\")\nelse:\n    tokenizer = download_and_save_tokenizer(MODEL, MODEL_LOC)\n    print(\"Tokenizer downloaded\")\n...\nwith strategy.scope():\n    if Path(MODEL_LOC).is_dir():\n        try:\n            transformer_layer = TFAutoModel.from_pretrained(MODEL_LOC)\n            print(\"Model imported from local dir!\")\n        except:\n            transformer_layer = download_and_save_model(MODEL, MODEL_LOC)\n            print(\"Model downloaded!\")\n    else:\n        transformer_layer = download_and_save_model(MODEL, MODEL_LOC)\n        print(\"Model downloaded!\")\n    model = build_model(transformer_layer, maxlen=MAXLEN)\nmodel.summary()\n```</p>\n\n<p>This code worked for me. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "891252": "Hi everyone, I just started working on this competition. As a starter, I am using [this notebook](https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta). What I observed is that the notebook is downloading encoder and model for reach run. So I tried the following code to save model in *MODEL_LOC* directory.\n\n```python\nfrom pathlib import Path\nfrom transformers import TFAutoModel, AutoTokenizer\n...\ndef download_and_save_tokenizer(MODEL, MODEL_LOC):\n    tokenizer = AutoTokenizer.from_pretrained(MODEL)\n    Path(MODEL_LOC).mkdir(exist_ok=True, parents=True)\n    tokenizer.save_pretrained(MODEL_LOC)\n    return tokenizer\n\nMODEL = \"jplu/tf-xlm-roberta-large\"\nMODEL_LOC = \"../input/tf-xlm-roberta-large/\"\nif Path(MODEL_LOC).is_dir():\n    tokenizer = AutoTokenizer.from_pretrained(MODEL_LOC)\nelse:\n    tokenizer = download_and_save_tokenizer(MODEL, MODEL_LOC)\n...\n```\n\nBut when I am trying this code, I am getting the following error:\n```python\nOSError: Can't load config for '../input/tf-xlm-roberta-large/'. Make sure that:\n\n- '../input/tf-xlm-roberta-large/' is a correct model identifier listed on 'https://huggingface.co/models'\n\n- or '../input/tf-xlm-roberta-large/' is the correct path to a directory containing a config.json file\n```\n\nThe following files have been downloaded:\n1. sentencepiece.bpe.model\n2. special_tokens_map.json\n3. tokenizer_config.json \n\nbut there is no config.json file. \n\nWhat am I doing wrong. Is it possible to load tokenizer from local saved directory? Any comments is appreciated. \n\nThanks to @xhlulu for [this notebook with good explanation](https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta).",
    "891327": "I was able to solve the problem. In order to use a downloaded tokenizer, we need to have config.json file. This file will only be created once we have downloaded the model. \n\nFor those interested in using the downloaded tokenizer, you can use the following code. \n\n```python\nfrom pathlib import Path\nfrom transformers import TFAutoModel, AutoTokenizer\n\ndef download_and_save_tokenizer(MODEL, MODEL_LOC):\n    tokenizer = AutoTokenizer.from_pretrained(MODEL)\n    Path(MODEL_LOC).mkdir(exist_ok=True, parents=True)\n    tokenizer.save_pretrained(MODEL_LOC)\n    return tokenizer\n\ndef download_and_save_model(MODEL, MODEL_LOC):\n    transformer_layer = TFAutoModel.from_pretrained(MODEL)\n    Path(MODEL_LOC).mkdir(exist_ok=True, parents=True)\n    transformer_layer.save_pretrained(MODEL_LOC)\n    return transformer_layer\n\nMODEL = \"jplu/tf-xlm-roberta-large\"\nMODEL_LOC = \"../input/tfxlmrobertalarge/tf-xlm-roberta-large\"  # location of saved dir\n\nif Path(MODEL_LOC).is_dir():\n    try:\n        tokenizer = AutoTokenizer.from_pretrained(MODEL_LOC)\n        print(\"Tokenizer loaded from local dir\")\n    except:\n        tokenizer = download_and_save_tokenizer(MODEL, MODEL_LOC)\n        print(\"Tokenizer downloaded\")\nelse:\n    tokenizer = download_and_save_tokenizer(MODEL, MODEL_LOC)\n    print(\"Tokenizer downloaded\")\n...\nwith strategy.scope():\n    if Path(MODEL_LOC).is_dir():\n        try:\n            transformer_layer = TFAutoModel.from_pretrained(MODEL_LOC)\n            print(\"Model imported from local dir!\")\n        except:\n            transformer_layer = download_and_save_model(MODEL, MODEL_LOC)\n            print(\"Model downloaded!\")\n    else:\n        transformer_layer = download_and_save_model(MODEL, MODEL_LOC)\n        print(\"Model downloaded!\")\n    model = build_model(transformer_layer, maxlen=MAXLEN)\nmodel.summary()\n```\n\nThis code worked for me."
  },
  "source": "meta"
}