{
  "id": 116153,
  "title": "Using/installing custom packages",
  "url": "/competitions/tensorflow2-question-answering/discussion/116153",
  "author_name": "",
  "post_date": "2019-11-07T09:56:04.463273100Z",
  "votes": 8,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Dear Kaggle team,\nI completely understand the reasons for a code competition and appreciate the effort to make kaggle fair, reproducable and comparable. \nThough these limitations can be quite frustrating. I now spent more time in getting my ML packages to run than actually developing and training my model. In <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/115046#662040\">a discussion</a> one of your team members stated:</p>\n\n<blockquote>\n  <p>you are allowed to use external data, input notebooks as a data source, or utility scripts. This means you can bring in huggingface with those methods, including training locally and then importing the model into your Kaggle notebook as an external data source</p>\n</blockquote>\n\n<p>So I went on a journey to try different methods, because copying one script (or the huggingface transformers folder) does not do the trick. These libraries depend themselves on other libraries, so we need to package all these libraries. </p>\n\n<ol>\n<li><p>I used \"pip download\" to get all dependencies. I zipped these dependecies and had to rename the zip extension, so the kaggle notebook would not hang during downloading the automatically unzipped folder. I still could not get these dependencies installed into my kernel with exotic errors popping up at various stages.</p></li>\n<li><p>I found somebody on kaggle stating \"When I want to do a kernel commit, I use a python package called stickytape to compile my codebase into a single script that can be copy and pasted into a script kernel.\". Compiling 500MB of dependencies into a single script sounds like a fun idea, I did not even try.</p></li>\n<li><p>There is a <a href=\"https://medium.com/@amimahloof/how-to-package-a-python-project-with-all-of-its-dependencies-for-offline-install-7eb240b27418\">medium article</a> about packaging python for offline install.... based upon python 2. I tried converting that code to py 3.6 but it failed silently. </p></li>\n<li><p>When copying whole folders (with my custom package) into the working directory, dashes inside filenames are removed, underscores are transformed to dashes (mhhkay?). I finally got my code to run through copying the site-packages folder (remember to rename the .zip extension, otherwise the kernel hangs) of my local virtualenv to my /kaggle/working dir but when I commit the notebook for submission it says: </p>\n\n<blockquote>\n  <p>Failure Message: Too many output files (max 500)</p>\n</blockquote></li>\n</ol>\n\n<p><strong>What?</strong></p>\n\n<p>This whole process is really frustrating. Could you please supply detailed explanations of how to package all needed dependencies (for example in the <a href=\"https://www.kaggle.com/docs/competitions#kernels-only-FAQ\">code competition FAQ</a>) for a python library like huggingface and how to get it to run on your kernels? Thanks </p>",
  "messages": [
    {
      "id": "667500",
      "postDate": "11/07/2019 09:56:04",
      "content": "<p>Dear Kaggle team,\nI completely understand the reasons for a code competition and appreciate the effort to make kaggle fair, reproducable and comparable. \nThough these limitations can be quite frustrating. I now spent more time in getting my ML packages to run than actually developing and training my model. In <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/115046#662040\">a discussion</a> one of your team members stated:</p>\n\n<blockquote>\n  <p>you are allowed to use external data, input notebooks as a data source, or utility scripts. This means you can bring in huggingface with those methods, including training locally and then importing the model into your Kaggle notebook as an external data source</p>\n</blockquote>\n\n<p>So I went on a journey to try different methods, because copying one script (or the huggingface transformers folder) does not do the trick. These libraries depend themselves on other libraries, so we need to package all these libraries. </p>\n\n<ol>\n<li><p>I used \"pip download\" to get all dependencies. I zipped these dependecies and had to rename the zip extension, so the kaggle notebook would not hang during downloading the automatically unzipped folder. I still could not get these dependencies installed into my kernel with exotic errors popping up at various stages.</p></li>\n<li><p>I found somebody on kaggle stating \"When I want to do a kernel commit, I use a python package called stickytape to compile my codebase into a single script that can be copy and pasted into a script kernel.\". Compiling 500MB of dependencies into a single script sounds like a fun idea, I did not even try.</p></li>\n<li><p>There is a <a href=\"https://medium.com/@amimahloof/how-to-package-a-python-project-with-all-of-its-dependencies-for-offline-install-7eb240b27418\">medium article</a> about packaging python for offline install.... based upon python 2. I tried converting that code to py 3.6 but it failed silently. </p></li>\n<li><p>When copying whole folders (with my custom package) into the working directory, dashes inside filenames are removed, underscores are transformed to dashes (mhhkay?). I finally got my code to run through copying the site-packages folder (remember to rename the .zip extension, otherwise the kernel hangs) of my local virtualenv to my /kaggle/working dir but when I commit the notebook for submission it says: </p>\n\n<blockquote>\n  <p>Failure Message: Too many output files (max 500)</p>\n</blockquote></li>\n</ol>\n\n<p><strong>What?</strong></p>\n\n<p>This whole process is really frustrating. Could you please supply detailed explanations of how to package all needed dependencies (for example in the <a href=\"https://www.kaggle.com/docs/competitions#kernels-only-FAQ\">code competition FAQ</a>) for a python library like huggingface and how to get it to run on your kernels? Thanks </p>",
      "rawMarkdown": "Dear Kaggle team,\nI completely understand the reasons for a code competition and appreciate the effort to make kaggle fair, reproducable and comparable. \nThough these limitations can be quite frustrating. I now spent more time in getting my ML packages to run than actually developing and training my model. In [a discussion](https://www.kaggle.com/c/tensorflow2-question-answering/discussion/115046#662040) one of your team members stated:\n&gt; you are allowed to use external data, input notebooks as a data source, or utility scripts. This means you can bring in huggingface with those methods, including training locally and then importing the model into your Kaggle notebook as an external data source\n\nSo I went on a journey to try different methods, because copying one script (or the huggingface transformers folder) does not do the trick. These libraries depend themselves on other libraries, so we need to package all these libraries. \n\n1. I used \"pip download\" to get all dependencies. I zipped these dependecies and had to rename the zip extension, so the kaggle notebook would not hang during downloading the automatically unzipped folder. I still could not get these dependencies installed into my kernel with exotic errors popping up at various stages.\n\n2.  I found somebody on kaggle stating \"When I want to do a kernel commit, I use a python package called stickytape to compile my codebase into a single script that can be copy and pasted into a script kernel.\". Compiling 500MB of dependencies into a single script sounds like a fun idea, I did not even try.\n\n3. There is a [medium article](https://medium.com/@amimahloof/how-to-package-a-python-project-with-all-of-its-dependencies-for-offline-install-7eb240b27418) about packaging python for offline install.... based upon python 2. I tried converting that code to py 3.6 but it failed silently. \n\n4. When copying whole folders (with my custom package) into the working directory, dashes inside filenames are removed, underscores are transformed to dashes (mhhkay?). I finally got my code to run through copying the site-packages folder (remember to rename the .zip extension, otherwise the kernel hangs) of my local virtualenv to my /kaggle/working dir but when I commit the notebook for submission it says: \n&gt; Failure Message: Too many output files (max 500)\n\n**What?**\n\nThis whole process is really frustrating. Could you please supply detailed explanations of how to package all needed dependencies (for example in the [code competition FAQ](https://www.kaggle.com/docs/competitions#kernels-only-FAQ)) for a python library like huggingface and how to get it to run on your kernels? Thanks",
      "votes": null
    },
    {
      "id": "671350",
      "postDate": "11/12/2019 15:01:40",
      "content": "<p>Really surprised that <a href=\"https://huggingface.co/transformers\">https://huggingface.co/transformers</a> is not allowed. Could someone clarify?</p>\n\n<p><a href=\"/timoeller\">@timoeller</a> Did you have trouble installing huggingface? I haven't tried it myself but our team plans to try out huggingface.</p>",
      "rawMarkdown": "Really surprised that https://huggingface.co/transformers is not allowed. Could someone clarify?\n\n@timoeller Did you have trouble installing huggingface? I haven't tried it myself but our team plans to try out huggingface.",
      "votes": null
    },
    {
      "id": "673154",
      "postDate": "11/14/2019 15:34:18",
      "content": "<p>Sorry for potential misunderstanding.\nHuggingface or other libraries are allowed. You just cannot install them through \"pip install transformers\" but have to add all dependencies manually. This manual adding of dependencies seems quite difficult.</p>",
      "rawMarkdown": "Sorry for potential misunderstanding.\nHuggingface or other libraries are allowed. You just cannot install them through \"pip install transformers\" but have to add all dependencies manually. This manual adding of dependencies seems quite difficult.",
      "votes": null
    },
    {
      "id": "673183",
      "postDate": "11/14/2019 16:15:10",
      "content": "<p>Thanks for clarifying.</p>",
      "rawMarkdown": "Thanks for clarifying.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 671350,
      "author_name": "arvindpdmn",
      "author_url": "",
      "post_date": "11/12/2019 15:01:40",
      "content": "<p>Really surprised that <a href=\"https://huggingface.co/transformers\">https://huggingface.co/transformers</a> is not allowed. Could someone clarify?</p>\n\n<p><a href=\"/timoeller\">@timoeller</a> Did you have trouble installing huggingface? I haven't tried it myself but our team plans to try out huggingface.</p>",
      "votes": null,
      "replies": [
        {
          "id": 673154,
          "author_name": "timoeller",
          "author_url": "",
          "post_date": "11/14/2019 15:34:18",
          "content": "<p>Sorry for potential misunderstanding.\nHuggingface or other libraries are allowed. You just cannot install them through \"pip install transformers\" but have to add all dependencies manually. This manual adding of dependencies seems quite difficult.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 673183,
          "author_name": "arvindpdmn",
          "author_url": "",
          "post_date": "11/14/2019 16:15:10",
          "content": "<p>Thanks for clarifying.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "667500": "Dear Kaggle team,\nI completely understand the reasons for a code competition and appreciate the effort to make kaggle fair, reproducable and comparable. \nThough these limitations can be quite frustrating. I now spent more time in getting my ML packages to run than actually developing and training my model. In [a discussion](https://www.kaggle.com/c/tensorflow2-question-answering/discussion/115046#662040) one of your team members stated:\n&gt; you are allowed to use external data, input notebooks as a data source, or utility scripts. This means you can bring in huggingface with those methods, including training locally and then importing the model into your Kaggle notebook as an external data source\n\nSo I went on a journey to try different methods, because copying one script (or the huggingface transformers folder) does not do the trick. These libraries depend themselves on other libraries, so we need to package all these libraries. \n\n1. I used \"pip download\" to get all dependencies. I zipped these dependecies and had to rename the zip extension, so the kaggle notebook would not hang during downloading the automatically unzipped folder. I still could not get these dependencies installed into my kernel with exotic errors popping up at various stages.\n\n2.  I found somebody on kaggle stating \"When I want to do a kernel commit, I use a python package called stickytape to compile my codebase into a single script that can be copy and pasted into a script kernel.\". Compiling 500MB of dependencies into a single script sounds like a fun idea, I did not even try.\n\n3. There is a [medium article](https://medium.com/@amimahloof/how-to-package-a-python-project-with-all-of-its-dependencies-for-offline-install-7eb240b27418) about packaging python for offline install.... based upon python 2. I tried converting that code to py 3.6 but it failed silently. \n\n4. When copying whole folders (with my custom package) into the working directory, dashes inside filenames are removed, underscores are transformed to dashes (mhhkay?). I finally got my code to run through copying the site-packages folder (remember to rename the .zip extension, otherwise the kernel hangs) of my local virtualenv to my /kaggle/working dir but when I commit the notebook for submission it says: \n&gt; Failure Message: Too many output files (max 500)\n\n**What?**\n\nThis whole process is really frustrating. Could you please supply detailed explanations of how to package all needed dependencies (for example in the [code competition FAQ](https://www.kaggle.com/docs/competitions#kernels-only-FAQ)) for a python library like huggingface and how to get it to run on your kernels? Thanks",
    "671350": "Really surprised that https://huggingface.co/transformers is not allowed. Could someone clarify?\n\n@timoeller Did you have trouble installing huggingface? I haven't tried it myself but our team plans to try out huggingface.",
    "673154": "Sorry for potential misunderstanding.\nHuggingface or other libraries are allowed. You just cannot install them through \"pip install transformers\" but have to add all dependencies manually. This manual adding of dependencies seems quite difficult.",
    "673183": "Thanks for clarifying."
  },
  "source": "meta"
}