{
  "id": 121695,
  "title": "First steps to get started",
  "url": "/competitions/deepfake-detection-challenge/discussion/121695",
  "author_name": "Carlos Souza",
  "post_date": "2019-12-14T21:33:58.618000",
  "votes": 95,
  "comment_count": 39,
  "views": 0,
  "content": "<p>Here are the steps I followed to get started:</p>\n\n<p><strong>1) Setup a VM instance</strong>\n- As I don't have local GPUs at my disposal, this was a must;\n- I sent an email to AWS requesting the $1000 in credits. While they don't answer, I got started in Google Cloud, which have cheaper &amp; better GPUs/TPUs;\n- While setting up the VM, I've chosen 1TB SSD disk storage: less than that would not allow me download and unzip the complete 472 GB full dataset at once;</p>\n\n<p><strong>2) Download the full dataset</strong>\n- Downloading the full dataset was surprisingly fast: it was constantly &gt;250 MB/s;\n- I did it using <em>wget</em>:\n<code>\nwget --load-cookies cookies.txt https://www.kaggle.com/c/16880/datadownload/dfdc_train_all.zip\n</code>\n- I had to save the cookies from my local machine with Kaggle's credentials in the <em>cookies.txt</em> in the VM. Without that, it won't work;</p>\n\n<p><strong>3) Unzip everything</strong>\n- The file *dfdc_train_all.zip* is like a Russian doll: it contains 49 zip files;\n- After unzipping and deleting it, I applied the following script to unzip each of the 49 files:\n```python\nimport os\nfrom tqdm import tqdm\nfrom zipfile import ZipFile</p>\n\n<p>def get_zipfiles(directory):\n    list_files = []\n    for filename in os.listdir(home):\n        if filename.endswith(\".zip\"): \n            list_files.append(os.path.join(home, filename))\n    return list_files</p>\n\n<p>zipfiles = get_zipfiles('/home/carlossouza')\nzipfiles.sort()</p>\n\n<p>for zipfile in zipfiles:\n    print(f'Extracting {zipfile}...')\n    with ZipFile(file=zipfile) as zip_file:\n        for file in tqdm(iterable=zip_file.namelist(), total=len(zip_file.namelist())):\n            zip_file.extract(member=file)</p>\n\n<pre><code>os.remove(zipfile)\n</code></pre>\n\n<p>```</p>\n\n<p>If you used a different/better procedure to get started, I'd love to hear and learn ;)\nCheers\nCarlos</p>",
  "messages": [
    {
      "id": 695263,
      "postDate": "2019-12-14T21:33:58.620Z",
      "content": "<p>Here are the steps I followed to get started:</p>\n\n<p><strong>1) Setup a VM instance</strong>\n- As I don't have local GPUs at my disposal, this was a must;\n- I sent an email to AWS requesting the $1000 in credits. While they don't answer, I got started in Google Cloud, which have cheaper &amp; better GPUs/TPUs;\n- While setting up the VM, I've chosen 1TB SSD disk storage: less than that would not allow me download and unzip the complete 472 GB full dataset at once;</p>\n\n<p><strong>2) Download the full dataset</strong>\n- Downloading the full dataset was surprisingly fast: it was constantly &gt;250 MB/s;\n- I did it using <em>wget</em>:\n<code>\nwget --load-cookies cookies.txt https://www.kaggle.com/c/16880/datadownload/dfdc_train_all.zip\n</code>\n- I had to save the cookies from my local machine with Kaggle's credentials in the <em>cookies.txt</em> in the VM. Without that, it won't work;</p>\n\n<p><strong>3) Unzip everything</strong>\n- The file *dfdc_train_all.zip* is like a Russian doll: it contains 49 zip files;\n- After unzipping and deleting it, I applied the following script to unzip each of the 49 files:\n```python\nimport os\nfrom tqdm import tqdm\nfrom zipfile import ZipFile</p>\n\n<p>def get_zipfiles(directory):\n    list_files = []\n    for filename in os.listdir(home):\n        if filename.endswith(\".zip\"): \n            list_files.append(os.path.join(home, filename))\n    return list_files</p>\n\n<p>zipfiles = get_zipfiles('/home/carlossouza')\nzipfiles.sort()</p>\n\n<p>for zipfile in zipfiles:\n    print(f'Extracting {zipfile}...')\n    with ZipFile(file=zipfile) as zip_file:\n        for file in tqdm(iterable=zip_file.namelist(), total=len(zip_file.namelist())):\n            zip_file.extract(member=file)</p>\n\n<pre><code>os.remove(zipfile)\n</code></pre>\n\n<p>```</p>\n\n<p>If you used a different/better procedure to get started, I'd love to hear and learn ;)\nCheers\nCarlos</p>",
      "rawMarkdown": "Here are the steps I followed to get started:\n\n**1) Setup a VM instance**\n- As I don't have local GPUs at my disposal, this was a must;\n- I sent an email to AWS requesting the $1000 in credits. While they don't answer, I got started in Google Cloud, which have cheaper &amp; better GPUs/TPUs;\n- While setting up the VM, I've chosen 1TB SSD disk storage: less than that would not allow me download and unzip the complete 472 GB full dataset at once;\n\n\n**2) Download the full dataset**\n- Downloading the full dataset was surprisingly fast: it was constantly &gt;250 MB/s;\n- I did it using *wget*:\n```\nwget --load-cookies cookies.txt https://www.kaggle.com/c/16880/datadownload/dfdc_train_all.zip\n```\n- I had to save the cookies from my local machine with Kaggle's credentials in the *cookies.txt* in the VM. Without that, it won't work;\n\n\n**3) Unzip everything**\n- The file *dfdc_train_all.zip* is like a Russian doll: it contains 49 zip files;\n- After unzipping and deleting it, I applied the following script to unzip each of the 49 files:\n```python\nimport os\nfrom tqdm import tqdm\nfrom zipfile import ZipFile\n\ndef get_zipfiles(directory):\n    list_files = []\n    for filename in os.listdir(home):\n        if filename.endswith(\".zip\"): \n            list_files.append(os.path.join(home, filename))\n    return list_files\n\nzipfiles = get_zipfiles('/home/carlossouza')\nzipfiles.sort()\n\nfor zipfile in zipfiles:\n    print(f'Extracting {zipfile}...')\n    with ZipFile(file=zipfile) as zip_file:\n        for file in tqdm(iterable=zip_file.namelist(), total=len(zip_file.namelist())):\n            zip_file.extract(member=file)\n    \n    os.remove(zipfile)\n```\n\nIf you used a different/better procedure to get started, I'd love to hear and learn ;)\nCheers\nCarlos",
      "votes": 94
    },
    {
      "id": 701600,
      "postDate": "2019-12-23T16:42:50.660Z",
      "content": "<p>My alternative to step 3 delineated by Carlos Above:</p>\n\n<p>&gt; 3) Unzip everything</p>\n\n<pre><code>The file dfdctrainall.zip is like a Russian doll: it contains 49 zip files;\nAfter unzipping and deleting it, I applied the following script to unzip each of the 49 files:\n</code></pre>\n\n<p>Since unzip is a disk/processor bottleneck the best way to use concurrency is to use all processors. ( i.e. Multiprocessing and not Multithreading )</p>\n\n<p>```python\nimport multiprocessing <br>\nfrom pathlib import Path <br>\nfrom time import time <br>\nfrom zipfile import ZipFile <br>\nimport logging\nfrom typing import Union                                                                                                                                                    </p>\n\n<p>DATA = Path(\"path_to_the_49_zipped files\") <br>\nlogging.basicConfig(filename=\"extract.log\", level=logging.INFO)\nzipfiles = sorted(list(DATA.glob(\"*<em>/</em>.zip\")), key=lambda x: x.stem)</p>\n\n<p>def extract_zip(zipfile: Union[str, Path])-&gt;None:\n    with ZipFile(zipfile) as zip_file:\n        for file in zip_file.namelist():\n            zip_file.extract(member=file, path=zipfile.parent)\n    logging.info(f'Finished extracting {zipfile.stem}')</p>\n\n<p>start = time()\nwith multiprocessing.Pool() as pool: # use all cores available\n    pool.map(extract_zip, zipfiles)</p>\n\n<p>logging.info(f\"Extracted all zip files in {time() - start} seconds!\")\n```</p>\n\n<p>I didn't use <code>tqdm</code> because in multiprocessing it would show 8 tqdm results at the same time so it didn't make sense.</p>\n\n<p>This created in <code>DATA</code> directory 49 folders. In order to work with all the files in a same folder I just created a folder with symlinks to each video file</p>\n\n<p>```python\nDATA = Path(\"path_to_the_49_zipped files\")\nTARGET = DATA.parent / \"full\"</p>\n\n<p>def path_walk(top: Union[str,Path], topdown: bool = False, followlinks:bool = False):\n    \"\"\"See Python docs for os.walk, exact same behavior but it yields Path() instances instead\n    \"\"\"</p>\n\n<pre><code>dirs = (node for node in top.iterdir() if node.is_dir())\nnondirs =(node for node in top.iterdir() if not node.is_dir())\n\nif topdown:\n    yield top, dirs, nondirs\n\nfor name in dirs:\n    if followlinks or name.is_symlink() is False:\n        for x in path_walk(name, topdown, followlinks):\n            yield x\n\nif topdown is not True:\n    yield top, dirs, nondirs    \n</code></pre>\n\n<h1>flatten the list of lists</h1>\n\n<p>files = (f for top, dirs, files  in path_walk(DATA) for f in files if f.suffix == \".mp4\") \nfor file in files:\n    (TARGET / file.name).symlink_to(file)\n<code>``\nSo in the</code>TARGET` I have all the <em>119146</em> video files to work on</p>\n\n<p>I hope this helps!</p>\n\n<p>Cheers</p>",
      "rawMarkdown": "My alternative to step 3 delineated by Carlos Above:\n\n&gt; 3) Unzip everything\n\n    The file dfdctrainall.zip is like a Russian doll: it contains 49 zip files;\n    After unzipping and deleting it, I applied the following script to unzip each of the 49 files:\n\nSince unzip is a disk/processor bottleneck the best way to use concurrency is to use all processors. ( i.e. Multiprocessing and not Multithreading )\n\n```python\nimport multiprocessing                                                                                                                                            \nfrom pathlib import Path                                                                                                                                          \nfrom time import time                                                                                                                                             \nfrom zipfile import ZipFile                                                                                                                                       \nimport logging\nfrom typing import Union                                                                                                                                                    \n                                                                                                                                                                   \nDATA = Path(\"path_to_the_49_zipped files\")                                                                                                    \nlogging.basicConfig(filename=\"extract.log\", level=logging.INFO)\nzipfiles = sorted(list(DATA.glob(\"**/*.zip\")), key=lambda x: x.stem)\n\ndef extract_zip(zipfile: Union[str, Path])-&gt;None:\n    with ZipFile(zipfile) as zip_file:\n        for file in zip_file.namelist():\n            zip_file.extract(member=file, path=zipfile.parent)\n    logging.info(f'Finished extracting {zipfile.stem}')\n\n\nstart = time()\nwith multiprocessing.Pool() as pool: # use all cores available\n    pool.map(extract_zip, zipfiles)\n\nlogging.info(f\"Extracted all zip files in {time() - start} seconds!\")\n```\n\nI didn't use `tqdm` because in multiprocessing it would show 8 tqdm results at the same time so it didn't make sense.\n\nThis created in `DATA` directory 49 folders. In order to work with all the files in a same folder I just created a folder with symlinks to each video file\n\n```python\nDATA = Path(\"path_to_the_49_zipped files\")\nTARGET = DATA.parent / \"full\"\n\ndef path_walk(top: Union[str,Path], topdown: bool = False, followlinks:bool = False):\n    \"\"\"See Python docs for os.walk, exact same behavior but it yields Path() instances instead\n    \"\"\"\n    \n    dirs = (node for node in top.iterdir() if node.is_dir())\n    nondirs =(node for node in top.iterdir() if not node.is_dir())\n\n    if topdown:\n        yield top, dirs, nondirs\n\n    for name in dirs:\n        if followlinks or name.is_symlink() is False:\n            for x in path_walk(name, topdown, followlinks):\n                yield x\n\n    if topdown is not True:\n        yield top, dirs, nondirs    \n\n#flatten the list of lists\nfiles = (f for top, dirs, files  in path_walk(DATA) for f in files if f.suffix == \".mp4\") \nfor file in files:\n    (TARGET / file.name).symlink_to(file)\n```\nSo in the `TARGET` I have all the *119146* video files to work on\n\nI hope this helps!\n\nCheers",
      "votes": 8,
      "replies": [
        {
          "id": 736084,
          "postDate": "2020-02-03T19:46:21.433Z",
          "content": "<p>just a small modification so each processed zip is deleted. I had limited storage...\nimport os\nimport multiprocessing\nfrom pathlib import Path\nfrom time import time\nfrom zipfile import ZipFile\nimport logging\nfrom typing import Union\n`\nDATA = Path(\"/data\")\nlogging.basicConfig(filename=\"extract.log\", level=logging.INFO)\nzipfiles = sorted(list(DATA.glob(\"*<em>/</em>.zip\")), key=lambda x: x.stem)</p>\n\n<p>def extract_zip(zipfile: Union[str, Path])-&gt;None:\n    with ZipFile(zipfile) as zip_file:\n        for file in zip_file.namelist():\n            zip_file.extract(member=file, path=zipfile.parent)\n    logging.info(f'Finished extracting {zipfile.stem}')\n    os.remove(zipfile)\n    logging.info(f'Finished deleting {zipfile.stem}')</p>\n\n<p>start = time()\nwith multiprocessing.Pool() as pool: # use all cores available\n    pool.map(extract_zip, zipfiles)</p>\n\n<p>logging.info(f\"Extracted all zip files in {time() - start} seconds!\")</p>",
          "rawMarkdown": "just a small modification so each processed zip is deleted. I had limited storage...\nimport os\nimport multiprocessing\nfrom pathlib import Path\nfrom time import time\nfrom zipfile import ZipFile\nimport logging\nfrom typing import Union\n`\nDATA = Path(\"/data\")\nlogging.basicConfig(filename=\"extract.log\", level=logging.INFO)\nzipfiles = sorted(list(DATA.glob(\"**/*.zip\")), key=lambda x: x.stem)\n\ndef extract_zip(zipfile: Union[str, Path])-&gt;None:\n    with ZipFile(zipfile) as zip_file:\n        for file in zip_file.namelist():\n            zip_file.extract(member=file, path=zipfile.parent)\n    logging.info(f'Finished extracting {zipfile.stem}')\n    os.remove(zipfile)\n    logging.info(f'Finished deleting {zipfile.stem}')\n\nstart = time()\nwith multiprocessing.Pool() as pool: # use all cores available\n    pool.map(extract_zip, zipfiles)\n\nlogging.info(f\"Extracted all zip files in {time() - start} seconds!\")",
          "votes": 1
        },
        {
          "id": 744687,
          "postDate": "2020-02-13T03:13:54.203Z",
          "content": "<p>This includes a few more tweaks. It moves the mp4 files into a <code>videos</code> directory, renames the metadata files to identify the zip file they came from, adds the zip file number to each video file entry and then saves the metadata files to a <code>metadata</code> directory. It also deletes each zip file once its files have been extracted. The <code>DEST</code> and <code>DEST/'videos'</code> and <code>DEST/'metadata'</code> directories need to exist before running.</p>\n\n<p>``` Python\nimport multiprocessing\nimport os\nfrom zipfile import ZipFile <br>\nimport logging\nfrom typing import Union\nfrom time import time\nimport pandas as pd</p>\n\n<p>DATA = Path()\nDEST = Path('destination_directory')\nlogging.basicConfig(filename='extract.log', level=logging.INFO)\nzipfiles = sorted(list(DATA.glob('dfdc_train_part_*.zip')), key=lambda x: x.stem)</p>\n\n<p>def extract_zip(zipfile: Union[str, Path])-&gt;None:\n    zip_no = zipfile.stem[-2:]\n    with ZipFile(zipfile) as zip_file:\n        for file in [Path(fn) for fn in zip_file.namelist()]:\n            try:\n                zip_info = zip_file.getinfo(str(file))\n                if file.suffix == '.json':\n                    zip_info.filename = f'{file.stem}_{zip_no}{file.suffix}'\n                    dest = DEST/'metadata'\n                else:\n                    zip_info.filename = file.name\n                    dest = DEST/'videos'\n                zip_file.extract(zip_info, path=dest)\n            except:\n                logging.error(f'error extracing {file.stem} from {zipfile.stem}')</p>\n\n<pre><code># add zip file number to metadata\nmeta_fn = DEST/'metadata'/f'metadata_{zip_no}.json'\ndf_meta = pd.read_json(meta_fn).T\ndf_meta['zip_no'] = zip_no\ndf_meta.to_json(meta_fn)\n\n# delete zip file\nzipfile.unlink()\n\nlogging.info(f'Finished extracting and deleted {zipfile.stem}')\n</code></pre>\n\n<p>start = int(time.time())\nwith multiprocessing.Pool() as pool: # use all cores available\n    pool.map(extract_zip, zipfiles)</p>\n\n<p>logging.info(f\"Extracted all zip files in {int(time.time()) - start} seconds!\")\n```</p>",
          "rawMarkdown": "This includes a few more tweaks. It moves the mp4 files into a `videos` directory, renames the metadata files to identify the zip file they came from, adds the zip file number to each video file entry and then saves the metadata files to a `metadata` directory. It also deletes each zip file once its files have been extracted. The `DEST` and `DEST/'videos'` and `DEST/'metadata'` directories need to exist before running.\n\n``` Python\nimport multiprocessing\nimport os\nfrom zipfile import ZipFile                                                                                                                                       \nimport logging\nfrom typing import Union\nfrom time import time\nimport pandas as pd\n\nDATA = Path()\nDEST = Path('destination_directory')\nlogging.basicConfig(filename='extract.log', level=logging.INFO)\nzipfiles = sorted(list(DATA.glob('dfdc_train_part_*.zip')), key=lambda x: x.stem)\n\ndef extract_zip(zipfile: Union[str, Path])-&gt;None:\n    zip_no = zipfile.stem[-2:]\n    with ZipFile(zipfile) as zip_file:\n        for file in [Path(fn) for fn in zip_file.namelist()]:\n            try:\n                zip_info = zip_file.getinfo(str(file))\n                if file.suffix == '.json':\n                    zip_info.filename = f'{file.stem}_{zip_no}{file.suffix}'\n                    dest = DEST/'metadata'\n                else:\n                    zip_info.filename = file.name\n                    dest = DEST/'videos'\n                zip_file.extract(zip_info, path=dest)\n            except:\n                logging.error(f'error extracing {file.stem} from {zipfile.stem}')\n            \n    # add zip file number to metadata\n    meta_fn = DEST/'metadata'/f'metadata_{zip_no}.json'\n    df_meta = pd.read_json(meta_fn).T\n    df_meta['zip_no'] = zip_no\n    df_meta.to_json(meta_fn)\n    \n    # delete zip file\n    zipfile.unlink()\n    \n    logging.info(f'Finished extracting and deleted {zipfile.stem}')\n    \nstart = int(time.time())\nwith multiprocessing.Pool() as pool: # use all cores available\n    pool.map(extract_zip, zipfiles)\n\nlogging.info(f\"Extracted all zip files in {int(time.time()) - start} seconds!\")\n```"
        }
      ]
    },
    {
      "id": 747307,
      "postDate": "2020-02-16T08:27:34.717Z",
      "content": "<p><a href=\"/carlossouza\">@carlossouza</a> \nit dsnt works for me ,not sure what m i missing .Please help\n1) i generated the cookies using my login credentials \nwget -qO- --keep-session-cookies --save-cookies cookies.txt --post-data  'user=username&amp;password=password' <a href=\"https://www.kaggle.com/account/login\">https://www.kaggle.com/account/login</a>?\n2) your  wget command. \n3) it downloads just the html files </p>",
      "rawMarkdown": "@carlossouza \nit dsnt works for me ,not sure what m i missing .Please help\n1) i generated the cookies using my login credentials \nwget -qO- --keep-session-cookies --save-cookies cookies.txt --post-data  'user=username&amp;password=password' https://www.kaggle.com/account/login?\n2) your  wget command. \n3) it downloads just the html files \n\n\n",
      "votes": 1
    },
    {
      "id": 696834,
      "postDate": "2019-12-17T05:54:17.740Z",
      "content": "<p>May I ask how many videos are in the dfdc_train_all folder? (not include .json file)\nI want to confirm, thank you!  </p>",
      "rawMarkdown": "May I ask how many videos are in the dfdc_train_all folder? (not include .json file)\nI want to confirm, thank you!  ",
      "votes": 1,
      "replies": [
        {
          "id": 696983,
          "postDate": "2019-12-17T10:17:44.500Z",
          "content": "<p>There are 119146 mp4 files.</p>",
          "rawMarkdown": "There are 119146 mp4 files.",
          "votes": 3
        }
      ]
    },
    {
      "id": 696084,
      "postDate": "2019-12-16T05:15:29.637Z",
      "content": "<p><a href=\"/diegojohnson\">@diegojohnson</a> ,\nRight now I’m playing with a setup that costs $0.8/hour. I’m budgeting ~10hs/week, which would give ~$8/week = $40/month.\nWhen I manage to fully utilize this setup (i.e. code optimized to run on multiple GPUs &amp; cores), and if I manage to produce competitive results, I might decide to beef up the setup and increase my budget limit.</p>",
      "rawMarkdown": "@diegojohnson ,\nRight now I’m playing with a setup that costs $0.8/hour. I’m budgeting ~10hs/week, which would give ~$8/week = $40/month.\nWhen I manage to fully utilize this setup (i.e. code optimized to run on multiple GPUs &amp; cores), and if I manage to produce competitive results, I might decide to beef up the setup and increase my budget limit.",
      "votes": 1,
      "replies": [
        {
          "id": 696098,
          "postDate": "2019-12-16T05:44:47.110Z",
          "content": "<p>Thanks very much😃 </p>",
          "rawMarkdown": "Thanks very much😃 "
        },
        {
          "id": 709125,
          "postDate": "2020-01-03T04:13:50.583Z",
          "content": "<p>Google cloud also have $300 free tier. Which will cover more than 7 month for you.</p>",
          "rawMarkdown": "Google cloud also have $300 free tier. Which will cover more than 7 month for you.",
          "votes": 1
        }
      ]
    },
    {
      "id": 695670,
      "postDate": "2019-12-15T12:51:26.103Z",
      "content": "<p>To unzip all the files at once, you can also do <code>unzip '*.zip'</code>.</p>",
      "rawMarkdown": "To unzip all the files at once, you can also do `unzip '*.zip'`.",
      "votes": 1,
      "replies": [
        {
          "id": 695675,
          "postDate": "2019-12-15T12:58:43.173Z",
          "content": "<p>Yes, but I don’t recommend. Unzipping these files took forever to finish. With the code above, you will see the progress. Unzip * won’t show any progress, you won’t know how long to finish and whether it was successful or not...</p>",
          "rawMarkdown": "Yes, but I don’t recommend. Unzipping these files took forever to finish. With the code above, you will see the progress. Unzip * won’t show any progress, you won’t know how long to finish and whether it was successful or not...",
          "votes": 2
        }
      ]
    },
    {
      "id": 702450,
      "postDate": "2019-12-24T17:32:43.643Z",
      "content": "<p>deserves the gold :) thanks fr quality content , upvoted!</p>",
      "rawMarkdown": "deserves the gold :) thanks fr quality content , upvoted!",
      "votes": 2
    },
    {
      "id": 777682,
      "postDate": "2020-03-17T21:12:44.907Z",
      "content": "<p>Could you please teach me how to save my kaggle credentials in cookies.txt?</p>",
      "rawMarkdown": "Could you please teach me how to save my kaggle credentials in cookies.txt?"
    },
    {
      "id": 759721,
      "postDate": "2020-02-29T11:37:02.623Z",
      "content": "<p>Very helpful! Thanks for sharing your scripts. A quick question: how do you get the data that you put inside your cookies.txt file? Using your browser cookies info I guess or something more elaborate? </p>",
      "rawMarkdown": "Very helpful! Thanks for sharing your scripts. A quick question: how do you get the data that you put inside your cookies.txt file? Using your browser cookies info I guess or something more elaborate? "
    },
    {
      "id": 739347,
      "postDate": "2020-02-07T18:23:32.747Z",
      "content": "<p>Hi, I was using this <a href=\"https://www.kaggle.com/frsanchez/download-kaggle-files-from-notebook\">kernel</a> to download files from a notebook, however it worked days ago and it doesn't work anymore, I have a 404 Not Found error.</p>\n\n<p>I tried using the <code>wget</code> command and loading cookies but I didn't get exactly what must be into the <code>cookies.txt</code> file. I tried to put the <code>cURL</code> command as explained <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121194#695299\">here</a> but it's not working.</p>\n\n<p>Have you already experienced a 404 Not Found error with the Kaggle API (given that the same code was working before)? For the <code>wget</code> command, is it the cURL command that should be into the <code>cookies.txt</code>?</p>\n\n<p>Thanks for your help!</p>",
      "rawMarkdown": "Hi, I was using this [kernel](https://www.kaggle.com/frsanchez/download-kaggle-files-from-notebook) to download files from a notebook, however it worked days ago and it doesn't work anymore, I have a 404 Not Found error.\n\nI tried using the `wget` command and loading cookies but I didn't get exactly what must be into the `cookies.txt` file. I tried to put the `cURL` command as explained [here](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121194#695299) but it's not working.\n\nHave you already experienced a 404 Not Found error with the Kaggle API (given that the same code was working before)? For the `wget` command, is it the cURL command that should be into the `cookies.txt`?\n\nThanks for your help!"
    },
    {
      "id": 716142,
      "postDate": "2020-01-11T10:05:20.510Z",
      "content": "<p>I could not figure out the way to use kaggle api to download so wget is so far the best way...And unzipping while deleting codes...Thanks for sharing!! </p>",
      "rawMarkdown": "I could not figure out the way to use kaggle api to download so wget is so far the best way...And unzipping while deleting codes...Thanks for sharing!! ",
      "replies": [
        {
          "id": 719871,
          "postDate": "2020-01-16T00:07:17.017Z",
          "content": "<p>I have implemented a notebook that can download programmatically any file, which is very useful if you want to download them from Colab: <a href=\"https://www.kaggle.com/frsanchez/download-kaggle-files-from-notebook\">https://www.kaggle.com/frsanchez/download-kaggle-files-from-notebook</a></p>",
          "rawMarkdown": "I have implemented a notebook that can download programmatically any file, which is very useful if you want to download them from Colab: https://www.kaggle.com/frsanchez/download-kaggle-files-from-notebook",
          "votes": 2
        },
        {
          "id": 720677,
          "postDate": "2020-01-16T16:27:58.560Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 701124,
      "postDate": "2019-12-23T06:00:09.983Z",
      "content": "<p>Hello Carlos,</p>\n\n<p>Wtih Wget and CurlWget the download stops after about 3%.\nSaving to: ‘dfdc_train_all.zip’</p>\n\n<p>dfdc_train_all.zip                 3%[=&gt;                                                      ]  17.42G  78.3MB/s    in 2m 50s  </p>\n\n<p>Cannot write to ‘dfdc_train_all.zip’ (Success).</p>\n\n<p>Any idea?\nthank you,</p>",
      "rawMarkdown": "Hello Carlos,\n\nWtih Wget and CurlWget the download stops after about 3%.\nSaving to: ‘dfdc_train_all.zip’\n\ndfdc_train_all.zip                 3%[=&gt;                                                      ]  17.42G  78.3MB/s    in 2m 50s  \n\n\nCannot write to ‘dfdc_train_all.zip’ (Success).\n\nAny idea?\nthank you,",
      "replies": [
        {
          "id": 709123,
          "postDate": "2020-01-03T04:11:09.710Z",
          "content": "<p>You have used all of your disk space. Switching to a disk that have more space solved the problem for me. You can check it by executing:\n<code>df</code>\nand your root volume will be 100% full. \nHope this helps!</p>",
          "rawMarkdown": "You have used all of your disk space. Switching to a disk that have more space solved the problem for me. You can check it by executing:\n`df`\nand your root volume will be 100% full. \nHope this helps!"
        }
      ]
    },
    {
      "id": 700098,
      "postDate": "2019-12-21T13:08:57.130Z",
      "content": "<p>Can someone give md5-sums of the files?</p>",
      "rawMarkdown": "Can someone give md5-sums of the files?"
    },
    {
      "id": 698176,
      "postDate": "2019-12-18T22:29:11.190Z",
      "content": "<p>hey <a href=\"/carlossouza\">@carlossouza</a>,</p>\n\n<p>When trying to download the data with wget and my own cookies everything seems ok but the download stops instantly without any error nor data downloaded (only zip file created). Any idea what's happening?</p>\n\n<p>Here is some more information : </p>\n\n<p>When running:\n<code>\nwget --load-cookies cookies.txt https://www.kaggle.com/c/16880/datadownload/dfdc_train_all.zip\n</code> \nmy terminal says this and then nothing:\n```\n--2019-12-18 23:13:26--  <a href=\"https://www.kaggle.com/c/16880/datadownload/dfdc_train_all.zip\">https://www.kaggle.com/c/16880/datadownload/dfdc_train_all.zip</a>\nResolving www.kaggle.com (www.kaggle.com)... \nConnecting to www.kaggle.com (www.kaggle.com)|:443... connected.\nHTTP request sent, awaiting response... 302 Found\nLocation: <a href=\"https://www.kaggle.com/account/login?ReturnUrl=%2Fc%2F16880%2Fdatadownload%2Fdfdc_train_all.zip\">https://www.kaggle.com/account/login?ReturnUrl=%2Fc%2F16880%2Fdatadownload%2Fdfdc_train_all.zip</a> [following]\n--2019-12-18 23:13:27--  <a href=\"https://www.kaggle.com/account/login?ReturnUrl=%2Fc%2F16880%2Fdatadownload%2Fdfdc_train_all.zip\">https://www.kaggle.com/account/login?ReturnUrl=%2Fc%2F16880%2Fdatadownload%2Fdfdc_train_all.zip</a>\nReusing existing connection to www.kaggle.com:443.\nHTTP request sent, awaiting response... 200 OK\nLength: unspecified [text/html]\nSaving to: ‘dfdc_train_all.zip’</p>\n\n<p>dfdc_train_all.zip                                     [ &lt;=&gt;                                                                                                             ]   8,96K  --.-KB/s    in 0,01s   </p>\n\n<p>2019-12-18 23:13:27 (940 KB/s) - ‘dfdc_train_all.zip’ saved [9170]\n```</p>",
      "rawMarkdown": "hey @carlossouza,\n\nWhen trying to download the data with wget and my own cookies everything seems ok but the download stops instantly without any error nor data downloaded (only zip file created). Any idea what's happening?\n\n\nHere is some more information : \n\nWhen running:\n```\nwget --load-cookies cookies.txt https://www.kaggle.com/c/16880/datadownload/dfdc_train_all.zip\n``` \nmy terminal says this and then nothing:\n```\n--2019-12-18 23:13:26--  https://www.kaggle.com/c/16880/datadownload/dfdc_train_all.zip\nResolving www.kaggle.com (www.kaggle.com)... \nConnecting to www.kaggle.com (www.kaggle.com)|:443... connected.\nHTTP request sent, awaiting response... 302 Found\nLocation: https://www.kaggle.com/account/login?ReturnUrl=%2Fc%2F16880%2Fdatadownload%2Fdfdc_train_all.zip [following]\n--2019-12-18 23:13:27--  https://www.kaggle.com/account/login?ReturnUrl=%2Fc%2F16880%2Fdatadownload%2Fdfdc_train_all.zip\nReusing existing connection to www.kaggle.com:443.\nHTTP request sent, awaiting response... 200 OK\nLength: unspecified [text/html]\nSaving to: ‘dfdc_train_all.zip’\n\ndfdc_train_all.zip                                     [ &lt;=&gt;                                                                                                             ]   8,96K  --.-KB/s    in 0,01s   \n\n2019-12-18 23:13:27 (940 KB/s) - ‘dfdc_train_all.zip’ saved [9170]\n```",
      "replies": [
        {
          "id": 698178,
          "postDate": "2019-12-18T22:37:00.427Z",
          "content": "<p>Probably something wrong with the cookies.. This is exactly what happens if you try without them.</p>",
          "rawMarkdown": "Probably something wrong with the cookies.. This is exactly what happens if you try without them.",
          "votes": 3
        },
        {
          "id": 698448,
          "postDate": "2019-12-19T09:00:46.657Z",
          "content": "<p>You were right, I uploaded the wrong cookies, did not know they were tab dependents. Working like a charm... well ETA in 15h40min ^^\nThanks!</p>",
          "rawMarkdown": "You were right, I uploaded the wrong cookies, did not know they were tab dependents. Working like a charm... well ETA in 15h40min ^^\nThanks!"
        },
        {
          "id": 701506,
          "postDate": "2019-12-23T15:04:05.103Z",
          "content": "<p>Hello, I have the same problem, but I can't edit the cookie correctly Could you please provide the structure of how you set cookie.txt. That would be of great help.</p>",
          "rawMarkdown": "Hello, I have the same problem, but I can't edit the cookie correctly Could you please provide the structure of how you set cookie.txt. That would be of great help.",
          "votes": 1
        },
        {
          "id": 701575,
          "postDate": "2019-12-23T16:22:11.390Z",
          "content": "<p>I just installed this on chrome : <a href=\"https://chrome.google.com/webstore/detail/cookiestxt/njabckikapfpffapmjgojcnbfjonfjfg?hl=en\">https://chrome.google.com/webstore/detail/cookiestxt/njabckikapfpffapmjgojcnbfjonfjfg?hl=en</a>\nThen you can extract the cookies by clicking on the cookie drawing at top right of your browser, you just need to copy paste this into a file you name cookies.txt. Just make sure that you are on a Kaggle web page (and connected) while extracting cookies as they are specific for each tab and what you need are Kaggle's cookies.</p>",
          "rawMarkdown": "I just installed this on chrome : https://chrome.google.com/webstore/detail/cookiestxt/njabckikapfpffapmjgojcnbfjonfjfg?hl=en\nThen you can extract the cookies by clicking on the cookie drawing at top right of your browser, you just need to copy paste this into a file you name cookies.txt. Just make sure that you are on a Kaggle web page (and connected) while extracting cookies as they are specific for each tab and what you need are Kaggle's cookies.",
          "votes": 10
        },
        {
          "id": 701595,
          "postDate": "2019-12-23T16:38:11.130Z",
          "content": "<p>Thank you so much, now it worked !!</p>",
          "rawMarkdown": "Thank you so much, now it worked !!",
          "votes": 1
        },
        {
          "id": 707228,
          "postDate": "2019-12-31T12:55:14.660Z",
          "content": "<p>This works like a charm, Thank you 👍 </p>",
          "rawMarkdown": "This works like a charm, Thank you 👍 ",
          "votes": 1
        },
        {
          "id": 711716,
          "postDate": "2020-01-06T13:14:51.297Z",
          "content": "<p>I tried as <a href=\"/optimo\">@optimo</a> suggested, and uploaded both types of cookies - for kaggle tab, and also entirely. In both situations the wget starts downloading the file, and stops immediately. I'm using AWS EC2 instance, with 1TB of storage. any suggestions?</p>",
          "rawMarkdown": "I tried as @optimo suggested, and uploaded both types of cookies - for kaggle tab, and also entirely. In both situations the wget starts downloading the file, and stops immediately. I'm using AWS EC2 instance, with 1TB of storage. any suggestions?"
        },
        {
          "id": 712325,
          "postDate": "2020-01-07T05:03:53.020Z",
          "content": "<p>Install <a href=\"https://www.google.com/url?sa=t&amp;rct=j&amp;q=&amp;esrc=s&amp;source=web&amp;cd=1&amp;cad=rja&amp;uact=8&amp;ved=2ahUKEwjo3qaa2fDmAhUdH7kGHehKAa0QFjAAegQIBRAB&amp;url=https%3A%2F%2Faddons.mozilla.org%2Fpt-BR%2Ffirefox%2Faddon%2Fcliget%2F&amp;usg=AOvVaw22TdCwxvos28AdJBBktX-G\">cliget</a> for Firefox or <a href=\"https://chrome.google.com/webstore/detail/curlwget/jmocjfidanebdlinpbcdkcmgdifblncg?hl=en\">CurlWget</a> for Chrome.</p>\n\n<p>Once installed, click on the link to download. Wait for the dialog to appear, cancel the download.</p>\n\n<p>Click on the icon in the extension, copy the code displayed  and paste in the terminal in your AWS instance. The download should start immediately.</p>\n\n<p>I suggest to open a tmux session just to download this, and let it running in the background:\n<code>bash\ntmux new -s download #starts a session in tmux\n</code>\nCopy the code from the extension and hit Enter.\nOnce it starts downloading: \n<code>bash\n&lt;Ctrl&gt;&lt;b&gt;&lt;d&gt; # detach from the session\n</code>\nTo check once in a while how the download is going:\n<code>bash\ntmux attach -t download\n</code>\nor only <code>tmux attach</code></p>\n\n<p>I hope it helps </p>",
          "rawMarkdown": "Install [cliget](https://www.google.com/url?sa=t&amp;rct=j&amp;q=&amp;esrc=s&amp;source=web&amp;cd=1&amp;cad=rja&amp;uact=8&amp;ved=2ahUKEwjo3qaa2fDmAhUdH7kGHehKAa0QFjAAegQIBRAB&amp;url=https%3A%2F%2Faddons.mozilla.org%2Fpt-BR%2Ffirefox%2Faddon%2Fcliget%2F&amp;usg=AOvVaw22TdCwxvos28AdJBBktX-G) for Firefox or [CurlWget](https://chrome.google.com/webstore/detail/curlwget/jmocjfidanebdlinpbcdkcmgdifblncg?hl=en) for Chrome.\n\nOnce installed, click on the link to download. Wait for the dialog to appear, cancel the download.\n\nClick on the icon in the extension, copy the code displayed  and paste in the terminal in your AWS instance. The download should start immediately.\n\nI suggest to open a tmux session just to download this, and let it running in the background:\n```bash\ntmux new -s download #starts a session in tmux\n```\nCopy the code from the extension and hit Enter.\nOnce it starts downloading: \n```bash\n",
          "votes": 1
        },
        {
          "id": 747182,
          "postDate": "2020-02-16T04:39:37.837Z",
          "content": "<p>Works like a charm!</p>",
          "rawMarkdown": "Works like a charm!"
        }
      ]
    },
    {
      "id": 696833,
      "postDate": "2019-12-17T05:47:40.033Z",
      "content": "<p>Hi Carlos , This is really helpful . Do you mean VM Instance setup in Google Cloud ? </p>\n\n<p>From the list of steps above \"While setting up the VM, I've chosen 1TB SSD disk storage: less than that would not allow me download and unzip the complete 472 GB full dataset at once;\"</p>",
      "rawMarkdown": "Hi Carlos , This is really helpful . Do you mean VM Instance setup in Google Cloud ? \n\nFrom the list of steps above \"While setting up the VM, I've chosen 1TB SSD disk storage: less than that would not allow me download and unzip the complete 472 GB full dataset at once;\""
    },
    {
      "id": 696058,
      "postDate": "2019-12-16T04:20:58.233Z",
      "content": "<p>Hi, I have no idea about cost of Google Cloud, could you give me some information about it? How much does it cost in a month to complete this competition? thanks. <a href=\"/carlossouza\">@carlossouza</a> </p>",
      "rawMarkdown": "Hi, I have no idea about cost of Google Cloud, could you give me some information about it? How much does it cost in a month to complete this competition? thanks. @carlossouza ",
      "replies": [
        {
          "id": 709074,
          "postDate": "2020-01-03T02:49:00.547Z",
          "rawMarkdown": ""
        }
      ]
    },
    {
      "id": 720672,
      "postDate": "2020-01-16T16:24:19.573Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 720671,
      "postDate": "2020-01-16T16:22:30.557Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 708822,
      "postDate": "2020-01-02T18:04:54.077Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 700077,
      "postDate": "2019-12-21T12:25:10.560Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 696097,
      "postDate": "2019-12-16T05:44:23.980Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 701600,
      "author_name": "Ronaldo S.A. Batista",
      "author_url": "",
      "post_date": "2019-12-23T16:42:50.660000",
      "content": "<p>My alternative to step 3 delineated by Carlos Above:</p>\n\n<p>&gt; 3) Unzip everything</p>\n\n<pre><code>The file dfdctrainall.zip is like a Russian doll: it contains 49 zip files;\nAfter unzipping and deleting it, I applied the following script to unzip each of the 49 files:\n</code></pre>\n\n<p>Since unzip is a disk/processor bottleneck the best way to use concurrency is to use all processors. ( i.e. Multiprocessing and not Multithreading )</p>\n\n<p>```python\nimport multiprocessing <br>\nfrom pathlib import Path <br>\nfrom time import time <br>\nfrom zipfile import ZipFile <br>\nimport logging\nfrom typing import Union                                                                                                                                                    </p>\n\n<p>DATA = Path(\"path_to_the_49_zipped files\") <br>\nlogging.basicConfig(filename=\"extract.log\", level=logging.INFO)\nzipfiles = sorted(list(DATA.glob(\"*<em>/</em>.zip\")), key=lambda x: x.stem)</p>\n\n<p>def extract_zip(zipfile: Union[str, Path])-&gt;None:\n    with ZipFile(zipfile) as zip_file:\n        for file in zip_file.namelist():\n            zip_file.extract(member=file, path=zipfile.parent)\n    logging.info(f'Finished extracting {zipfile.stem}')</p>\n\n<p>start = time()\nwith multiprocessing.Pool() as pool: # use all cores available\n    pool.map(extract_zip, zipfiles)</p>\n\n<p>logging.info(f\"Extracted all zip files in {time() - start} seconds!\")\n```</p>\n\n<p>I didn't use <code>tqdm</code> because in multiprocessing it would show 8 tqdm results at the same time so it didn't make sense.</p>\n\n<p>This created in <code>DATA</code> directory 49 folders. In order to work with all the files in a same folder I just created a folder with symlinks to each video file</p>\n\n<p>```python\nDATA = Path(\"path_to_the_49_zipped files\")\nTARGET = DATA.parent / \"full\"</p>\n\n<p>def path_walk(top: Union[str,Path], topdown: bool = False, followlinks:bool = False):\n    \"\"\"See Python docs for os.walk, exact same behavior but it yields Path() instances instead\n    \"\"\"</p>\n\n<pre><code>dirs = (node for node in top.iterdir() if node.is_dir())\nnondirs =(node for node in top.iterdir() if not node.is_dir())\n\nif topdown:\n    yield top, dirs, nondirs\n\nfor name in dirs:\n    if followlinks or name.is_symlink() is False:\n        for x in path_walk(name, topdown, followlinks):\n            yield x\n\nif topdown is not True:\n    yield top, dirs, nondirs    \n</code></pre>\n\n<h1>flatten the list of lists</h1>\n\n<p>files = (f for top, dirs, files  in path_walk(DATA) for f in files if f.suffix == \".mp4\") \nfor file in files:\n    (TARGET / file.name).symlink_to(file)\n<code>``\nSo in the</code>TARGET` I have all the <em>119146</em> video files to work on</p>\n\n<p>I hope this helps!</p>\n\n<p>Cheers</p>",
      "votes": 8,
      "replies": [
        {
          "id": 736084,
          "author_name": "rrkitlo",
          "author_url": "",
          "post_date": "2020-02-03T19:46:21.433000",
          "content": "<p>just a small modification so each processed zip is deleted. I had limited storage...\nimport os\nimport multiprocessing\nfrom pathlib import Path\nfrom time import time\nfrom zipfile import ZipFile\nimport logging\nfrom typing import Union\n`\nDATA = Path(\"/data\")\nlogging.basicConfig(filename=\"extract.log\", level=logging.INFO)\nzipfiles = sorted(list(DATA.glob(\"*<em>/</em>.zip\")), key=lambda x: x.stem)</p>\n\n<p>def extract_zip(zipfile: Union[str, Path])-&gt;None:\n    with ZipFile(zipfile) as zip_file:\n        for file in zip_file.namelist():\n            zip_file.extract(member=file, path=zipfile.parent)\n    logging.info(f'Finished extracting {zipfile.stem}')\n    os.remove(zipfile)\n    logging.info(f'Finished deleting {zipfile.stem}')</p>\n\n<p>start = time()\nwith multiprocessing.Pool() as pool: # use all cores available\n    pool.map(extract_zip, zipfiles)</p>\n\n<p>logging.info(f\"Extracted all zip files in {time() - start} seconds!\")</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 744687,
          "author_name": "Caleb",
          "author_url": "",
          "post_date": "2020-02-13T03:13:54.203000",
          "content": "<p>This includes a few more tweaks. It moves the mp4 files into a <code>videos</code> directory, renames the metadata files to identify the zip file they came from, adds the zip file number to each video file entry and then saves the metadata files to a <code>metadata</code> directory. It also deletes each zip file once its files have been extracted. The <code>DEST</code> and <code>DEST/'videos'</code> and <code>DEST/'metadata'</code> directories need to exist before running.</p>\n\n<p>``` Python\nimport multiprocessing\nimport os\nfrom zipfile import ZipFile <br>\nimport logging\nfrom typing import Union\nfrom time import time\nimport pandas as pd</p>\n\n<p>DATA = Path()\nDEST = Path('destination_directory')\nlogging.basicConfig(filename='extract.log', level=logging.INFO)\nzipfiles = sorted(list(DATA.glob('dfdc_train_part_*.zip')), key=lambda x: x.stem)</p>\n\n<p>def extract_zip(zipfile: Union[str, Path])-&gt;None:\n    zip_no = zipfile.stem[-2:]\n    with ZipFile(zipfile) as zip_file:\n        for file in [Path(fn) for fn in zip_file.namelist()]:\n            try:\n                zip_info = zip_file.getinfo(str(file))\n                if file.suffix == '.json':\n                    zip_info.filename = f'{file.stem}_{zip_no}{file.suffix}'\n                    dest = DEST/'metadata'\n                else:\n                    zip_info.filename = file.name\n                    dest = DEST/'videos'\n                zip_file.extract(zip_info, path=dest)\n            except:\n                logging.error(f'error extracing {file.stem} from {zipfile.stem}')</p>\n\n<pre><code># add zip file number to metadata\nmeta_fn = DEST/'metadata'/f'metadata_{zip_no}.json'\ndf_meta = pd.read_json(meta_fn).T\ndf_meta['zip_no'] = zip_no\ndf_meta.to_json(meta_fn)\n\n# delete zip file\nzipfile.unlink()\n\nlogging.info(f'Finished extracting and deleted {zipfile.stem}')\n</code></pre>\n\n<p>start = int(time.time())\nwith multiprocessing.Pool() as pool: # use all cores available\n    pool.map(extract_zip, zipfiles)</p>\n\n<p>logging.info(f\"Extracted all zip files in {int(time.time()) - start} seconds!\")\n```</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 747307,
      "author_name": "Jaideep",
      "author_url": "",
      "post_date": "2020-02-16T08:27:34.717000",
      "content": "<p><a href=\"/carlossouza\">@carlossouza</a> \nit dsnt works for me ,not sure what m i missing .Please help\n1) i generated the cookies using my login credentials \nwget -qO- --keep-session-cookies --save-cookies cookies.txt --post-data  'user=username&amp;password=password' <a href=\"https://www.kaggle.com/account/login\">https://www.kaggle.com/account/login</a>?\n2) your  wget command. \n3) it downloads just the html files </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 696834,
      "author_name": "jerry lo",
      "author_url": "",
      "post_date": "2019-12-17T05:54:17.740000",
      "content": "<p>May I ask how many videos are in the dfdc_train_all folder? (not include .json file)\nI want to confirm, thank you!  </p>",
      "votes": 1,
      "replies": [
        {
          "id": 696983,
          "author_name": "Human Analog",
          "author_url": "",
          "post_date": "2019-12-17T10:17:44.500000",
          "content": "<p>There are 119146 mp4 files.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 696084,
      "author_name": "Carlos Souza",
      "author_url": "",
      "post_date": "2019-12-16T05:15:29.637000",
      "content": "<p><a href=\"/diegojohnson\">@diegojohnson</a> ,\nRight now I’m playing with a setup that costs $0.8/hour. I’m budgeting ~10hs/week, which would give ~$8/week = $40/month.\nWhen I manage to fully utilize this setup (i.e. code optimized to run on multiple GPUs &amp; cores), and if I manage to produce competitive results, I might decide to beef up the setup and increase my budget limit.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 696098,
          "author_name": "DiegoJohnson",
          "author_url": "",
          "post_date": "2019-12-16T05:44:47.110000",
          "content": "<p>Thanks very much😃 </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 709125,
          "author_name": "Shangqiu Li",
          "author_url": "",
          "post_date": "2020-01-03T04:13:50.583000",
          "content": "<p>Google cloud also have $300 free tier. Which will cover more than 7 month for you.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 695670,
      "author_name": "Human Analog",
      "author_url": "",
      "post_date": "2019-12-15T12:51:26.103000",
      "content": "<p>To unzip all the files at once, you can also do <code>unzip '*.zip'</code>.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 695675,
          "author_name": "Carlos Souza",
          "author_url": "",
          "post_date": "2019-12-15T12:58:43.173000",
          "content": "<p>Yes, but I don’t recommend. Unzipping these files took forever to finish. With the code above, you will see the progress. Unzip * won’t show any progress, you won’t know how long to finish and whether it was successful or not...</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 702450,
      "author_name": "ravi tanwar",
      "author_url": "",
      "post_date": "2019-12-24T17:32:43.643000",
      "content": "<p>deserves the gold :) thanks fr quality content , upvoted!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 777682,
      "author_name": "madmax0404",
      "author_url": "",
      "post_date": "2020-03-17T21:12:44.907000",
      "content": "<p>Could you please teach me how to save my kaggle credentials in cookies.txt?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 759721,
      "author_name": "Yassine Alouini",
      "author_url": "",
      "post_date": "2020-02-29T11:37:02.623000",
      "content": "<p>Very helpful! Thanks for sharing your scripts. A quick question: how do you get the data that you put inside your cookies.txt file? Using your browser cookies info I guess or something more elaborate? </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 739347,
      "author_name": "Michaël Karpe",
      "author_url": "",
      "post_date": "2020-02-07T18:23:32.747000",
      "content": "<p>Hi, I was using this <a href=\"https://www.kaggle.com/frsanchez/download-kaggle-files-from-notebook\">kernel</a> to download files from a notebook, however it worked days ago and it doesn't work anymore, I have a 404 Not Found error.</p>\n\n<p>I tried using the <code>wget</code> command and loading cookies but I didn't get exactly what must be into the <code>cookies.txt</code> file. I tried to put the <code>cURL</code> command as explained <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121194#695299\">here</a> but it's not working.</p>\n\n<p>Have you already experienced a 404 Not Found error with the Kaggle API (given that the same code was working before)? For the <code>wget</code> command, is it the cURL command that should be into the <code>cookies.txt</code>?</p>\n\n<p>Thanks for your help!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 716142,
      "author_name": "Owen Xu",
      "author_url": "",
      "post_date": "2020-01-11T10:05:20.510000",
      "content": "<p>I could not figure out the way to use kaggle api to download so wget is so far the best way...And unzipping while deleting codes...Thanks for sharing!! </p>",
      "votes": 0,
      "replies": [
        {
          "id": 719871,
          "author_name": "F.J. Sanchez",
          "author_url": "",
          "post_date": "2020-01-16T00:07:17.017000",
          "content": "<p>I have implemented a notebook that can download programmatically any file, which is very useful if you want to download them from Colab: <a href=\"https://www.kaggle.com/frsanchez/download-kaggle-files-from-notebook\">https://www.kaggle.com/frsanchez/download-kaggle-files-from-notebook</a></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 720677,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-01-16T16:27:58.560000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 701124,
      "author_name": "Sam K",
      "author_url": "",
      "post_date": "2019-12-23T06:00:09.983000",
      "content": "<p>Hello Carlos,</p>\n\n<p>Wtih Wget and CurlWget the download stops after about 3%.\nSaving to: ‘dfdc_train_all.zip’</p>\n\n<p>dfdc_train_all.zip                 3%[=&gt;                                                      ]  17.42G  78.3MB/s    in 2m 50s  </p>\n\n<p>Cannot write to ‘dfdc_train_all.zip’ (Success).</p>\n\n<p>Any idea?\nthank you,</p>",
      "votes": 0,
      "replies": [
        {
          "id": 709123,
          "author_name": "Shangqiu Li",
          "author_url": "",
          "post_date": "2020-01-03T04:11:09.710000",
          "content": "<p>You have used all of your disk space. Switching to a disk that have more space solved the problem for me. You can check it by executing:\n<code>df</code>\nand your root volume will be 100% full. \nHope this helps!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 700098,
      "author_name": "Kinimod",
      "author_url": "",
      "post_date": "2019-12-21T13:08:57.130000",
      "content": "<p>Can someone give md5-sums of the files?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 698176,
      "author_name": "Optimo",
      "author_url": "",
      "post_date": "2019-12-18T22:29:11.190000",
      "content": "<p>hey <a href=\"/carlossouza\">@carlossouza</a>,</p>\n\n<p>When trying to download the data with wget and my own cookies everything seems ok but the download stops instantly without any error nor data downloaded (only zip file created). Any idea what's happening?</p>\n\n<p>Here is some more information : </p>\n\n<p>When running:\n<code>\nwget --load-cookies cookies.txt https://www.kaggle.com/c/16880/datadownload/dfdc_train_all.zip\n</code> \nmy terminal says this and then nothing:\n```\n--2019-12-18 23:13:26--  <a href=\"https://www.kaggle.com/c/16880/datadownload/dfdc_train_all.zip\">https://www.kaggle.com/c/16880/datadownload/dfdc_train_all.zip</a>\nResolving www.kaggle.com (www.kaggle.com)... \nConnecting to www.kaggle.com (www.kaggle.com)|:443... connected.\nHTTP request sent, awaiting response... 302 Found\nLocation: <a href=\"https://www.kaggle.com/account/login?ReturnUrl=%2Fc%2F16880%2Fdatadownload%2Fdfdc_train_all.zip\">https://www.kaggle.com/account/login?ReturnUrl=%2Fc%2F16880%2Fdatadownload%2Fdfdc_train_all.zip</a> [following]\n--2019-12-18 23:13:27--  <a href=\"https://www.kaggle.com/account/login?ReturnUrl=%2Fc%2F16880%2Fdatadownload%2Fdfdc_train_all.zip\">https://www.kaggle.com/account/login?ReturnUrl=%2Fc%2F16880%2Fdatadownload%2Fdfdc_train_all.zip</a>\nReusing existing connection to www.kaggle.com:443.\nHTTP request sent, awaiting response... 200 OK\nLength: unspecified [text/html]\nSaving to: ‘dfdc_train_all.zip’</p>\n\n<p>dfdc_train_all.zip                                     [ &lt;=&gt;                                                                                                             ]   8,96K  --.-KB/s    in 0,01s   </p>\n\n<p>2019-12-18 23:13:27 (940 KB/s) - ‘dfdc_train_all.zip’ saved [9170]\n```</p>",
      "votes": 0,
      "replies": [
        {
          "id": 698178,
          "author_name": "Carlos Souza",
          "author_url": "",
          "post_date": "2019-12-18T22:37:00.427000",
          "content": "<p>Probably something wrong with the cookies.. This is exactly what happens if you try without them.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 698448,
          "author_name": "Optimo",
          "author_url": "",
          "post_date": "2019-12-19T09:00:46.657000",
          "content": "<p>You were right, I uploaded the wrong cookies, did not know they were tab dependents. Working like a charm... well ETA in 15h40min ^^\nThanks!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 701506,
          "author_name": "Luciano Batista",
          "author_url": "",
          "post_date": "2019-12-23T15:04:05.103000",
          "content": "<p>Hello, I have the same problem, but I can't edit the cookie correctly Could you please provide the structure of how you set cookie.txt. That would be of great help.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 701575,
          "author_name": "Optimo",
          "author_url": "",
          "post_date": "2019-12-23T16:22:11.390000",
          "content": "<p>I just installed this on chrome : <a href=\"https://chrome.google.com/webstore/detail/cookiestxt/njabckikapfpffapmjgojcnbfjonfjfg?hl=en\">https://chrome.google.com/webstore/detail/cookiestxt/njabckikapfpffapmjgojcnbfjonfjfg?hl=en</a>\nThen you can extract the cookies by clicking on the cookie drawing at top right of your browser, you just need to copy paste this into a file you name cookies.txt. Just make sure that you are on a Kaggle web page (and connected) while extracting cookies as they are specific for each tab and what you need are Kaggle's cookies.</p>",
          "votes": 10,
          "replies": []
        },
        {
          "id": 701595,
          "author_name": "Luciano Batista",
          "author_url": "",
          "post_date": "2019-12-23T16:38:11.130000",
          "content": "<p>Thank you so much, now it worked !!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 707228,
          "author_name": "Sundeep Pidugu",
          "author_url": "",
          "post_date": "2019-12-31T12:55:14.660000",
          "content": "<p>This works like a charm, Thank you 👍 </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 711716,
          "author_name": "autonomi",
          "author_url": "",
          "post_date": "2020-01-06T13:14:51.297000",
          "content": "<p>I tried as <a href=\"/optimo\">@optimo</a> suggested, and uploaded both types of cookies - for kaggle tab, and also entirely. In both situations the wget starts downloading the file, and stops immediately. I'm using AWS EC2 instance, with 1TB of storage. any suggestions?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 712325,
          "author_name": "Ronaldo S.A. Batista",
          "author_url": "",
          "post_date": "2020-01-07T05:03:53.020000",
          "content": "<p>Install <a href=\"https://www.google.com/url?sa=t&amp;rct=j&amp;q=&amp;esrc=s&amp;source=web&amp;cd=1&amp;cad=rja&amp;uact=8&amp;ved=2ahUKEwjo3qaa2fDmAhUdH7kGHehKAa0QFjAAegQIBRAB&amp;url=https%3A%2F%2Faddons.mozilla.org%2Fpt-BR%2Ffirefox%2Faddon%2Fcliget%2F&amp;usg=AOvVaw22TdCwxvos28AdJBBktX-G\">cliget</a> for Firefox or <a href=\"https://chrome.google.com/webstore/detail/curlwget/jmocjfidanebdlinpbcdkcmgdifblncg?hl=en\">CurlWget</a> for Chrome.</p>\n\n<p>Once installed, click on the link to download. Wait for the dialog to appear, cancel the download.</p>\n\n<p>Click on the icon in the extension, copy the code displayed  and paste in the terminal in your AWS instance. The download should start immediately.</p>\n\n<p>I suggest to open a tmux session just to download this, and let it running in the background:\n<code>bash\ntmux new -s download #starts a session in tmux\n</code>\nCopy the code from the extension and hit Enter.\nOnce it starts downloading: \n<code>bash\n&lt;Ctrl&gt;&lt;b&gt;&lt;d&gt; # detach from the session\n</code>\nTo check once in a while how the download is going:\n<code>bash\ntmux attach -t download\n</code>\nor only <code>tmux attach</code></p>\n\n<p>I hope it helps </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 747182,
          "author_name": "robotD",
          "author_url": "",
          "post_date": "2020-02-16T04:39:37.837000",
          "content": "<p>Works like a charm!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 696833,
      "author_name": "Srinivas M Besthar",
      "author_url": "",
      "post_date": "2019-12-17T05:47:40.033000",
      "content": "<p>Hi Carlos , This is really helpful . Do you mean VM Instance setup in Google Cloud ? </p>\n\n<p>From the list of steps above \"While setting up the VM, I've chosen 1TB SSD disk storage: less than that would not allow me download and unzip the complete 472 GB full dataset at once;\"</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 696058,
      "author_name": "DiegoJohnson",
      "author_url": "",
      "post_date": "2019-12-16T04:20:58.233000",
      "content": "<p>Hi, I have no idea about cost of Google Cloud, could you give me some information about it? How much does it cost in a month to complete this competition? thanks. <a href=\"/carlossouza\">@carlossouza</a> </p>",
      "votes": 0,
      "replies": [
        {
          "id": 709074,
          "author_name": "jieming yang",
          "author_url": "",
          "post_date": "2020-01-03T02:49:00.547000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 720672,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-16T16:24:19.573000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 720671,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-16T16:22:30.557000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 708822,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-02T18:04:54.077000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 700077,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-12-21T12:25:10.560000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 696097,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-12-16T05:44:23.980000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "695263": "Here are the steps I followed to get started:\n\n**1) Setup a VM instance**\n- As I don't have local GPUs at my disposal, this was a must;\n- I sent an email to AWS requesting the $1000 in credits. While they don't answer, I got started in Google Cloud, which have cheaper &amp; better GPUs/TPUs;\n- While setting up the VM, I've chosen 1TB SSD disk storage: less than that would not allow me download and unzip the complete 472 GB full dataset at once;\n\n\n**2) Download the full dataset**\n- Downloading the full dataset was surprisingly fast: it was constantly &gt;250 MB/s;\n- I did it using *wget*:\n```\nwget --load-cookies cookies.txt https://www.kaggle.com/c/16880/datadownload/dfdc_train_all.zip\n```\n- I had to save the cookies from my local machine with Kaggle's credentials in the *cookies.txt* in the VM. Without that, it won't work;\n\n\n**3) Unzip everything**\n- The file *dfdc_train_all.zip* is like a Russian doll: it contains 49 zip files;\n- After unzipping and deleting it, I applied the following script to unzip each of the 49 files:\n```python\nimport os\nfrom tqdm import tqdm\nfrom zipfile import ZipFile\n\ndef get_zipfiles(directory):\n    list_files = []\n    for filename in os.listdir(home):\n        if filename.endswith(\".zip\"): \n            list_files.append(os.path.join(home, filename))\n    return list_files\n\nzipfiles = get_zipfiles('/home/carlossouza')\nzipfiles.sort()\n\nfor zipfile in zipfiles:\n    print(f'Extracting {zipfile}...')\n    with ZipFile(file=zipfile) as zip_file:\n        for file in tqdm(iterable=zip_file.namelist(), total=len(zip_file.namelist())):\n            zip_file.extract(member=file)\n    \n    os.remove(zipfile)\n```\n\nIf you used a different/better procedure to get started, I'd love to hear and learn ;)\nCheers\nCarlos",
    "701600": "My alternative to step 3 delineated by Carlos Above:\n\n&gt; 3) Unzip everything\n\n    The file dfdctrainall.zip is like a Russian doll: it contains 49 zip files;\n    After unzipping and deleting it, I applied the following script to unzip each of the 49 files:\n\nSince unzip is a disk/processor bottleneck the best way to use concurrency is to use all processors. ( i.e. Multiprocessing and not Multithreading )\n\n```python\nimport multiprocessing                                                                                                                                            \nfrom pathlib import Path                                                                                                                                          \nfrom time import time                                                                                                                                             \nfrom zipfile import ZipFile                                                                                                                                       \nimport logging\nfrom typing import Union                                                                                                                                                    \n                                                                                                                                                                   \nDATA = Path(\"path_to_the_49_zipped files\")                                                                                                    \nlogging.basicConfig(filename=\"extract.log\", level=logging.INFO)\nzipfiles = sorted(list(DATA.glob(\"**/*.zip\")), key=lambda x: x.stem)\n\ndef extract_zip(zipfile: Union[str, Path])-&gt;None:\n    with ZipFile(zipfile) as zip_file:\n        for file in zip_file.namelist():\n            zip_file.extract(member=file, path=zipfile.parent)\n    logging.info(f'Finished extracting {zipfile.stem}')\n\n\nstart = time()\nwith multiprocessing.Pool() as pool: # use all cores available\n    pool.map(extract_zip, zipfiles)\n\nlogging.info(f\"Extracted all zip files in {time() - start} seconds!\")\n```\n\nI didn't use `tqdm` because in multiprocessing it would show 8 tqdm results at the same time so it didn't make sense.\n\nThis created in `DATA` directory 49 folders. In order to work with all the files in a same folder I just created a folder with symlinks to each video file\n\n```python\nDATA = Path(\"path_to_the_49_zipped files\")\nTARGET = DATA.parent / \"full\"\n\ndef path_walk(top: Union[str,Path], topdown: bool = False, followlinks:bool = False):\n    \"\"\"See Python docs for os.walk, exact same behavior but it yields Path() instances instead\n    \"\"\"\n    \n    dirs = (node for node in top.iterdir() if node.is_dir())\n    nondirs =(node for node in top.iterdir() if not node.is_dir())\n\n    if topdown:\n        yield top, dirs, nondirs\n\n    for name in dirs:\n        if followlinks or name.is_symlink() is False:\n            for x in path_walk(name, topdown, followlinks):\n                yield x\n\n    if topdown is not True:\n        yield top, dirs, nondirs    \n\n#flatten the list of lists\nfiles = (f for top, dirs, files  in path_walk(DATA) for f in files if f.suffix == \".mp4\") \nfor file in files:\n    (TARGET / file.name).symlink_to(file)\n```\nSo in the `TARGET` I have all the *119146* video files to work on\n\nI hope this helps!\n\nCheers",
    "747307": "@carlossouza \nit dsnt works for me ,not sure what m i missing .Please help\n1) i generated the cookies using my login credentials \nwget -qO- --keep-session-cookies --save-cookies cookies.txt --post-data  'user=username&amp;password=password' https://www.kaggle.com/account/login?\n2) your  wget command. \n3) it downloads just the html files \n\n\n",
    "696834": "May I ask how many videos are in the dfdc_train_all folder? (not include .json file)\nI want to confirm, thank you!  ",
    "696084": "@diegojohnson ,\nRight now I’m playing with a setup that costs $0.8/hour. I’m budgeting ~10hs/week, which would give ~$8/week = $40/month.\nWhen I manage to fully utilize this setup (i.e. code optimized to run on multiple GPUs &amp; cores), and if I manage to produce competitive results, I might decide to beef up the setup and increase my budget limit.",
    "695670": "To unzip all the files at once, you can also do `unzip '*.zip'`.",
    "702450": "deserves the gold :) thanks fr quality content , upvoted!",
    "777682": "Could you please teach me how to save my kaggle credentials in cookies.txt?",
    "759721": "Very helpful! Thanks for sharing your scripts. A quick question: how do you get the data that you put inside your cookies.txt file? Using your browser cookies info I guess or something more elaborate? ",
    "739347": "Hi, I was using this [kernel](https://www.kaggle.com/frsanchez/download-kaggle-files-from-notebook) to download files from a notebook, however it worked days ago and it doesn't work anymore, I have a 404 Not Found error.\n\nI tried using the `wget` command and loading cookies but I didn't get exactly what must be into the `cookies.txt` file. I tried to put the `cURL` command as explained [here](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121194#695299) but it's not working.\n\nHave you already experienced a 404 Not Found error with the Kaggle API (given that the same code was working before)? For the `wget` command, is it the cURL command that should be into the `cookies.txt`?\n\nThanks for your help!",
    "716142": "I could not figure out the way to use kaggle api to download so wget is so far the best way...And unzipping while deleting codes...Thanks for sharing!! ",
    "701124": "Hello Carlos,\n\nWtih Wget and CurlWget the download stops after about 3%.\nSaving to: ‘dfdc_train_all.zip’\n\ndfdc_train_all.zip                 3%[=&gt;                                                      ]  17.42G  78.3MB/s    in 2m 50s  \n\n\nCannot write to ‘dfdc_train_all.zip’ (Success).\n\nAny idea?\nthank you,",
    "700098": "Can someone give md5-sums of the files?",
    "698176": "hey @carlossouza,\n\nWhen trying to download the data with wget and my own cookies everything seems ok but the download stops instantly without any error nor data downloaded (only zip file created). Any idea what's happening?\n\n\nHere is some more information : \n\nWhen running:\n```\nwget --load-cookies cookies.txt https://www.kaggle.com/c/16880/datadownload/dfdc_train_all.zip\n``` \nmy terminal says this and then nothing:\n```\n--2019-12-18 23:13:26--  https://www.kaggle.com/c/16880/datadownload/dfdc_train_all.zip\nResolving www.kaggle.com (www.kaggle.com)... \nConnecting to www.kaggle.com (www.kaggle.com)|:443... connected.\nHTTP request sent, awaiting response... 302 Found\nLocation: https://www.kaggle.com/account/login?ReturnUrl=%2Fc%2F16880%2Fdatadownload%2Fdfdc_train_all.zip [following]\n--2019-12-18 23:13:27--  https://www.kaggle.com/account/login?ReturnUrl=%2Fc%2F16880%2Fdatadownload%2Fdfdc_train_all.zip\nReusing existing connection to www.kaggle.com:443.\nHTTP request sent, awaiting response... 200 OK\nLength: unspecified [text/html]\nSaving to: ‘dfdc_train_all.zip’\n\ndfdc_train_all.zip                                     [ &lt;=&gt;                                                                                                             ]   8,96K  --.-KB/s    in 0,01s   \n\n2019-12-18 23:13:27 (940 KB/s) - ‘dfdc_train_all.zip’ saved [9170]\n```",
    "696833": "Hi Carlos , This is really helpful . Do you mean VM Instance setup in Google Cloud ? \n\nFrom the list of steps above \"While setting up the VM, I've chosen 1TB SSD disk storage: less than that would not allow me download and unzip the complete 472 GB full dataset at once;\"",
    "696058": "Hi, I have no idea about cost of Google Cloud, could you give me some information about it? How much does it cost in a month to complete this competition? thanks. @carlossouza ",
    "720672": "",
    "720671": "",
    "708822": "",
    "700077": "",
    "696097": ""
  }
}