{
  "id": 134485,
  "title": "More timeout issues",
  "url": "/competitions/deepfake-detection-challenge/discussion/134485",
  "author_name": "",
  "post_date": "2020-03-08T13:14:21.992062800Z",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Kernel runs in 25 minutes on 400 videos - should scale up to a little over 4 hours on 4,000 videos. My workflow entails extracting faces, serializing them to disk and then predicting them. I had to do it this way to avoid GPU memory issues between my face extraction model in Pytorch and my prediction model in Tensorflow.</p>\n\n<p>I have error handling in place for the face extraction and serialization.</p>\n\n<p>How can I make sure that the GPU is on for the submission?</p>\n\n<p>Are the limits for working disk space the same as for working kernels at 5 GB? I calculate less than 2GB of usage.</p>\n\n<p>I'm deleting all working files except submission.csv at the end of my process, but it is leaving one .pyc file. What are the requirements to clean up working directory?</p>\n\n<p>Any other suggestions on how to debug would be greatly appreciated.</p>",
  "messages": [
    {
      "id": "766620",
      "postDate": "03/08/2020 13:14:21",
      "content": "<p>Kernel runs in 25 minutes on 400 videos - should scale up to a little over 4 hours on 4,000 videos. My workflow entails extracting faces, serializing them to disk and then predicting them. I had to do it this way to avoid GPU memory issues between my face extraction model in Pytorch and my prediction model in Tensorflow.</p>\n\n<p>I have error handling in place for the face extraction and serialization.</p>\n\n<p>How can I make sure that the GPU is on for the submission?</p>\n\n<p>Are the limits for working disk space the same as for working kernels at 5 GB? I calculate less than 2GB of usage.</p>\n\n<p>I'm deleting all working files except submission.csv at the end of my process, but it is leaving one .pyc file. What are the requirements to clean up working directory?</p>\n\n<p>Any other suggestions on how to debug would be greatly appreciated.</p>",
      "rawMarkdown": "Kernel runs in 25 minutes on 400 videos - should scale up to a little over 4 hours on 4,000 videos. My workflow entails extracting faces, serializing them to disk and then predicting them. I had to do it this way to avoid GPU memory issues between my face extraction model in Pytorch and my prediction model in Tensorflow.\n\nI have error handling in place for the face extraction and serialization.\n\nHow can I make sure that the GPU is on for the submission?\n\nAre the limits for working disk space the same as for working kernels at 5 GB? I calculate less than 2GB of usage.\n\nI'm deleting all working files except submission.csv at the end of my process, but it is leaving one .pyc file. What are the requirements to clean up working directory?\n\nAny other suggestions on how to debug would be greatly appreciated.",
      "votes": null
    },
    {
      "id": "766935",
      "postDate": "03/09/2020 00:30:34",
      "content": "<p>Be careful about video resolutions(&gt;1920x1080) and number of faces(&gt;3) per frame. They might be large then the training set.</p>",
      "rawMarkdown": "Be careful about video resolutions(&gt;1920x1080) and number of faces(&gt;3) per frame. They might be large then the training set.",
      "votes": null
    },
    {
      "id": "766980",
      "postDate": "03/09/2020 02:12:17",
      "content": "<p>Thanks, I was thinking that if the GPU ran out of memory, I would get an error instead of timing out.</p>",
      "rawMarkdown": "Thanks, I was thinking that if the GPU ran out of memory, I would get an error instead of timing out.",
      "votes": null
    },
    {
      "id": "775596",
      "postDate": "03/16/2020 21:06:12",
      "content": "<p>Have you figured out what the reason is? Your workflow seems similar to ours and the timeout occurred to us as well. I also tried to do the inference on only 400 videos (set 0.5 for others) but still \"notebook timeout\". Just wondering if you have got any clues.</p>",
      "rawMarkdown": "Have you figured out what the reason is? Your workflow seems similar to ours and the timeout occurred to us as well. I also tried to do the inference on only 400 videos (set 0.5 for others) but still \"notebook timeout\". Just wondering if you have got any clues.",
      "votes": null
    },
    {
      "id": "775816",
      "postDate": "03/17/2020 01:22:23",
      "content": "<p>I noticed yesterday in my submission that mine times out. Upon running the script multiple times on the set of 400, it seems to hang during a function call that is called via thread pool. It only seems to happen about 1 in 5 times so I didn't notice it the first time I ran it. The example below is the thread pool method I use.</p>\n\n<p><code>\nfrom concurrent.futures import ThreadPoolExecutor\nwith ThreadPoolExecutor(max_workers=3) as ex:\n        res = ex.map(process_vid, test_vid_paths)\n</code></p>\n\n<p>It does not hang when I don't use the thread pool but I havn't submitted it yet. I'm not sure if this is just an issue on my end or a global problem. But I thought I'd mention it in case your issue is similar to mine.</p>",
      "rawMarkdown": "I noticed yesterday in my submission that mine times out. Upon running the script multiple times on the set of 400, it seems to hang during a function call that is called via thread pool. It only seems to happen about 1 in 5 times so I didn't notice it the first time I ran it. The example below is the thread pool method I use.\n\n```\nfrom concurrent.futures import ThreadPoolExecutor\nwith ThreadPoolExecutor(max_workers=3) as ex:\n        res = ex.map(process_vid, test_vid_paths)\n```\n\nIt does not hang when I don't use the thread pool but I havn't submitted it yet. I'm not sure if this is just an issue on my end or a global problem. But I thought I'd mention it in case your issue is similar to mine.",
      "votes": null
    },
    {
      "id": "783049",
      "postDate": "03/22/2020 23:35:45",
      "content": "<p>You got top place within 7 days, it's impressive 0.0</p>",
      "rawMarkdown": "You got top place within 7 days, it's impressive 0.0",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 766935,
      "author_name": "wufanyou",
      "author_url": "",
      "post_date": "03/09/2020 00:30:34",
      "content": "<p>Be careful about video resolutions(&gt;1920x1080) and number of faces(&gt;3) per frame. They might be large then the training set.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 766980,
      "author_name": "calebeverett",
      "author_url": "",
      "post_date": "03/09/2020 02:12:17",
      "content": "<p>Thanks, I was thinking that if the GPU ran out of memory, I would get an error instead of timing out.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 775596,
      "author_name": "lfcpeng17",
      "author_url": "",
      "post_date": "03/16/2020 21:06:12",
      "content": "<p>Have you figured out what the reason is? Your workflow seems similar to ours and the timeout occurred to us as well. I also tried to do the inference on only 400 videos (set 0.5 for others) but still \"notebook timeout\". Just wondering if you have got any clues.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 775816,
      "author_name": "davidmilam",
      "author_url": "",
      "post_date": "03/17/2020 01:22:23",
      "content": "<p>I noticed yesterday in my submission that mine times out. Upon running the script multiple times on the set of 400, it seems to hang during a function call that is called via thread pool. It only seems to happen about 1 in 5 times so I didn't notice it the first time I ran it. The example below is the thread pool method I use.</p>\n\n<p><code>\nfrom concurrent.futures import ThreadPoolExecutor\nwith ThreadPoolExecutor(max_workers=3) as ex:\n        res = ex.map(process_vid, test_vid_paths)\n</code></p>\n\n<p>It does not hang when I don't use the thread pool but I havn't submitted it yet. I'm not sure if this is just an issue on my end or a global problem. But I thought I'd mention it in case your issue is similar to mine.</p>",
      "votes": null,
      "replies": [
        {
          "id": 783049,
          "author_name": "yuanzhezhou",
          "author_url": "",
          "post_date": "03/22/2020 23:35:45",
          "content": "<p>You got top place within 7 days, it's impressive 0.0</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "766620": "Kernel runs in 25 minutes on 400 videos - should scale up to a little over 4 hours on 4,000 videos. My workflow entails extracting faces, serializing them to disk and then predicting them. I had to do it this way to avoid GPU memory issues between my face extraction model in Pytorch and my prediction model in Tensorflow.\n\nI have error handling in place for the face extraction and serialization.\n\nHow can I make sure that the GPU is on for the submission?\n\nAre the limits for working disk space the same as for working kernels at 5 GB? I calculate less than 2GB of usage.\n\nI'm deleting all working files except submission.csv at the end of my process, but it is leaving one .pyc file. What are the requirements to clean up working directory?\n\nAny other suggestions on how to debug would be greatly appreciated.",
    "766935": "Be careful about video resolutions(&gt;1920x1080) and number of faces(&gt;3) per frame. They might be large then the training set.",
    "766980": "Thanks, I was thinking that if the GPU ran out of memory, I would get an error instead of timing out.",
    "775596": "Have you figured out what the reason is? Your workflow seems similar to ours and the timeout occurred to us as well. I also tried to do the inference on only 400 videos (set 0.5 for others) but still \"notebook timeout\". Just wondering if you have got any clues.",
    "775816": "I noticed yesterday in my submission that mine times out. Upon running the script multiple times on the set of 400, it seems to hang during a function call that is called via thread pool. It only seems to happen about 1 in 5 times so I didn't notice it the first time I ran it. The example below is the thread pool method I use.\n\n```\nfrom concurrent.futures import ThreadPoolExecutor\nwith ThreadPoolExecutor(max_workers=3) as ex:\n        res = ex.map(process_vid, test_vid_paths)\n```\n\nIt does not hang when I don't use the thread pool but I havn't submitted it yet. I'm not sure if this is just an issue on my end or a global problem. But I thought I'd mention it in case your issue is similar to mine.",
    "783049": "You got top place within 7 days, it's impressive 0.0"
  },
  "source": "meta"
}