{
  "id": 124590,
  "title": "Someone successfully submitted an audio classifier？get Submission Scoring Error when I extract audios on test set",
  "url": "/competitions/deepfake-detection-challenge/discussion/124590",
  "author_name": "",
  "post_date": "2020-01-05T05:58:43.839001500Z",
  "votes": null,
  "comment_count": 10,
  "views": 0,
  "content": "<p>When I submit an audio classifier, the <code>Submission Scoring Error</code> returned for many times...</p>\n\n<p>So I use simply use <code>ffmpeg</code> to extract audios on test set. But <code>Submission Scoring Error</code> still exist...</p>\n\n<p>There is my code, I cant find anywhere cause the 'Submission Scoring Error'</p>\n\n<p>```python\nimport soundfile as sf\nimport torch\nimport math\nimport numpy as np\nimport os\nfrom pathlib import Path\nimport subprocess\nimport pandas as pd</p>\n\n<p>output_dir = Path(f\"wavs\")\nPath(output_dir).mkdir(exist_ok=True, parents=True)\nsubmit_ = []\nfor vi in os.listdir('/kaggle/input/deepfake-detection-challenge/test_videos'):\n    sub_list = [vi,0.5]\n    video_path = os.path.join(\"/kaggle/input/deepfake-detection-challenge/test_videos\", vi)\n    command = f\"../working/ffmpeg-git-20191209-amd64-static/ffmpeg -i {video_path} -ab 192000 -ac 2 -ar 44100 -vn {output_dir/vi[:-4]}.wav\"\n    subprocess.call(command, shell=True)\n    if not os.path.exists(os.path.join(output_dir, vi[:-4] + '.wav')):\n        submit_.append(sub_list)\n        continue\n    submit_.append(sub_list)</p>\n\n<p>submission = pd.DataFrame(submit_, columns=['filename', 'label'])\nsubmission.sort_values('filename').to_csv('submission.csv', index=False)\n```</p>\n\n<p>Wish you can give me some advices.</p>",
  "messages": [
    {
      "id": "710721",
      "postDate": "01/05/2020 05:58:43",
      "content": "<p>When I submit an audio classifier, the <code>Submission Scoring Error</code> returned for many times...</p>\n\n<p>So I use simply use <code>ffmpeg</code> to extract audios on test set. But <code>Submission Scoring Error</code> still exist...</p>\n\n<p>There is my code, I cant find anywhere cause the 'Submission Scoring Error'</p>\n\n<p>```python\nimport soundfile as sf\nimport torch\nimport math\nimport numpy as np\nimport os\nfrom pathlib import Path\nimport subprocess\nimport pandas as pd</p>\n\n<p>output_dir = Path(f\"wavs\")\nPath(output_dir).mkdir(exist_ok=True, parents=True)\nsubmit_ = []\nfor vi in os.listdir('/kaggle/input/deepfake-detection-challenge/test_videos'):\n    sub_list = [vi,0.5]\n    video_path = os.path.join(\"/kaggle/input/deepfake-detection-challenge/test_videos\", vi)\n    command = f\"../working/ffmpeg-git-20191209-amd64-static/ffmpeg -i {video_path} -ab 192000 -ac 2 -ar 44100 -vn {output_dir/vi[:-4]}.wav\"\n    subprocess.call(command, shell=True)\n    if not os.path.exists(os.path.join(output_dir, vi[:-4] + '.wav')):\n        submit_.append(sub_list)\n        continue\n    submit_.append(sub_list)</p>\n\n<p>submission = pd.DataFrame(submit_, columns=['filename', 'label'])\nsubmission.sort_values('filename').to_csv('submission.csv', index=False)\n```</p>\n\n<p>Wish you can give me some advices.</p>",
      "rawMarkdown": "When I submit an audio classifier, the `Submission Scoring Error` returned for many times...\n\nSo I use simply use `ffmpeg` to extract audios on test set. But `Submission Scoring Error` still exist...\n\nThere is my code, I cant find anywhere cause the 'Submission Scoring Error'\n\n```python\nimport soundfile as sf\nimport torch\nimport math\nimport numpy as np\nimport os\nfrom pathlib import Path\nimport subprocess\nimport pandas as pd\n\noutput_dir = Path(f\"wavs\")\nPath(output_dir).mkdir(exist_ok=True, parents=True)\nsubmit_ = []\nfor vi in os.listdir('/kaggle/input/deepfake-detection-challenge/test_videos'):\n    sub_list = [vi,0.5]\n    video_path = os.path.join(\"/kaggle/input/deepfake-detection-challenge/test_videos\", vi)\n    command = f\"../working/ffmpeg-git-20191209-amd64-static/ffmpeg -i {video_path} -ab 192000 -ac 2 -ar 44100 -vn {output_dir/vi[:-4]}.wav\"\n    subprocess.call(command, shell=True)\n    if not os.path.exists(os.path.join(output_dir, vi[:-4] + '.wav')):\n        submit_.append(sub_list)\n        continue\n    submit_.append(sub_list)\n\nsubmission = pd.DataFrame(submit_, columns=['filename', 'label'])\nsubmission.sort_values('filename').to_csv('submission.csv', index=False)\n```\n\nWish you can give me some advices.",
      "votes": null
    },
    {
      "id": "710748",
      "postDate": "01/05/2020 06:56:56",
      "content": "<p>did you load ffmpg and just not show us that in your code?</p>\n\n<p>Easier for folks to help if you have a public kernel with all the code so I don't have to ask what might be this dumb question</p>\n\n<p>! tar xvf ../input/ffmpeg-static-build/ffmpeg-git-amd64-static.tar.xz</p>",
      "rawMarkdown": "did you load ffmpg and just not show us that in your code?\n\nEasier for folks to help if you have a public kernel with all the code so I don't have to ask what might be this dumb question\n\n! tar xvf ../input/ffmpeg-static-build/ffmpeg-git-amd64-static.tar.xz",
      "votes": null
    },
    {
      "id": "710757",
      "postDate": "01/05/2020 07:12:13",
      "content": "<p>public kernel is at: <a href=\"https://www.kaggle.com/xiaofengmao/kernel553c006564/notebook\">https://www.kaggle.com/xiaofengmao/kernel553c006564/notebook</a>\nAs my code shows, I just define the model, and not use any model to give a prediction. Instead I use 0.5 as my submission score.\nSo this report error <code>Submission Scoring Error</code> puzzled me a lot. I cannot find anywhere to report this error. </p>",
      "rawMarkdown": "public kernel is at: https://www.kaggle.com/xiaofengmao/kernel553c006564/notebook\nAs my code shows, I just define the model, and not use any model to give a prediction. Instead I use 0.5 as my submission score.\nSo this report error `Submission Scoring Error` puzzled me a lot. I cannot find anywhere to report this error.",
      "votes": null
    },
    {
      "id": "710762",
      "postDate": "01/05/2020 07:20:07",
      "content": "<p>Hmmm - can't run a fork to try as the first three cells are inputs that don't exist on my fork.  Don't see any obvious.</p>",
      "rawMarkdown": "Hmmm - can't run a fork to try as the first three cells are inputs that don't exist on my fork.  Don't see any obvious.",
      "votes": null
    },
    {
      "id": "710763",
      "postDate": "01/05/2020 07:30:38",
      "content": "<p>Sorry that they are from my private dataset. \nTry this clean kernel:\n<a href=\"https://www.kaggle.com/xiaofengmao/kernel553c006564\">https://www.kaggle.com/xiaofengmao/kernel553c006564</a>\nI delete these dataset here.</p>",
      "rawMarkdown": "Sorry that they are from my private dataset. \nTry this clean kernel:\nhttps://www.kaggle.com/xiaofengmao/kernel553c006564\nI delete these dataset here.",
      "votes": null
    },
    {
      "id": "711268",
      "postDate": "01/05/2020 21:32:55",
      "content": "<p>OK - it ran ok on the commit.   Just submitted now - guessing it might be an hour or so before I know if success.  One observation that I want to check out - a lot of files get written during the process - I wonder if the 10X increase in the files being evaluated during the submission leads to using more HDD than we are allowed.</p>",
      "rawMarkdown": "OK - it ran ok on the commit.   Just submitted now - guessing it might be an hour or so before I know if success.  One observation that I want to check out - a lot of files get written during the process - I wonder if the 10X increase in the files being evaluated during the submission leads to using more HDD than we are allowed.",
      "votes": null
    },
    {
      "id": "711269",
      "postDate": "01/05/2020 21:34:49",
      "content": "<p>Got your same error - looks like error occurred about 6 minutes in.  </p>",
      "rawMarkdown": "Got your same error - looks like error occurred about 6 minutes in.",
      "votes": null
    },
    {
      "id": "711271",
      "postDate": "01/05/2020 21:41:49",
      "content": "<p>Size of written 158 mb - so 10X might be 1580 mb - think we are allowed 5GB during runs so this not it.</p>",
      "rawMarkdown": "Size of written 158 mb - so 10X might be 1580 mb - think we are allowed 5GB during runs so this not it.",
      "votes": null
    },
    {
      "id": "711274",
      "postDate": "01/05/2020 21:46:33",
      "content": "<p>There is also a limit on number of output files, don't remember exactly, something like 500. Try cleaning with <code>shutil.rmtree()</code></p>",
      "rawMarkdown": "There is also a limit on number of output files, don't remember exactly, something like 500. Try cleaning with `shutil.rmtree()`",
      "votes": null
    },
    {
      "id": "711386",
      "postDate": "01/06/2020 02:41:46",
      "content": "<p>Thank you <a href=\"/zaharch\">@zaharch</a> <a href=\"/pcjimmmy\">@pcjimmmy</a> \nI add <code>os.remove()</code> at end of the loop to clean generated <code>.wav</code>. Now it works ! There exactly have a limitation on number of output files. I forget to clean them during my kernel running. </p>",
      "rawMarkdown": "Thank you @zaharch @pcjimmmy \nI add `os.remove()` at end of the loop to clean generated `.wav`. Now it works ! There exactly have a limitation on number of output files. I forget to clean them during my kernel running.",
      "votes": null
    },
    {
      "id": "711848",
      "postDate": "01/06/2020 15:21:39",
      "content": "<p>Glad you got it - nice to learn about the limit on output files !  </p>\n\n<p>Just hit submit myself - interested in how long this loop of yours takes on the full test.</p>",
      "rawMarkdown": "Glad you got it - nice to learn about the limit on output files !  \n\nJust hit submit myself - interested in how long this loop of yours takes on the full test.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 710748,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "01/05/2020 06:56:56",
      "content": "<p>did you load ffmpg and just not show us that in your code?</p>\n\n<p>Easier for folks to help if you have a public kernel with all the code so I don't have to ask what might be this dumb question</p>\n\n<p>! tar xvf ../input/ffmpeg-static-build/ffmpeg-git-amd64-static.tar.xz</p>",
      "votes": null,
      "replies": [
        {
          "id": 710757,
          "author_name": "xiaofengmao",
          "author_url": "",
          "post_date": "01/05/2020 07:12:13",
          "content": "<p>public kernel is at: <a href=\"https://www.kaggle.com/xiaofengmao/kernel553c006564/notebook\">https://www.kaggle.com/xiaofengmao/kernel553c006564/notebook</a>\nAs my code shows, I just define the model, and not use any model to give a prediction. Instead I use 0.5 as my submission score.\nSo this report error <code>Submission Scoring Error</code> puzzled me a lot. I cannot find anywhere to report this error. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 710762,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "01/05/2020 07:20:07",
          "content": "<p>Hmmm - can't run a fork to try as the first three cells are inputs that don't exist on my fork.  Don't see any obvious.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 710763,
          "author_name": "xiaofengmao",
          "author_url": "",
          "post_date": "01/05/2020 07:30:38",
          "content": "<p>Sorry that they are from my private dataset. \nTry this clean kernel:\n<a href=\"https://www.kaggle.com/xiaofengmao/kernel553c006564\">https://www.kaggle.com/xiaofengmao/kernel553c006564</a>\nI delete these dataset here.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 711268,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "01/05/2020 21:32:55",
          "content": "<p>OK - it ran ok on the commit.   Just submitted now - guessing it might be an hour or so before I know if success.  One observation that I want to check out - a lot of files get written during the process - I wonder if the 10X increase in the files being evaluated during the submission leads to using more HDD than we are allowed.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 711269,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "01/05/2020 21:34:49",
          "content": "<p>Got your same error - looks like error occurred about 6 minutes in.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 711271,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "01/05/2020 21:41:49",
          "content": "<p>Size of written 158 mb - so 10X might be 1580 mb - think we are allowed 5GB during runs so this not it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 711274,
          "author_name": "zaharch",
          "author_url": "",
          "post_date": "01/05/2020 21:46:33",
          "content": "<p>There is also a limit on number of output files, don't remember exactly, something like 500. Try cleaning with <code>shutil.rmtree()</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 711386,
          "author_name": "xiaofengmao",
          "author_url": "",
          "post_date": "01/06/2020 02:41:46",
          "content": "<p>Thank you <a href=\"/zaharch\">@zaharch</a> <a href=\"/pcjimmmy\">@pcjimmmy</a> \nI add <code>os.remove()</code> at end of the loop to clean generated <code>.wav</code>. Now it works ! There exactly have a limitation on number of output files. I forget to clean them during my kernel running. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 711848,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "01/06/2020 15:21:39",
          "content": "<p>Glad you got it - nice to learn about the limit on output files !  </p>\n\n<p>Just hit submit myself - interested in how long this loop of yours takes on the full test.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "710721": "When I submit an audio classifier, the `Submission Scoring Error` returned for many times...\n\nSo I use simply use `ffmpeg` to extract audios on test set. But `Submission Scoring Error` still exist...\n\nThere is my code, I cant find anywhere cause the 'Submission Scoring Error'\n\n```python\nimport soundfile as sf\nimport torch\nimport math\nimport numpy as np\nimport os\nfrom pathlib import Path\nimport subprocess\nimport pandas as pd\n\noutput_dir = Path(f\"wavs\")\nPath(output_dir).mkdir(exist_ok=True, parents=True)\nsubmit_ = []\nfor vi in os.listdir('/kaggle/input/deepfake-detection-challenge/test_videos'):\n    sub_list = [vi,0.5]\n    video_path = os.path.join(\"/kaggle/input/deepfake-detection-challenge/test_videos\", vi)\n    command = f\"../working/ffmpeg-git-20191209-amd64-static/ffmpeg -i {video_path} -ab 192000 -ac 2 -ar 44100 -vn {output_dir/vi[:-4]}.wav\"\n    subprocess.call(command, shell=True)\n    if not os.path.exists(os.path.join(output_dir, vi[:-4] + '.wav')):\n        submit_.append(sub_list)\n        continue\n    submit_.append(sub_list)\n\nsubmission = pd.DataFrame(submit_, columns=['filename', 'label'])\nsubmission.sort_values('filename').to_csv('submission.csv', index=False)\n```\n\nWish you can give me some advices.",
    "710748": "did you load ffmpg and just not show us that in your code?\n\nEasier for folks to help if you have a public kernel with all the code so I don't have to ask what might be this dumb question\n\n! tar xvf ../input/ffmpeg-static-build/ffmpeg-git-amd64-static.tar.xz",
    "710757": "public kernel is at: https://www.kaggle.com/xiaofengmao/kernel553c006564/notebook\nAs my code shows, I just define the model, and not use any model to give a prediction. Instead I use 0.5 as my submission score.\nSo this report error `Submission Scoring Error` puzzled me a lot. I cannot find anywhere to report this error.",
    "710762": "Hmmm - can't run a fork to try as the first three cells are inputs that don't exist on my fork.  Don't see any obvious.",
    "710763": "Sorry that they are from my private dataset. \nTry this clean kernel:\nhttps://www.kaggle.com/xiaofengmao/kernel553c006564\nI delete these dataset here.",
    "711268": "OK - it ran ok on the commit.   Just submitted now - guessing it might be an hour or so before I know if success.  One observation that I want to check out - a lot of files get written during the process - I wonder if the 10X increase in the files being evaluated during the submission leads to using more HDD than we are allowed.",
    "711269": "Got your same error - looks like error occurred about 6 minutes in.",
    "711271": "Size of written 158 mb - so 10X might be 1580 mb - think we are allowed 5GB during runs so this not it.",
    "711274": "There is also a limit on number of output files, don't remember exactly, something like 500. Try cleaning with `shutil.rmtree()`",
    "711386": "Thank you @zaharch @pcjimmmy \nI add `os.remove()` at end of the loop to clean generated `.wav`. Now it works ! There exactly have a limitation on number of output files. I forget to clean them during my kernel running.",
    "711848": "Glad you got it - nice to learn about the limit on output files !  \n\nJust hit submit myself - interested in how long this loop of yours takes on the full test."
  },
  "source": "meta"
}