{"cells":[{"metadata":{},"cell_type":"markdown","source":"It's fork from https://www.kaggle.com/zaharch/data-leak-in-metadata @nosound\n\nThe original author has computed the data distribution, so we can compute Public LB score directly.\n\nThis leak will be repaired shortly afterwards."},{"metadata":{},"cell_type":"markdown","source":"# If you like it, please give it an upvote."},{"metadata":{"trusted":true,"_kg_hide-input":true},"cell_type":"code","source":"from math import log\n\n# loss1 - None\nloss1 = -1/4000 * (174 * log(174/1477) + (1477 - 174) * log(1 - 174/1477))\n\n\n# loss2 - 16:9\n# cd == '1/48000':\nloss2_1 = -1/4000 * (1407 * log(1407/1873) + (1873 - 1407) * log(1 - 1407/1873))\n# cd == else:\nloss2_2 = -1/4000 * (156 * log(156/206) + (206 - 156) * log(1 - 156/206))\n\n\n# loss3 - 9:16\n# cd == '1/48000':\nloss3_1 = -1/4000 * (70 * log(70/178) + (178 - 70) * log(1 - 70/178))\n# cd == else:\nloss3_2 = -1/4000 * (182 * log(182/241) + (241 - 182) * log(1 - 182/241))\n\n\n# others\nothers = -1/4000 * (11 * log(11/25) + 14 * log(14/25))\n\nscore = loss1 + loss2_1 + loss2_2 + loss3_1 + loss3_2 + others","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_kg_hide-input":false},"cell_type":"code","source":"print('Public LB score: ', int(score * 100000) / 100000)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Metadata is leaking targets"},{"metadata":{},"cell_type":"markdown","source":"This notebook uses display_aspect_ratio metadata field as a great fake video predictor."},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"_kg_hide-output":true,"collapsed":true},"cell_type":"code","source":"import pandas as pd\nimport glob\nimport os\nimport subprocess as sp\nimport tqdm.notebook as tqdm\nfrom collections import defaultdict\nimport json\n\n! tar xvf ../input/ffmpeg-static-build/ffmpeg-git-amd64-static.tar.xz","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Getting the metadata field with ffprobe"},{"metadata":{},"cell_type":"markdown","source":"The code below is from\n\nhttp://www.scikit-video.org/stable/io.html\n\nand specifically assembled from the following source files\n\nhttps://github.com/scikit-video/scikit-video/blob/master/skvideo/io/ffprobe.py\n\nhttps://github.com/scikit-video/scikit-video/blob/master/skvideo/utils/__init__.py\n\nThanks to [btk1](https://www.kaggle.com/rakibilly) for the [ffmpeg Static Build dataset](https://www.kaggle.com/rakibilly/ffmpeg-static-build)"},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"def check_output(*popenargs, **kwargs):\n    closeNULL = 0\n    try:\n        from subprocess import DEVNULL\n        closeNULL = 0\n    except ImportError:\n        import os\n        DEVNULL = open(os.devnull, 'wb')\n        closeNULL = 1\n\n    process = sp.Popen(stdout=sp.PIPE, stderr=DEVNULL, *popenargs, **kwargs)\n    output, unused_err = process.communicate()\n    retcode = process.poll()\n\n    if closeNULL:\n        DEVNULL.close()\n\n    if retcode:\n        cmd = kwargs.get(\"args\")\n        if cmd is None:\n            cmd = popenargs[0]\n        error = sp.CalledProcessError(retcode, cmd)\n        error.output = output\n        raise error\n    return output\n\ndef ffprobe(filename):\n    \n    command = [\"../working/ffmpeg-git-20191209-amd64-static/ffprobe\", \"-v\", \"error\", \"-show_streams\", \"-print_format\", \"xml\", filename]\n\n    xml = check_output(command)\n    \n    return xml\n\ndef get_markers(video_file):\n\n    xml = ffprobe(str(video_file))\n    \n    found = str(xml).find('display_aspect_ratio')\n    if found >= 0:\n        ar = str(xml)[found+22:found+26]\n    else:\n        ar = None\n        \n    found = str(xml).find('\"audio\" codec_time_base')\n    if found >= 0:\n        cd = str(xml)[found+25:found+32]\n    else:\n        cd = None\n    \n    return ar, cd","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"video_file = '/kaggle/input/deepfake-detection-challenge/test_videos/gunamloolc.mp4'\nget_markers(video_file)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Motivation"},{"metadata":{"trusted":true},"cell_type":"code","source":"filenames = glob.glob('/kaggle/input/deepfake-detection-challenge/train_sample_videos/*.mp4')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"my_dict = defaultdict()\nfor filename in tqdm.tqdm(filenames):\n    fn = filename.split('/')[-1]\n    ar, cd = get_markers(filename)\n    my_dict[fn] = ar","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"display_aspect_ratios = pd.DataFrame.from_dict(my_dict, orient='index')\ndisplay_aspect_ratios.columns = ['display_aspect_ratio']\ndisplay_aspect_ratios = display_aspect_ratios.fillna('NONE')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"labels = json.load(open('/kaggle/input/deepfake-detection-challenge/train_sample_videos/metadata.json', encoding=\"utf8\"))\n\nlabels = pd.DataFrame(labels).transpose()\nlabels = labels.reset_index()\nlabels = labels.join(display_aspect_ratios, on='index')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"labels.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"pd.crosstab(labels.display_aspect_ratio, labels.label)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Make predictions"},{"metadata":{"trusted":true},"cell_type":"code","source":"filenames = glob.glob('/kaggle/input/deepfake-detection-challenge/test_videos/*.mp4')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sub = pd.read_csv('/kaggle/input/deepfake-detection-challenge/sample_submission.csv')\nsub.label = 11/25\nsub = sub.set_index('filename',drop=False)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"In the public test datatset it is not strictly true that if display_aspect_ratio field is missing it is a real video, and if it equals 16:9 it is a fake. There are some exceptions for whatever reason, but not too many. By selecting a proper threshold below we can decrease log-loss substantially. The thresholds are selected by Gurobi solver running on available submissions.\n\nIn more details, there are 4 groups that I am looking at, depending on the values of display_aspect_ratio, see below. In each of the groups there are [1303,  516,  167,   14] real samples and [ 174, 1563,  252,   11] fakes samples. For each group I select probability of a fake for that group, - this is the value that minimizes log-loss. How did I get the numbers for each group? I had some submissions with scores already, where I put different values for those groups. With these constraints it is possible to find the numbers, even manually. But manually is a little bit tiresome, so I wrote a mixed integer programming formulation for that problem, and used Gurobi to solve."},{"metadata":{"trusted":true},"cell_type":"code","source":"for filename in tqdm.tqdm(filenames):\n    \n    fn = filename.split('/')[-1]\n    ar, cd = get_markers(filename)\n    \n    if ar is None:\n        sub.loc[fn, 'label'] = 174/1477\n    if cd == '1/48000':\n        if ar == '16:9':\n            sub.loc[fn, 'label'] = 1407/1873\n        if ar == '9:16':\n            sub.loc[fn, 'label'] = 70/178\n    else:\n        if ar == '16:9':\n            sub.loc[fn, 'label'] = 156/206\n        if ar == '9:16':\n            sub.loc[fn, 'label'] = 182/241","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sub.label.value_counts()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sub.to_csv('submission.csv', index=False)","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":1}