{"cells":[{"metadata":{},"cell_type":"markdown","source":"# Does the private test set have the same image size distribution as the public test set?\nAccording to [this really cool kernel](https://www.kaggle.com/fhopfmueller/removing-unwanted-correlations-in-training-public), 72.770% of the public test images has the size of 640x480. Is this ratio the same in the private test set? The answer to this question would give us some clues about how much difference there is between the public test set and the private test set. In this kernel, inspired by [this discussion](https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/97652#latest-596913) and [this kernel](https://www.kaggle.com/cdeotte/private-lb-probing-0-950), the image size distribution of public + private test set will be probed.\n\n# Approach\nLet's assume that we have N-(submission.csv, public LB score) pairs from already submitted kernels. By changing the submission file according to the desired information about the private test set, we can get log(N) bits information from one submission. Here \"change submission\" means to copy the targets of the public test set from the known submissions and insert arbitrary dummy targets to the remaining private test set. By doing so, we can control the public LB score according to what we want to know.\n\n# Answer\nThe answer is no. While the ratio of 640x480 images in the public test set is 72.770%, the ratio of 640x480 images in the public + private test set is 30-40%. There are some differences between the public test set and the private test set. I hope this information would be useful in choosing the final submission(s)."},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"from pathlib import Path\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport cv2","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# private + public test\ntest_csv_path = \"../input/aptos2019-blindness-detection/test.csv\"\ndf = pd.read_csv(test_csv_path)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Let's count the number of images with the size of 640x480, and calculate the ratio of them. `target_ratio_code 0` indicates 0-10% ratio, `target_ratio_code 1` indicates 10-20% ratio, and so forth. In case of the public test set, as the ratio is 0.72770, `target_ratio_code` is 7."},{"metadata":{"trusted":true},"cell_type":"code","source":"test_image_dir = Path(\"../input/aptos2019-blindness-detection/test_images\")\nprivate_img_cnt = 0\ntarget_img_cnt = 0\n\nfor _, row in df.iterrows():\n    id_code = row[\"id_code\"]\n    img_path = test_image_dir.joinpath(f\"{id_code}.png\")\n    img = cv2.imread(str(img_path), 1)\n    h, w, _ = img.shape\n\n    if w == 640 and h == 480:\n        target_img_cnt += 1\n        \n    private_img_cnt += 1\n    \ntarget_ratio = target_img_cnt / private_img_cnt\ntarget_ratio_code = int(target_ratio * 10)\nprint(target_ratio)\nprint(target_ratio_code)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"I have ten submissions with known public LB scores. Let's select a submission file according to `target_ratio_code` and copy the targets to the submission of this kernel."},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"submissions = [\n    \"submission_0.683.csv\",\n    \"submission_0.694.csv\",\n    \"submission_0.709.csv\",\n    \"submission_0.711.csv\",\n    \"submission_0.739.csv\",\n    \"submission_0.751.csv\",\n    \"submission_0.755.csv\",\n    \"submission_0.766.csv\",\n    \"submission_0.768.csv\",\n    \"submission_0.785.csv\"\n]\n\n# select a submission file according to the ratio of target image size count.\npub_test_csv_path = \"../input/aptos10submissions/\" + submissions[target_ratio_code]\npub_df = pd.read_csv(pub_test_csv_path)\nid_to_diagnosis = {id_code: diag for id_code, diag in zip(pub_df.id_code, pub_df.diagnosis)}","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"all_diagnosis = []\n\nfor _, row in df.iterrows():\n    id_code = row[\"id_code\"]\n    \n    if id_code in id_to_diagnosis:\n        all_diagnosis.append(id_to_diagnosis[id_code])\n    else:\n        all_diagnosis.append(0)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"id_codes = df.id_code.values\nnew_df = pd.DataFrame.from_dict(data={\"id_code\": df.id_code.values, \"diagnosis\": all_diagnosis})\nnew_df.to_csv(\"submission.csv\", index=False)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Alright, what is the public LB score? I got the score of 0.711, which indicates that `target_ratio_code` of the public + private test set is 3; the ratio is 30-40%. This is a little bit different from that of the public test set. How do you think this?"}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":1}