{"cells":[{"metadata":{},"cell_type":"markdown","source":"# Number of public melanoma is 78 (or 77)\n\nThis is a simple simulation showing that the number of melanoma in public test is 78 or 77.  This result was already shared by Sirish Somanchi, based on submitting a probe where all predictions are 0 except for one known melanoma case where he submitted 1.  He got a public LB of 0.5064.  I was puzzled by this and decided to check it in this notebook.\n\nIt confirms that Sirish Somanchi was (almost) right. I say almost because 77 is also a legit number of public melanoma.","execution_count":null},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import numpy as np # linear algebra\nfrom sklearn.metrics import roc_auc_score","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Let's simulate what happens when we submit one positive when there are 78 positive in ground truth.  We do this for, say, 1000 examples.","execution_count":null},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"# 1000 samples\nn = 1000\n\n# ground truth\ny = np.zeros(n)\ny[-78:] = 1\n\n# probing prediction with one known positive set to 1, rest set to 0\nyhat = np.zeros(n)\nyhat[-1] = 1\n\n# score\nroc_auc_score(y, yhat)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"This seems right, but wait  aminute, public test rather has 3300 samples or so.  Let's see what happens with that number of samples.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# 3300 samples\nn = 3300\n\n# ground truth\ny = np.zeros(n)\ny[-78:] = 1\n\n# probing prediction with one known positive set to 1, rest set to 0\nyhat = np.zeros(n)\nyhat[-1] = 1\n\n# score\nroc_auc_score(y, yhat)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The result is unchanged!\n\nLet's see if 78 is the right number still.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# 3300 samples\nn = 3300\n\n# ground truth with 77 positive\ny = np.zeros(n)\ny[-77:] = 1\n\n# probing prediction with one known positive set to 1, rest set to 0\nyhat = np.zeros(n)\nyhat[-1] = 1\n\n# score\nroc_auc_score(y, yhat)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We see that 77 positives is also consitent with a public score truncated to 0.5064","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"# 3300 samples\nn = 3300\n\n# ground truth with 79 positive\ny = np.zeros(n)\ny[-79:] = 1\n\n# probing prediction with one known positive set to 1, rest set to 0\nyhat = np.zeros(n)\nyhat[-1] = 1\n\n# score\nroc_auc_score(y, yhat)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We see that 79 positives is not compatible with the public score.","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"Conclusion: there are either 77 or 78 positives in the public test.  ","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}