{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"### UltraMNIST baseline: All class 20\n\nThis is a simple test to take a look at the distribution of the [UltraMNIST Classification Challenge](https://www.kaggle.com/c/ultra-mnist/) test data. There are 28 classes $[0,27]$ and if there are an equal number of \ntest images belonging to each class then this test should have an expectation value of $ 1/28 \\approx 0.03571 $. However, the class of each image corresponds to the sum of a set of 3-5 digits. In the topic [\"Digits sum analysis for 3, 4, 5 digits\"](https://www.kaggle.com/c/ultra-mnist/discussion/312702) by towzeur it can be seen that, for example, there are only three sets that belong to the class \"1\", those sets being $\\{0,0,1\\}$, $\\{0,0,0,1\\}$ and $\\{0,0,0,0,1\\}$, whereas there are 147 sets that belong to the class \"20\". In this competition there a total of 2365 possible sets. If there are an equal number of \ntest images corresponding to each set, rather than belonging to each class, then this test should have an expectation value of $ 147/2365 \\approx 0.06215 $.\n\nA baseline score of $ \\approx 0.03571 $ would indicate that test data has a similar composition to the training data. In the training data there are 1000 images associated with each class. This would imply that, say for the class \"1\" which has three sets, on average there are around $ 1000/3 \\approx 333 $ training images per set. However, for the class \"20\" there are only around $ 1000/147 \\approx 6.8 $ training images per set.","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"code","source":"import pandas as pd\nsample_submission = pd.read_csv(\"../input/ultra-mnist/sample_submission.csv\")\nsample_submission[\"digit_sum\"] = 20\nsample_submission.to_csv('submission.csv',index=False)","metadata":{"execution":{"iopub.status.busy":"2022-03-15T17:41:47.920871Z","iopub.execute_input":"2022-03-15T17:41:47.921252Z","iopub.status.idle":"2022-03-15T17:41:48.013153Z","shell.execute_reply.started":"2022-03-15T17:41:47.921217Z","shell.execute_reply":"2022-03-15T17:41:48.011967Z"},"trusted":true},"execution_count":null,"outputs":[]}]}