{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Baseline: Preserve Train Proportions\n\nIt is always important to establish a simple baseline before building more sophisticated models. This way you have a reference to assess the true value added of your complex models. Often, it's hard to outperform these simple baseline models or the added complexity of other models is hard to justify.\n\nSo lets establish a simple baseline model that does random predictions that preserve the proportion of class predictions from the training data. This is of course not a very useful model but gives us something to compare to.","metadata":{}},{"cell_type":"markdown","source":"# Imports","metadata":{}},{"cell_type":"code","source":"import random\nimport pandas as pd\nimport seaborn as sns\nimport jo_wilder","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-02-13T18:34:22.012022Z","iopub.execute_input":"2023-02-13T18:34:22.012454Z","iopub.status.idle":"2023-02-13T18:34:23.106763Z","shell.execute_reply.started":"2023-02-13T18:34:22.012369Z","shell.execute_reply":"2023-02-13T18:34:23.105183Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load data\nHere we load the training label data so we can see what the most common output class.","metadata":{}},{"cell_type":"code","source":"# load training lables\ntrain_df = pd.read_csv(\"/kaggle/input/predict-student-performance-from-game-play/train_labels.csv\")\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-02-13T18:34:23.108477Z","iopub.execute_input":"2023-02-13T18:34:23.109596Z","iopub.status.idle":"2023-02-13T18:34:23.393903Z","shell.execute_reply.started":"2023-02-13T18:34:23.109558Z","shell.execute_reply":"2023-02-13T18:34:23.393071Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Class proportion\nHere we have a look to see what the most common class (whether the user answered the quistion correctly). The assumption is that the training data is representative of the test data as well. If this holds true, there would be a similar proportion of this class in the test data.","metadata":{}},{"cell_type":"code","source":"# quick check to see if we have any null values\ntrain_df.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2023-02-13T18:34:23.397771Z","iopub.execute_input":"2023-02-13T18:34:23.399845Z","iopub.status.idle":"2023-02-13T18:34:23.421938Z","shell.execute_reply.started":"2023-02-13T18:34:23.399807Z","shell.execute_reply":"2023-02-13T18:34:23.421099Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# check class proportions\nclass_prop = train_df[\"correct\"].value_counts() / train_df.shape[0]\nclass_prop","metadata":{"execution":{"iopub.status.busy":"2023-02-13T18:34:23.426548Z","iopub.execute_input":"2023-02-13T18:34:23.428688Z","iopub.status.idle":"2023-02-13T18:34:23.448540Z","shell.execute_reply.started":"2023-02-13T18:34:23.428651Z","shell.execute_reply":"2023-02-13T18:34:23.447709Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# nice plot of class counts as a proportion of the total\nsns.countplot(\n    data=train_df,\n    x=\"correct\"\n)","metadata":{"execution":{"iopub.status.busy":"2023-02-13T18:34:23.452678Z","iopub.execute_input":"2023-02-13T18:34:23.454924Z","iopub.status.idle":"2023-02-13T18:34:23.669041Z","shell.execute_reply.started":"2023-02-13T18:34:23.454881Z","shell.execute_reply":"2023-02-13T18:34:23.668172Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Make \"predictions\"\nAs you can see, around 70% of the quiestions in the training data was answered correctly. We could simply always predict `1` with this probability and `0` otherwise.","metadata":{}},{"cell_type":"code","source":"random.seed(123)\nthreshold = class_prop[1] # ~ 70%\n\nenv = jo_wilder.make_env()\niter_test = env.iter_test()\n\nprint(\"Predictions:\")\nfor (sample_submission, test) in iter_test:\n    r = random.random()\n    if r < threshold:\n        sample_submission[\"correct\"] = 1\n        print(\"1\", end=\" \")\n    else:\n        sample_submission[\"correct\"] = 0\n        print(\"0\", end=\" \")\n    env.predict(sample_submission)","metadata":{"execution":{"iopub.status.busy":"2023-02-13T18:34:23.672817Z","iopub.execute_input":"2023-02-13T18:34:23.675355Z","iopub.status.idle":"2023-02-13T18:34:23.736262Z","shell.execute_reply.started":"2023-02-13T18:34:23.675316Z","shell.execute_reply":"2023-02-13T18:34:23.735362Z"},"trusted":true},"execution_count":null,"outputs":[]}]}