{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Baseline: Most Common Class\n\nIt is always important to establish a simple baseline before building more sophisticated models. This way you have a reference to assess the true value added of your complex models. Often, it's hard to outperform these simple baseline models or the added complexity of other models is hard to justify.\n\nSo lets establish a simple baseline model that always predicts the most common class found in the training data. This is of course not a very useful model but gives us something to compare to.","metadata":{}},{"cell_type":"markdown","source":"# Imports","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport seaborn as sns\nimport jo_wilder","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-02-13T18:10:50.061316Z","iopub.execute_input":"2023-02-13T18:10:50.062192Z","iopub.status.idle":"2023-02-13T18:10:50.683577Z","shell.execute_reply.started":"2023-02-13T18:10:50.062082Z","shell.execute_reply":"2023-02-13T18:10:50.682265Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load data\nHere we load the training label data so we can see what the most common output class.","metadata":{}},{"cell_type":"code","source":"# load training lables\ntrain_df = pd.read_csv(\"/kaggle/input/predict-student-performance-from-game-play/train_labels.csv\")\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-02-13T18:10:50.685442Z","iopub.execute_input":"2023-02-13T18:10:50.685764Z","iopub.status.idle":"2023-02-13T18:10:50.981969Z","shell.execute_reply.started":"2023-02-13T18:10:50.685735Z","shell.execute_reply":"2023-02-13T18:10:50.981176Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Most common output class\nHere we have a look to see what the most common class (whether the user answered the quistion correctly). The assumption is that the training data is representative of the test data as well. If this holds true, there would be a similar proportion of this class in the test data.","metadata":{}},{"cell_type":"code","source":"# quick check to see if we have any null values\ntrain_df.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2023-02-13T18:10:50.983089Z","iopub.execute_input":"2023-02-13T18:10:50.983748Z","iopub.status.idle":"2023-02-13T18:10:51.001947Z","shell.execute_reply.started":"2023-02-13T18:10:50.983717Z","shell.execute_reply":"2023-02-13T18:10:51.000825Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# quick inspection of class counts\ntrain_df[\"correct\"].value_counts()","metadata":{"execution":{"iopub.status.busy":"2023-02-13T18:10:51.004310Z","iopub.execute_input":"2023-02-13T18:10:51.004804Z","iopub.status.idle":"2023-02-13T18:10:51.022489Z","shell.execute_reply.started":"2023-02-13T18:10:51.004767Z","shell.execute_reply":"2023-02-13T18:10:51.021401Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# as proportion\ntrain_df[\"correct\"].value_counts() / train_df.shape[0]","metadata":{"execution":{"iopub.status.busy":"2023-02-13T18:10:51.024016Z","iopub.execute_input":"2023-02-13T18:10:51.024416Z","iopub.status.idle":"2023-02-13T18:10:51.034964Z","shell.execute_reply.started":"2023-02-13T18:10:51.024386Z","shell.execute_reply":"2023-02-13T18:10:51.033776Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# nice plot of class counts as a proportion of the total\nsns.countplot(\n    data=train_df,\n    x=\"correct\"\n)","metadata":{"execution":{"iopub.status.busy":"2023-02-13T18:10:51.036447Z","iopub.execute_input":"2023-02-13T18:10:51.036822Z","iopub.status.idle":"2023-02-13T18:10:51.223751Z","shell.execute_reply.started":"2023-02-13T18:10:51.036789Z","shell.execute_reply":"2023-02-13T18:10:51.222836Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Make \"predictions\"\nAs you can see, around 70% of the quiestions in the training data was answered correctly. We could simply always predict `1` and we should be correct quite often.","metadata":{}},{"cell_type":"code","source":"env = jo_wilder.make_env()\niter_test = env.iter_test()\n\nfor (sample_submission, test) in iter_test:\n    sample_submission[\"correct\"] = 1\n    env.predict(sample_submission)","metadata":{"execution":{"iopub.status.busy":"2023-02-13T18:10:51.227807Z","iopub.execute_input":"2023-02-13T18:10:51.228636Z","iopub.status.idle":"2023-02-13T18:10:51.294431Z","shell.execute_reply.started":"2023-02-13T18:10:51.228597Z","shell.execute_reply":"2023-02-13T18:10:51.293460Z"},"trusted":true},"execution_count":null,"outputs":[]}]}