{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# TPSOCT22 Baseline Log-loss score\n\n\n**Log-loss** is a common metric used in classification problem, which measures how close the prediction probability is to the corresponding actual/true value (0 or 1 in case of binary classification). \nThe more the predicted probability diverges from the actual value, the higher is the log-loss value.\n\n\n**Baseline log-loss** score for a dataset is determined from the naïve classification model, which simply pegs all the observations with a constant probability equal to % of data with class 1 observations. For a balanced dataset with a 51:49 ratio of class 0 to class 1, a naïve model with constant probability of 0.49 will yield log-loss score of 0.693, which is regarded as the lowest baseline score for that dataset.\n\nIn this TPS OCT dataset, the class is slightly imbalanced, discussed [here](https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/357924), with the ratio of 16:1, which means that out of 17 row of dataset, there is 16 no score vs 1 score. So, simply predict all zeros will give us 16 correct and 1 false prediction(but logloss is used to penalize wrong prediction with huge loss)\n\n\nIn this notebook, I am exploring the baseline log-loss score through optimum prediction probability, which can be the base comparison on how well our model performs as compared to random guess.\n\n**Conclusion**\n\nBaseline log-loss score from test dataset is **0.22538**, which is close to train dataset log-loss, which proves our assumption that test dataset has the similar imbalanced ratio. If you're not confident about the prediction, 0.06 would be the safer option. \n\n**Reference:**\n\nhttps://towardsdatascience.com/intuition-behind-log-loss-score-4e0c9979680a\n\nhttps://towardsdatascience.com/estimate-model-performance-with-log-loss-like-a-pro-9f47d13c8865\n\nhttps://stats.stackexchange.com/questions/276067/whats-considered-a-good-log-loss","metadata":{}},{"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nfrom sklearn.metrics import log_loss\nimport matplotlib.pyplot as plt\n\nimport os\n# for dirname, _, filenames in os.walk('/kaggle/input'):\n#     for filename in filenames:\n#         print(os.path.join(dirname, filename))","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-10-08T10:44:49.949461Z","iopub.execute_input":"2022-10-08T10:44:49.949997Z","iopub.status.idle":"2022-10-08T10:44:49.956429Z","shell.execute_reply.started":"2022-10-08T10:44:49.949957Z","shell.execute_reply":"2022-10-08T10:44:49.955212Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train0_df = pd.read_parquet('/kaggle/input/tps-rocket-league-data-float16-parquet-format/train_0.parquet.gzip')","metadata":{"execution":{"iopub.status.busy":"2022-10-08T10:44:49.961576Z","iopub.execute_input":"2022-10-08T10:44:49.961983Z","iopub.status.idle":"2022-10-08T10:44:51.696220Z","shell.execute_reply.started":"2022-10-08T10:44:49.961946Z","shell.execute_reply":"2022-10-08T10:44:51.694928Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"TARGET = ['team_A_scoring_within_10sec','team_B_scoring_within_10sec']\nx = train0_df.drop(columns=TARGET)\ny = train0_df[TARGET]","metadata":{"execution":{"iopub.status.busy":"2022-10-08T10:44:51.698039Z","iopub.execute_input":"2022-10-08T10:44:51.698401Z","iopub.status.idle":"2022-10-08T10:44:52.106091Z","shell.execute_reply.started":"2022-10-08T10:44:51.698365Z","shell.execute_reply":"2022-10-08T10:44:52.104965Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Imbalance labels with a ratio of ~17:1","metadata":{}},{"cell_type":"code","source":"y.value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-10-08T10:44:52.107948Z","iopub.execute_input":"2022-10-08T10:44:52.108326Z","iopub.status.idle":"2022-10-08T10:44:52.298550Z","shell.execute_reply.started":"2022-10-08T10:44:52.108290Z","shell.execute_reply":"2022-10-08T10:44:52.297437Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Baseline score for Team A - 0.22232","metadata":{}},{"cell_type":"code","source":"feature = 'team_A_scoring_within_10sec'\nloglossA = []\nfor i in range(101):\n    j = i/100\n    pred = np.full(len(y),j)\n    loss = log_loss(y[feature],pred)\n    loglossA.append(loss)\npredA = np.argmin(loglossA)/100\nminLossValue = min(loglossA)\nplt.plot(loglossA)\nplt.plot(predA,minLossValue,marker='o')\nplt.annotate(f'p = {predA}, logloss = {round(minLossValue,5)}',(predA,minLossValue))\nplt.title('Minimum Team A log loss with prediction probability of 0.06')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-10-08T10:44:52.300911Z","iopub.execute_input":"2022-10-08T10:44:52.301239Z","iopub.status.idle":"2022-10-08T10:45:38.312677Z","shell.execute_reply.started":"2022-10-08T10:44:52.301209Z","shell.execute_reply":"2022-10-08T10:45:38.311401Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Baseline score for Team B - 0.21528","metadata":{}},{"cell_type":"code","source":"feature = 'team_B_scoring_within_10sec'\nloglossB = []\nfor i in range(101):\n    j = i/100\n    pred = np.full(len(y),j)\n    loss = log_loss(y[feature],pred)\n    loglossB.append(loss)\npredB = np.argmin(loglossB)/100\nminLossValue = min(loglossB)\nplt.plot(loglossB)\nplt.plot(predB,minLossValue,marker='o')\nplt.annotate(f'p = {predB}, logloss = {round(minLossValue,5)}',(predB,minLossValue))\nplt.title('Minimum Team B log loss with prediction probability of 0.06')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-10-08T10:45:38.313883Z","iopub.execute_input":"2022-10-08T10:45:38.314233Z","iopub.status.idle":"2022-10-08T10:46:24.167779Z","shell.execute_reply.started":"2022-10-08T10:45:38.314200Z","shell.execute_reply":"2022-10-08T10:46:24.166553Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Baseline score submission\n\nIf the test dataset has imbalance distribution which is similar to train dataset, then we will get similar log-loss, as reported from LB.","metadata":{}},{"cell_type":"code","source":"test_df = pd.read_parquet('/kaggle/input/tps-rocket-league-data-float16-parquet-format/test.parquet.gzip')\nsubmitData = pd.read_csv('/kaggle/input/tabular-playground-series-oct-2022/sample_submission.csv')","metadata":{"execution":{"iopub.status.busy":"2022-10-08T10:46:24.169121Z","iopub.execute_input":"2022-10-08T10:46:24.169594Z","iopub.status.idle":"2022-10-08T10:46:24.826675Z","shell.execute_reply.started":"2022-10-08T10:46:24.169560Z","shell.execute_reply":"2022-10-08T10:46:24.825436Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predTest = submitData[TARGET].copy()\nprediction = [0.06, 0.06]\nfor i, feature in enumerate(TARGET): \n    predTest.loc[:,feature] = np.full(len(predTest),prediction[i])\npredTest.head(5)","metadata":{"execution":{"iopub.status.busy":"2022-10-08T10:46:33.075829Z","iopub.execute_input":"2022-10-08T10:46:33.076344Z","iopub.status.idle":"2022-10-08T10:46:33.105958Z","shell.execute_reply.started":"2022-10-08T10:46:33.076304Z","shell.execute_reply":"2022-10-08T10:46:33.104673Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submitData = pd.read_csv('/kaggle/input/tabular-playground-series-oct-2022/sample_submission.csv')\noutput = pd.DataFrame({'id': submitData.id, \n                       'team_A_scoring_within_10sec': predTest['team_A_scoring_within_10sec'],\n                       'team_B_scoring_within_10sec': predTest['team_B_scoring_within_10sec']})\noutput.to_csv('submission.csv', index=False)\nprint(\"Your submission was successfully saved!\")","metadata":{"execution":{"iopub.status.busy":"2022-10-08T10:46:36.587110Z","iopub.execute_input":"2022-10-08T10:46:36.587748Z","iopub.status.idle":"2022-10-08T10:46:38.080300Z","shell.execute_reply.started":"2022-10-08T10:46:36.587708Z","shell.execute_reply":"2022-10-08T10:46:38.078888Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"output.head(5)","metadata":{"execution":{"iopub.status.busy":"2022-10-08T10:46:39.386734Z","iopub.execute_input":"2022-10-08T10:46:39.387908Z","iopub.status.idle":"2022-10-08T10:46:39.400250Z","shell.execute_reply.started":"2022-10-08T10:46:39.387856Z","shell.execute_reply":"2022-10-08T10:46:39.399417Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}