{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"In the training data not all users are playing the same game events because there are four different scripts that are randomly chosen when the game starts.\n\nAs pointet out by @mattoglesby [here](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/387864#2147760) description of game types can be found on github [jo_wilder](https://github.com/fielddaylab/jo_wilder) on Script Versions paragraph :  \n\n\n> Script Versions\n> \n> Starting in v7, multiple scripts were added to the game for AB tests on snark and humor. The game will randomly choose between 4 different data files: data_dry.js, data_nohumor.js, data_nosnark.js, and the original data.js. They are each built from their respective data folder in assets. The type of script used is only logged once, in startgame.\n> \n> \n> - 0 dry: no humor or snark\n> - 1 nohumor: no humor (includes snark)\n> - 2 nosnark: no snark (includes humor). No snark can also be thought of as \"obedient\"\n> - 3 normal: base script (includes snark and humor)\n\nIf you want to play a specific game type you must use these custom links:\n* normal:  https://jowilder-master.netlify.app/?script_type=original\n* dry: https://jowilder-master.netlify.app/?script_type=dry\n* nohumor: https://jowilder-master.netlify.app/?script_type=nohumor\n* nosnark:  https://jowilder-master.netlify.app/?script_type=nosnark\n\nif you want to play a randomly chosen game use this link https://jowilder-master.netlify.app/ \n\nThere are differenes between texts in each of these versions.\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F214989%2Feba6581ebd028faff88f08a2b719758e%2Fsample.png?generation=1679297783311464&alt=media)\n\nIn this notebook I will try to generate a dataset that contains text differences  between types of game by analyzing the source files\n\nYou will find dataset [here](https://www.kaggle.com/datasets/steubk/meetings-are-boring)","metadata":{"papermill":{"duration":0.005123,"end_time":"2023-03-20T07:00:47.034967","exception":false,"start_time":"2023-03-20T07:00:47.029844","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"First off all we donwnolad sources of the 5 game types from github","metadata":{"papermill":{"duration":0.003791,"end_time":"2023-03-20T07:00:47.042979","exception":false,"start_time":"2023-03-20T07:00:47.039188","status":"completed"},"tags":[]}},{"cell_type":"code","source":"!wget https://github.com/fielddaylab/jo_wilder/raw/master/src/scenes/data.js\n!wget https://github.com/fielddaylab/jo_wilder/raw/master/src/scenes/data_dry.js\n!wget https://github.com/fielddaylab/jo_wilder/raw/master/src/scenes/data_nohumor.js\n!wget https://github.com/fielddaylab/jo_wilder/raw/master/src/scenes/data_nosnark.js","metadata":{"execution":{"iopub.execute_input":"2023-03-20T07:00:47.056247Z","iopub.status.busy":"2023-03-20T07:00:47.055287Z","iopub.status.idle":"2023-03-20T07:00:53.864905Z","shell.execute_reply":"2023-03-20T07:00:53.863107Z"},"papermill":{"duration":6.821981,"end_time":"2023-03-20T07:00:53.869722","exception":false,"start_time":"2023-03-20T07:00:47.047741","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":".. then we use `diff` command to grab diffences between source files","metadata":{"papermill":{"duration":0.006623,"end_time":"2023-03-20T07:00:53.881877","exception":false,"start_time":"2023-03-20T07:00:53.875254","status":"completed"},"tags":[]}},{"cell_type":"code","source":"!diff data.js data_dry.js | grep raw_atext | head","metadata":{"execution":{"iopub.execute_input":"2023-03-20T07:00:53.903116Z","iopub.status.busy":"2023-03-20T07:00:53.902569Z","iopub.status.idle":"2023-03-20T07:00:54.963506Z","shell.execute_reply":"2023-03-20T07:00:54.961816Z"},"papermill":{"duration":1.077241,"end_time":"2023-03-20T07:00:54.966728","exception":false,"start_time":"2023-03-20T07:00:53.889487","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we save differences between normal game and dry, nohumor, nosnark games in separate files ... ","metadata":{"papermill":{"duration":0.004756,"end_time":"2023-03-20T07:00:54.976859","exception":false,"start_time":"2023-03-20T07:00:54.972103","status":"completed"},"tags":[]}},{"cell_type":"code","source":"!diff data.js data_dry.js | grep raw_atext > data_dry.diff\n!diff data.js data_nohumor.js | grep raw_atext > data_nohumor.diff\n!diff data.js data_nosnark.js | grep raw_atext > data_nosnark.diff","metadata":{"execution":{"iopub.execute_input":"2023-03-20T07:00:54.989614Z","iopub.status.busy":"2023-03-20T07:00:54.989105Z","iopub.status.idle":"2023-03-20T07:00:58.125182Z","shell.execute_reply":"2023-03-20T07:00:58.123641Z"},"papermill":{"duration":3.146424,"end_time":"2023-03-20T07:00:58.128366","exception":false,"start_time":"2023-03-20T07:00:54.981942","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"... and we build parser to read these files","metadata":{"papermill":{"duration":0.004766,"end_time":"2023-03-20T07:00:58.138306","exception":false,"start_time":"2023-03-20T07:00:58.133540","status":"completed"},"tags":[]}},{"cell_type":"code","source":"import pandas as pd\n\ndef game_type_text_translation ( filename ):\n\n    with open(filename) as file:\n        d = {}\n        for line in file:\n            line = line.rstrip()\n            text = line.replace(\"tmp_speak_command.raw_atext =\",\"\")\n\n            f = text[0]\n            text = text.replace('<  \"',\"\").replace('>  \"',\"\").replace('\";',\"\")\n            if f == \"<\":\n                key = text\n            else:\n                value = text\n                d[key] = value \n    \n    return d\n\nd_dry = game_type_text_translation (\"data_dry.diff\")\nd_nohumor = game_type_text_translation (\"data_nohumor.diff\")\nd_nosnark = game_type_text_translation (\"data_nosnark.diff\")\n\n","metadata":{"execution":{"iopub.execute_input":"2023-03-20T07:00:58.152295Z","iopub.status.busy":"2023-03-20T07:00:58.151375Z","iopub.status.idle":"2023-03-20T07:00:58.161627Z","shell.execute_reply":"2023-03-20T07:00:58.160414Z"},"papermill":{"duration":0.021315,"end_time":"2023-03-20T07:00:58.164931","exception":false,"start_time":"2023-03-20T07:00:58.143616","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I's time to build our dataset:","metadata":{"papermill":{"duration":0.004813,"end_time":"2023-03-20T07:00:58.175062","exception":false,"start_time":"2023-03-20T07:00:58.170249","status":"completed"},"tags":[]}},{"cell_type":"code","source":"data = []\ndry = []\nnohumor = []\nnosnark = []\nfor key in set(d_dry.keys()).union(d_nohumor.keys()).union(d_nosnark.keys()):\n    data.append(key)\n    dry.append(d_dry.get(key,key))\n    nohumor.append(d_nohumor.get(key,key))\n    nosnark.append(d_nosnark.get(key,key))\n    \n    \ndf = pd.DataFrame({\n    \"normal\": data,\n    \"dry\": dry,\n    \"nohumor\": nohumor,\n    \"nosnark\": nosnark,\n}) \n\n\n\ndf.to_csv(\"game_type_text_translation.csv\", index=False)\n","metadata":{"execution":{"iopub.execute_input":"2023-03-20T07:00:58.188223Z","iopub.status.busy":"2023-03-20T07:00:58.187299Z","iopub.status.idle":"2023-03-20T07:00:58.209541Z","shell.execute_reply":"2023-03-20T07:00:58.208033Z"},"papermill":{"duration":0.033238,"end_time":"2023-03-20T07:00:58.213570","exception":false,"start_time":"2023-03-20T07:00:58.180332","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.head(20)","metadata":{"execution":{"iopub.execute_input":"2023-03-20T07:00:58.226892Z","iopub.status.busy":"2023-03-20T07:00:58.225973Z","iopub.status.idle":"2023-03-20T07:00:58.262921Z","shell.execute_reply":"2023-03-20T07:00:58.261505Z"},"papermill":{"duration":0.047272,"end_time":"2023-03-20T07:00:58.266149","exception":false,"start_time":"2023-03-20T07:00:58.218877","status":"completed"},"tags":[]},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"done!","metadata":{"papermill":{"duration":0.005432,"end_time":"2023-03-20T07:00:58.277485","exception":false,"start_time":"2023-03-20T07:00:58.272053","status":"completed"},"tags":[]}}]}