{"cells":[{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true,"_kg_hide-input":false},"cell_type":"code","source":"import pandas as pd\nimport numpy as np\n\ntrain = pd.read_csv('/kaggle/input/riiid-test-answer-prediction/train.csv',\n                   dtype={'row_id': 'int64',\n                          'timestamp': 'int64',\n                          'user_id': 'int32',\n                          'content_id': 'int16',\n                          'content_type_id': 'int8',\n                          'task_container_id': 'int16',\n                          'user_answer': 'int8',\n                          'answered_correctly':'int8',\n                          'prior_question_elapsed_time': 'float32',\n                          'prior_question_had_explanation': 'boolean'}\n                   )","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"prior_question_elapsed_time = train.prior_question_elapsed_time.dropna()\nprior_question_elapsed_time.mean()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Is this result correct?\n### **No, the actual average is twice as high as this result!　工ｴｴｪｪ Σ”(⚙♊⚙ﾉ)ﾉ ｪｪｴｴ工‼"},{"metadata":{"trusted":true},"cell_type":"code","source":"total = np.zeros(1,dtype=np.float128)\nfor v in prior_question_elapsed_time:\n    total[0] += v\nprint(total[0]/len(prior_question_elapsed_time))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"**prior_question_elapsed_time.mean() is wrong!!!**"},{"metadata":{"trusted":true},"cell_type":"code","source":"prior_question_elapsed_time","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"prior_question_elapsed_time.min(), prior_question_elapsed_time.max()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Even min and max are ok with float32, pandas series mean() returned wrong. \nIf we use float64, it looks OK."},{"metadata":{"trusted":true},"cell_type":"code","source":"prior_question_elapsed_time.astype('float64').mean()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"numpy.mean() looks OK too even with float32."},{"metadata":{"trusted":true},"cell_type":"code","source":"prior_question_elapsed_time.values","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"prior_question_elapsed_time.values.mean()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"I don't know the detail, but this problem happens baddly especially when the data is large.\nMay be it happens when the `sum` overflows flot32."},{"metadata":{"trusted":true},"cell_type":"code","source":"for i in range(9):\n    target_data = prior_question_elapsed_time[:10**i]\n    diff = target_data.astype('float64').mean() - target_data.mean()\n    print(f'{i} {abs(diff):.4f}')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"So we need to be carefull with pandas mean(). The results of this experiment are:\n<pre>\n.mean() with float32 is NG\n.mean() with float64 is OK\n.values.mean() OK (you need to dropna beforehand for np.mean)\n</pre>"}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}