{
  "id": 205830,
  "title": "[Float64 mean()] = ~2 x [Float32 mean()] for 'prior_question_elapsed_time'",
  "url": "/competitions/riiid-test-answer-prediction/discussion/205830",
  "author_name": "",
  "post_date": "2020-12-22T04:16:25.295593Z",
  "votes": 3,
  "comment_count": 2,
  "views": 0,
  "content": "<pre><code>data_type_dict = {\n    'timestamp': 'int64',\n    'user_id': 'int32',\n    'content_id': 'int16',\n    'content_type_id': 'int8',\n    'answered_correctly': 'int8',\n    'prior_question_elapsed_time': 'float32',\n}\ntarget = 'answered_correctly'\ntrain_df = dt.fread('../input/riiid-test-answer-prediction/train.csv',\n                   columns=set(data_type_dict.keys())).to_pandas()\nsum1 = train_df['prior_question_elapsed_time'].sum()\navg1 = train_df['prior_question_elapsed_time'].mean()\nn1 = train_df['prior_question_elapsed_time'].count()\nprint(sum1,avg1,n1,sum1/n1)\n\ntrain_df = train_df.astype(data_type_dict)\n\nsum2 = train_df['prior_question_elapsed_time'].sum()\navg2 = train_df['prior_question_elapsed_time'].mean()\nn2 = train_df['prior_question_elapsed_time'].count()\nprint(sum2,avg2,n2,sum2/n2)\n</code></pre>\n<p>Results in</p>\n<p>2513875675933.0 25423.810042960275 98878794 25423.810042960275<br>\n2513879000000.0 13005.0810546875 98878794 25423.842739505904</p>\n<p>Sum1 = Sum2, and  Count1 = Count2, but Mean1 != Mean2<br>\nAny insight would be appreciated. Thanks.</p>",
  "messages": [
    {
      "id": "1121968",
      "postDate": "12/22/2020 04:16:25",
      "content": "<pre><code>data_type_dict = {\n    'timestamp': 'int64',\n    'user_id': 'int32',\n    'content_id': 'int16',\n    'content_type_id': 'int8',\n    'answered_correctly': 'int8',\n    'prior_question_elapsed_time': 'float32',\n}\ntarget = 'answered_correctly'\ntrain_df = dt.fread('../input/riiid-test-answer-prediction/train.csv',\n                   columns=set(data_type_dict.keys())).to_pandas()\nsum1 = train_df['prior_question_elapsed_time'].sum()\navg1 = train_df['prior_question_elapsed_time'].mean()\nn1 = train_df['prior_question_elapsed_time'].count()\nprint(sum1,avg1,n1,sum1/n1)\n\ntrain_df = train_df.astype(data_type_dict)\n\nsum2 = train_df['prior_question_elapsed_time'].sum()\navg2 = train_df['prior_question_elapsed_time'].mean()\nn2 = train_df['prior_question_elapsed_time'].count()\nprint(sum2,avg2,n2,sum2/n2)\n</code></pre>\n<p>Results in</p>\n<p>2513875675933.0 25423.810042960275 98878794 25423.810042960275<br>\n2513879000000.0 13005.0810546875 98878794 25423.842739505904</p>\n<p>Sum1 = Sum2, and  Count1 = Count2, but Mean1 != Mean2<br>\nAny insight would be appreciated. Thanks.</p>",
      "rawMarkdown": "```\ndata_type_dict = {\n    'timestamp': 'int64',\n    'user_id': 'int32',\n    'content_id': 'int16',\n    'content_type_id': 'int8',\n    'answered_correctly': 'int8',\n    'prior_question_elapsed_time': 'float32',\n}\ntarget = 'answered_correctly'\ntrain_df = dt.fread('../input/riiid-test-answer-prediction/train.csv',\n                   columns=set(data_type_dict.keys())).to_pandas()\nsum1 = train_df['prior_question_elapsed_time'].sum()\navg1 = train_df['prior_question_elapsed_time'].mean()\nn1 = train_df['prior_question_elapsed_time'].count()\nprint(sum1,avg1,n1,sum1/n1)\n\ntrain_df = train_df.astype(data_type_dict)\n\nsum2 = train_df['prior_question_elapsed_time'].sum()\navg2 = train_df['prior_question_elapsed_time'].mean()\nn2 = train_df['prior_question_elapsed_time'].count()\nprint(sum2,avg2,n2,sum2/n2)\n```\n\nResults in\n\n2513875675933.0 25423.810042960275 98878794 25423.810042960275\n2513879000000.0 13005.0810546875 98878794 25423.842739505904\n\nSum1 = Sum2, and  Count1 = Count2, but Mean1 != Mean2\nAny insight would be appreciated. Thanks.",
      "votes": null
    },
    {
      "id": "1121969",
      "postDate": "12/22/2020 04:22:48",
      "content": "<p>Sum and count are integers, using floating point representation can be pretty accurate because the biggest number that can be represented FP32 is about 2^31.</p>\n<p>For decimals, FP32 has only 9 significant digits, therefore the round error is huge whenever summing up so many rows to get an average.</p>",
      "rawMarkdown": "Sum and count are integers, using floating point representation can be pretty accurate because the biggest number that can be represented FP32 is about 2^31.\n\nFor decimals, FP32 has only 9 significant digits, therefore the round error is huge whenever summing up so many rows to get an average.",
      "votes": null
    },
    {
      "id": "1121978",
      "postDate": "12/22/2020 04:47:26",
      "content": "<p>understood. thanks.</p>",
      "rawMarkdown": "understood. thanks.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1121969,
      "author_name": "scaomath",
      "author_url": "",
      "post_date": "12/22/2020 04:22:48",
      "content": "<p>Sum and count are integers, using floating point representation can be pretty accurate because the biggest number that can be represented FP32 is about 2^31.</p>\n<p>For decimals, FP32 has only 9 significant digits, therefore the round error is huge whenever summing up so many rows to get an average.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1121978,
      "author_name": "farstars",
      "author_url": "",
      "post_date": "12/22/2020 04:47:26",
      "content": "<p>understood. thanks.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1121968": "```\ndata_type_dict = {\n    'timestamp': 'int64',\n    'user_id': 'int32',\n    'content_id': 'int16',\n    'content_type_id': 'int8',\n    'answered_correctly': 'int8',\n    'prior_question_elapsed_time': 'float32',\n}\ntarget = 'answered_correctly'\ntrain_df = dt.fread('../input/riiid-test-answer-prediction/train.csv',\n                   columns=set(data_type_dict.keys())).to_pandas()\nsum1 = train_df['prior_question_elapsed_time'].sum()\navg1 = train_df['prior_question_elapsed_time'].mean()\nn1 = train_df['prior_question_elapsed_time'].count()\nprint(sum1,avg1,n1,sum1/n1)\n\ntrain_df = train_df.astype(data_type_dict)\n\nsum2 = train_df['prior_question_elapsed_time'].sum()\navg2 = train_df['prior_question_elapsed_time'].mean()\nn2 = train_df['prior_question_elapsed_time'].count()\nprint(sum2,avg2,n2,sum2/n2)\n```\n\nResults in\n\n2513875675933.0 25423.810042960275 98878794 25423.810042960275\n2513879000000.0 13005.0810546875 98878794 25423.842739505904\n\nSum1 = Sum2, and  Count1 = Count2, but Mean1 != Mean2\nAny insight would be appreciated. Thanks.",
    "1121969": "Sum and count are integers, using floating point representation can be pretty accurate because the biggest number that can be represented FP32 is about 2^31.\n\nFor decimals, FP32 has only 9 significant digits, therefore the round error is huge whenever summing up so many rows to get an average.",
    "1121978": "understood. thanks."
  },
  "source": "meta"
}