{"cells":[{"metadata":{"_uuid":"412cce80e1f8d64bc0df5fbc5120cb8de3d30eef"},"cell_type":"markdown","source":"**Earthquakes - Exploratory data Analysis**\n\nIn this competition we want to predict timing between earthquakes at specific locations. But first let's understand the data better!\n\nI highly recommend reading [the cited article](https://www.nature.com/articles/ncomms11104) to understand what we are talking here about at all.\n\nNow, let's load libraries:"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport os\nimport warnings\nwarnings.simplefilter(action='ignore', category=FutureWarning)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3eab31e06a1c4a74c82192877b3085f710c738f4"},"cell_type":"markdown","source":"Loading of data itself. As there is about 9GB of data I will load only a sample first (1 milion of rows)"},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"train = pd.read_csv(\"../input/train.csv\", dtype={'acoustic_data': np.int16, 'time_to_failure': np.float64}, nrows=2000000)\nprint(\"train shape\", train.shape)\npd.options.display.precision = 20\ntrain.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"06d7b01c450f320217f275135a49bdd6d5352566"},"cell_type":"markdown","source":"**Missing values**"},{"metadata":{"trusted":true,"_uuid":"81d1e4579805dbe32dc2a90a58da3c78280a9d3b"},"cell_type":"code","source":"train.isna().sum()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a9c7595f4f1c5644d89d5fd069c87cf97fad8457"},"cell_type":"markdown","source":"There are no missing values - very good!\n\n\n**Basic descriptive statistics**"},{"metadata":{"trusted":true,"_uuid":"b50c15a3d583e609a1801c0128485fc37c8f45de"},"cell_type":"code","source":"train[\"acoustic_data\"].describe()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"4236bc60a73130ac599c3749b4e74c254ce71fec"},"cell_type":"markdown","source":"First the most basic plot \"time_to_failure\" vs. \"acoustic_data\"."},{"metadata":{"trusted":true,"_uuid":"2dbe0380d75fd15e7c7bc182a66330fefac2c985"},"cell_type":"code","source":"plt.figure(figsize=(12,6))\nplt.title(\"time_to_failure histogram\")\nax = plt.plot(train[\"time_to_failure\"], train[\"acoustic_data\"])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"2db5091ecaa4b15e6eaafaa9aadb047f4d947080"},"cell_type":"markdown","source":"Distribution of \"acoustic_data\"."},{"metadata":{"_uuid":"8701adbcd746f4fa7ea01c9258e4b0bc860c93d0"},"cell_type":"markdown","source":""},{"metadata":{"trusted":true,"_uuid":"4937d02b40b6cef9a2d25af6e178920f3da9d501"},"cell_type":"code","source":"plt.figure(figsize=(12,6))\nplt.title(\"Acoustic data histogram\")\nax = sns.distplot(train[\"acoustic_data\"], label='Acustic data')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5f6e7bee0e217aebc8f09f150fbde7e346f5074b"},"cell_type":"markdown","source":"Most of data is gathered in a very narrow range. Therefore let's create a range = mean +/- 2 standard deviations (4 sigma) for better visualisation."},{"metadata":{"trusted":true,"_uuid":"38c0ae10f3347c50530df23821e353b73e7a932c"},"cell_type":"code","source":"upper = train[\"acoustic_data\"].mean()+2*train[\"acoustic_data\"].std()\nlower = train[\"acoustic_data\"].mean()-2*train[\"acoustic_data\"].std()\n\ntrain_subset = train[(train[\"acoustic_data\"]>lower) & (train[\"acoustic_data\"]<upper)]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b4e94c89639022cda0cc7e93d0c24b5504c8577e"},"cell_type":"code","source":"plt.figure(figsize=(12,6))\nplt.title(\"Acoustic data histogram\")\nax = sns.distplot(train_subset[\"acoustic_data\"], label='Acustic data', kde=False)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"3da95829e2d93ad028ad74ea02ca3ec7a9706c55"},"cell_type":"code","source":"plt.figure(figsize=(10,5))\nplt.title(\"time_to_failure histogram\")\nax = sns.distplot(train[\"time_to_failure\"], label='time_to_failure')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a61cd3a12784906dcdc50d0ecc703e0776bf4e55"},"cell_type":"markdown","source":"UNDER CONSTRUCTION"}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}