{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"In this notebook, I leverage the power of a **Spark cluster** to explore the large (~88GB) page_views.csv dataset and analyze its relationshiop with events.csv.  \nAfter some hours of processing, we've got answer to questions like: \n\n* **How to join page_views.csv and events.csv?**\n* **Is events.csv a subset of page_views.csv?**\n* **Are there additional page views for users in events.csv?**  ","metadata":{"_cell_guid":"5c42ed67-f3c3-f3f2-17f0-631d55c9adbd"}},{"cell_type":"markdown","source":"**IMPORTANT:** This Jupyter notebook was implemented in my own Spark cluster, and it was not possible to share the notebook as a Kernel without actually running it as a Kernel (which obviously would not be possible).  \n**Thus, you can see the [full notebook on my GitHub](https://github.com/gabrielspmoreira/static_resources/blob/gh-pages/Kaggle-Outbrain-PageViews_EventsAnalytics.ipynb).**","metadata":{"_cell_guid":"70ede1a4-aaa1-e163-c0f7-df025f594f58"}},{"cell_type":"markdown","source":"**If this Kaggle kernel helps you, don't forget to thumbs it up! ;)**","metadata":{"_cell_guid":"abd43ba8-5636-f1a5-31c6-6ec39930d889"}},{"cell_type":"markdown","source":"By [Gabriel S. P. Moreira](https://about.me/gspmoreira)","metadata":{"_cell_guid":"21253ccd-170d-155a-8c3f-2517de1bd710"}},{"cell_type":"markdown","source":"","metadata":{"_cell_guid":"e7699cf4-9c4c-5ae6-dcf3-9b8c04f14906"}}]}