{"cells":[{"metadata":{},"cell_type":"markdown","source":"#Why experts are rethinking how we teach statistics in schools.\n\nhttps://www.kaggle.com/mpwolke/cusersmarildownloadsetudiantscsv/discussion/124699\n\n\nMathematics is behind everything we do in an advanced society, and as we become more dependent on technology, it will underpin more jobs than ever before. And yet, fewer and fewer young people are pursuing mathematics in their later years of high school, let alone at university. \n\nThe shortage of workers who are skilled in mathematics and technology is making its mark on industries such as banks, tech firms, and even government agencies. Even if better teaching methods do not lead directly to a career in mathematics, they will foster more informed citizens in general.\n\n****TEACHERS MATTER\n\nMost only knew about the subject if they happened to do a statistics course as part of their degree, in science.\n\n***CONTEXT IS EVERYTHING\n\nThe students aren't using formal statistical inference, such as p-values. They’re asking questions, collecting data, drawing graphs, and analysing them to answer the questions, stating their confidence in the answer. They are using informal inference, which is based on the evidence they have collected. If they get to university, they will say, ‘Oh, I've done that before. They're just giving me a more theoretical way of doing it.'\nhttps://www.utas.edu.au/news/2018/9/10/802-why-experts-are-rethinking-how-we-teach-statistics-in-schools/\n"},{"metadata":{},"cell_type":"markdown","source":"![](https://www.utas.edu.au/tf-assets/media/images/Research_800x700px11_pCYOy2v.width-670.jpg)https://www.utas.edu.au/news/2018/9/10/802-why-experts-are-rethinking-how-we-teach-statistics-in-schools/"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"_kg_hide-output":true},"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport plotly.express as px\nimport plotly.graph_objects as go\nimport plotly.offline as py\nimport plotly.express as px\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 5GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"#Code by Heroseo  https://www.kaggle.com/piantic/riiid-answer-correctness-prediction-basic-eda"},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"#Code by Heroseo  https://www.kaggle.com/piantic/riiid-answer-correctness-prediction-basic-eda\n\nfrom colorama import Fore, Style\n\ndf = pd.read_csv('/kaggle/input/riiid-test-answer-prediction/train.csv', low_memory=False, nrows=10**3, \n                       dtype={'row_id': 'int64', 'timestamp': 'int64', 'user_id': 'int32', 'content_id': 'int16', 'content_type_id': 'int8',\n                              'task_container_id': 'int16', 'user_answer': 'int8', 'answered_correctly': 'int8', 'prior_question_elapsed_time': 'float32', \n                             'prior_question_had_explanation': 'boolean',\n                             }\n                      )\nprint(Fore.MAGENTA + 'Training data shape: ',Style.RESET_ALL,df.shape)\ndf","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"#It's a huge amount of rows. I changed only the Fore color and the number of rows. And `train_df` is just df, since I won't train anything. "},{"metadata":{},"cell_type":"markdown","source":"#Code from Tanay Mehta https://www.kaggle.com/heyytanay/police-shooting-eda-interactive-map/comments"},{"metadata":{"trusted":true},"cell_type":"code","source":"#Code from Tanay Mehta https://www.kaggle.com/heyytanay/police-shooting-eda-interactive-map/comments\n\n\ndef count(string: str, color=Fore.RED):\n    \"\"\"\n    Saves some work \n    \"\"\"\n    print(color+string+Style.RESET_ALL)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def statistics(dataframe, column):\n    count(f\"The Average value in {column} is: {dataframe[column].mean():.2f}\", Fore.RED)\n    count(f\"The Maximum value in {column} is: {dataframe[column].max()}\", Fore.BLUE)\n    count(f\"The Minimum value in {column} is: {dataframe[column].min()}\", Fore.YELLOW)\n    count(f\"The 25th Quantile of {column} is: {dataframe[column].quantile(0.25)}\", Fore.GREEN)\n    count(f\"The 50th Quantile of {column} is: {dataframe[column].quantile(0.50)}\", Fore.CYAN)\n    count(f\"The 75th Quantile of {column} is: {dataframe[column].quantile(0.75)}\", Fore.MAGENTA)","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"# Print Offset Column Statistics\nstatistics(df, 'timestamp')","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"# Let's plot Timestamp\nplt.style.use(\"classic\")\nsns.distplot(df['timestamp'], color='blue')\nplt.title(f\"Timestamp [\\u03BC : {df['timestamp'].mean():.2f} conditions | \\u03C3 : {df['timestamp'].std():.2f} conditions]\")\nplt.xlabel(\"Timestamp\")\nplt.ylabel(\"Count\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"# Print Offset Column Statistics\nstatistics(df, 'task_container_id')","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"# Let's plot Task Container Id\nplt.style.use(\"classic\")\nsns.distplot(df['task_container_id'], color='red')\nplt.title(f\"Task container id [\\u03BC : {df['task_container_id'].mean():.2f} status | \\u03C3 : {df['task_container_id'].std():.2f} status]\")\nplt.xlabel(\"Task container id\")\nplt.ylabel(\"Count\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"# Print Offset Column Statistics\nstatistics(df, 'content_id')","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"# Let's plot Content Id\nplt.style.use(\"classic\")\nsns.distplot(df['content_id'], color='green')\nplt.title(f\"Content Id [\\u03BC : {df['content_id'].mean():.2f} status | \\u03C3 : {df['content_id'].std():.2f} status]\")\nplt.xlabel(\"Content Id\")\nplt.ylabel(\"Count\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"#Code by Crislânio Macedo, a.k.a. Caesar Lupum "},{"metadata":{"trusted":true},"cell_type":"code","source":"#Code by Crislânio Macedo, a.k.a. Caesar Lupum \n\ndef plot_dist_col(column):\n    pos__df = df[df['answered_correctly'] ==1]\n    neg__df = df[df['answered_correctly'] ==0]\n\n    '''plot dist curves for train and test weather data for the given column name'''\n    fig, ax = plt.subplots(figsize=(8, 8))\n    sns.distplot(pos__df[column].dropna(), color='green', ax=ax).set_title(column, fontsize=16)\n    sns.distplot(neg__df[column].dropna(), color='purple', ax=ax).set_title(column, fontsize=16)\n    plt.xlabel(column, fontsize=15)\n    plt.legend(['answered_correctly', 'content_id'])\n    plt.show()\nplot_dist_col('content_id')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"##Code by Olga Belitskaya https://www.kaggle.com/olgabelitskaya/parts-of-speech"},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"#Code by Olga Belitskaya https://www.kaggle.com/olgabelitskaya/parts-of-speech\nplt.figure(figsize=(10,5))\nsns.countplot(y=\"prior_question_had_explanation\",data=df,\n             facecolor=(0,0,0,0),linewidth=5,\n             edgecolor=sns.color_palette(\"rainbow\"))\nplt.title('Prior Question Had Explanation?',\n         fontsize=15);","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"fig = px.bar(df, \n             x='answered_correctly', y='content_id', color_discrete_sequence=['crimson'],\n             title='Re-thinking Education?', text='timestamp')\n\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"sns.heatmap(df.corr(),cmap = 'summer',cbar = True)","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"df1 = pd.read_csv('../input/riiid-test-answer-prediction/questions.csv', encoding='ISO-8859-2')\ndf1.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"#Code by Carl MacBride Ellis https://www.kaggle.com/carlmcbrideellis/riiid-a-quick-eda-of-train-csv-using-dabl-and-rf"},{"metadata":{"_kg_hide-output":true,"trusted":true},"cell_type":"code","source":"!pip install dabl\nimport dabl","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"dabl.plot(df1, target_col=\"correct_answer\")","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"#Code by Olga Belitskaya https://www.kaggle.com/olgabelitskaya/sequential-data/comments\nfrom IPython.display import display,HTML\nc1,c2,f1,f2,fs1,fs2=\\\n'#3486eb','#3437eb','Akronim','Smokum',30,15\ndef dhtml(string,fontcolor=c1,font=f1,fontsize=fs1):\n    display(HTML(\"\"\"<style>\n    @import 'https://fonts.googleapis.com/css?family=\"\"\"\\\n    +font+\"\"\"&effect=3d-float';</style>\n    <h1 class='font-effect-3d-float' style='font-family:\"\"\"+\\\n    font+\"\"\"; color:\"\"\"+fontcolor+\"\"\"; font-size:\"\"\"+\\\n    str(fontsize)+\"\"\"px;'>%s</h1>\"\"\"%string))\n    \n    \ndhtml('To be a good teacher, you must have compassion, tolerance, knowledge of life, passion for sharing what you know.' )","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"\"To be a good teacher, you must have compassion, tolerance, patience, practical wisdom, knowledge of life, and a passion for sharing what you know.\" \n\n\"This is not guaranteed by the diploma/titles. In Barefoot College, everyone is a lifelong student because you never stop learning (or unlearning).\"\n\n#\"It is a place where the teacher is the student, and the student is the teacher.\" Words from SANJIT \"BUNKER\" ROY.\n\nhttps://www.kaggle.com/mpwolke/cusersmarildownloadsetudiantscsv/discussion/125097\n\nhttps://oglobo.globo.com/sociedade/educacao/educacao-360/para-ser-bom-professor-preciso-ter-compaixao-diz-ativista-indiano-23051077"},{"metadata":{},"cell_type":"markdown","source":"#Dedicated to All Kagglers that shared their works in Kaggle. They are sort of my Programming language teachers.\n\nSome of them are mentioned in this Kaggle Notebook as I usually do with my Pythonish Style. \n\n#Thank you Kagglers. Keep Sharing and Keep Learning!"}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}