{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Introduction\n\nI am a beginner at Kaggle. My current interested field is data visualization and machine learning.\nBecause I am considering deploying my work on the web, I decided to learn a Plotly library. \n\n\nPlotly is an interactive data visualization library. \nI think it is a good publishing method of visualization results on the web.\n\n\nIn addition, I think Kaggle's survey data is beneficial and insightful information for Kaggle beginners. After I finish this notebook, I  can figure out current ML and DS trends.\nThat's why I choose Plotly chart and this Kaggle survey dataset.\n\nI hope this notebook will help you.","metadata":{}},{"cell_type":"code","source":"%matplotlib inline\n\nimport pandas as pd\nimport numpy as np\nfrom scipy import stats\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport plotly.graph_objects as go\nimport plotly.express as px\n\nimport warnings\n\nplt.rcParams.update({'font.size' : 5})\nwarnings.filterwarnings('ignore')","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-07-12T13:56:45.676442Z","iopub.execute_input":"2022-07-12T13:56:45.676810Z","iopub.status.idle":"2022-07-12T13:56:47.749980Z","shell.execute_reply.started":"2022-07-12T13:56:45.676782Z","shell.execute_reply":"2022-07-12T13:56:47.748685Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df = pd.read_csv('../input/kaggle-survey-2021/kaggle_survey_2021_responses.csv')\ndf.shape","metadata":{"execution":{"iopub.status.busy":"2022-07-12T13:59:07.496678Z","iopub.execute_input":"2022-07-12T13:59:07.497060Z","iopub.status.idle":"2022-07-12T13:59:08.633422Z","shell.execute_reply.started":"2022-07-12T13:59:07.497029Z","shell.execute_reply":"2022-07-12T13:59:08.632098Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Understanding Data\nAs we can see the below colume names are the number of questions such as \"Q1\", \"Q2\"\nEach question is written in kaggle_survey_2021_answer_choices.pdf linked in this Kaggle notebook.\nYou can see full question statements from above pdf file.\n\nI summerized what I missed frequently and hope these notes will help you to understand data\n\n## 1. Sub columes\nWhile checking the name of columns, you can check a some \"Part_1\" suffix.\nThese column represents they are the result of multiple choice questions.\n\nFor example, question Q7 is \"What programming languages do you use on a regular basis? (Select all that apply)\"\nYou can see \"(Select all that apply)\" at the end of the question. \nSo selections are seperated on each sub columes \n\n## 2. Question Title\nThe first row of each columes has the question string.\nSo you need to remove the first row to visualize or rearrage data\n\n## 3. NaN\nBecause respondent selects some of multiple choice or some other reaseon, there are NaN data in the dataframe.\nThis NaN data should be removed or filled with proper values.\n","metadata":{}},{"cell_type":"code","source":"df.head(10)","metadata":{"execution":{"iopub.status.busy":"2022-07-12T13:59:13.616326Z","iopub.execute_input":"2022-07-12T13:59:13.616733Z","iopub.status.idle":"2022-07-12T13:59:13.655210Z","shell.execute_reply.started":"2022-07-12T13:59:13.616703Z","shell.execute_reply":"2022-07-12T13:59:13.654088Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Util Functions\n\nUtil Functions helps make a dataframe applying value_counts() function. \n\nIt returns keys of column values and count value of each keys.\nIn case of multiple choice columes, values are seperated with each sub columes.\nSo a make_df_multichoice() function concatenate all these seperated sub columes as one colume\n\nI used a different util function based on colume type and plotly uses a returned dataframe.","metadata":{}},{"cell_type":"code","source":"def make_df_countplot(df, isNormalize=False, *args):\n  src_key = args[0]\n  dst_key = args[1]\n  dst_value = args[2]\n  convert_type = args[3]\n  \n  if isNormalize == True:\n    df_temp = pd.DataFrame(round(df[src_key].value_counts(normalize=isNormalize)*100))\n  else:\n    df_temp = pd.DataFrame(df[src_key].value_counts())\n  df_temp[dst_key] = df_temp.index\n  df_temp.columns = [dst_value, dst_key]\n  df_temp = df_temp.reset_index().drop('index', axis=1)[[dst_key, dst_value]]\n  df_temp = df_temp[:-1]\n  df_temp = df_temp.astype({dst_key : convert_type})\n\n  return df_temp\n\ndef make_df_multichoice(df, isNormalize, ar):\n  src_key = ar[0]\n  dst_key = ar[1]\n  dst_value = ar[2]\n  convert_type = ar[3]\n  list_of_Q = ar[4]\n\n  df_choice = pd.concat(list_of_Q, axis=0, ignore_index=True).dropna().astype(convert_type)\n  df_temp = pd.DataFrame(round(df_choice.value_counts(normalize=True)*100))\n  df_temp[dst_key] = df_temp.index\n  df_temp.columns = [dst_value, dst_key]\n  df_temp = df_temp.reset_index().drop('index', axis=1)[[dst_key, dst_value]]\n  df_temp = df_temp.astype({dst_key :convert_type})\n  return df_temp\n\n","metadata":{"execution":{"iopub.status.busy":"2022-07-12T13:59:37.587774Z","iopub.execute_input":"2022-07-12T13:59:37.588213Z","iopub.status.idle":"2022-07-12T13:59:37.603768Z","shell.execute_reply.started":"2022-07-12T13:59:37.588178Z","shell.execute_reply":"2022-07-12T13:59:37.602283Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Q2 : What is your gender?","metadata":{}},{"cell_type":"code","source":"src_key = 'Q2'\ndst_key = 'gender'\ndst_value = 'count'\nconvert_type = 'string'\n\ndf_gender = make_df_countplot(df, False, src_key, dst_key, dst_value, convert_type)\ndf_gender","metadata":{"execution":{"iopub.status.busy":"2022-07-12T14:02:32.992143Z","iopub.execute_input":"2022-07-12T14:02:32.992564Z","iopub.status.idle":"2022-07-12T14:02:33.024088Z","shell.execute_reply.started":"2022-07-12T14:02:32.992530Z","shell.execute_reply":"2022-07-12T14:02:33.022956Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = px.bar(df_gender, color=dst_value, y=dst_value, x=dst_key, width=1000, text=dst_value )\nfig.update_xaxes(tickangle=-45)\nfig.update_layout(xaxis={'categoryorder':'total descending'} )\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-12T14:02:40.803022Z","iopub.execute_input":"2022-07-12T14:02:40.803444Z","iopub.status.idle":"2022-07-12T14:02:41.921454Z","shell.execute_reply.started":"2022-07-12T14:02:40.803409Z","shell.execute_reply":"2022-07-12T14:02:41.920010Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Q3 : In which country do you currently reside?","metadata":{}},{"cell_type":"code","source":"src_key = 'Q3'\ndst_key = 'country'\ndst_value = 'ratio'\nconvert_type = 'string'\n\ndf_country = make_df_countplot(df, True, src_key, dst_key, dst_value, convert_type)\ndf_country.sort_values(by=[dst_value, dst_key], ascending=[False, False] ).reset_index(drop=True).head(20)","metadata":{"execution":{"iopub.status.busy":"2022-07-12T14:09:04.638384Z","iopub.execute_input":"2022-07-12T14:09:04.638756Z","iopub.status.idle":"2022-07-12T14:09:04.666868Z","shell.execute_reply.started":"2022-07-12T14:09:04.638728Z","shell.execute_reply":"2022-07-12T14:09:04.665748Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = px.bar(df_country[:20], color=dst_value, y=dst_value, x=dst_key, width=1000, text=dst_value )\nfig.update_xaxes(tickangle=-45)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-12T14:09:13.916855Z","iopub.execute_input":"2022-07-12T14:09:13.917260Z","iopub.status.idle":"2022-07-12T14:09:13.990240Z","shell.execute_reply.started":"2022-07-12T14:09:13.917226Z","shell.execute_reply":"2022-07-12T14:09:13.989364Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Q4 : Education Degree","metadata":{}},{"cell_type":"code","source":"df['Q4'].replace([\"Bachelorâs degree\"], \"Bachelor's degree\",inplace=True)\ndf['Q4'].replace([\"Masterâs degree\"], \"Master's degree\",inplace=True)\ndf['Q4'].replace([\"Some college/university study without earning a bachelorâs degree\"], \"Some college/university study without earning a Bachelor's degree\",inplace=True)\n\nsrc_key = 'Q4'\ndst_key = 'education'\ndst_value = 'ratio'\nconvert_type = 'string'\n\ndf_degree = make_df_countplot(df, True, src_key, dst_key, dst_value, convert_type)\ndf_degree.sort_values(by=dst_value, ascending=False ).head(20)","metadata":{"execution":{"iopub.status.busy":"2022-07-12T14:10:14.254104Z","iopub.execute_input":"2022-07-12T14:10:14.255291Z","iopub.status.idle":"2022-07-12T14:10:14.295823Z","shell.execute_reply.started":"2022-07-12T14:10:14.255243Z","shell.execute_reply":"2022-07-12T14:10:14.294550Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = px.bar(df_degree, color=dst_value, y=dst_value, x=dst_key, width=1000, text=dst_value )\nfig.update_xaxes(tickangle=-45)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-12T14:10:25.455553Z","iopub.execute_input":"2022-07-12T14:10:25.455963Z","iopub.status.idle":"2022-07-12T14:10:25.534598Z","shell.execute_reply.started":"2022-07-12T14:10:25.455931Z","shell.execute_reply":"2022-07-12T14:10:25.533581Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Q5 : Select the title most similar to your current role (or most recent title if retired):","metadata":{}},{"cell_type":"code","source":"src_key = 'Q5'\ndst_key = 'role'\ndst_value = 'ratio'\nconvert_type = 'string'\n\ndf_role = make_df_countplot(df, True, src_key, dst_key, dst_value, convert_type)\ndf_role","metadata":{"execution":{"iopub.status.busy":"2022-07-12T14:10:55.197392Z","iopub.execute_input":"2022-07-12T14:10:55.197778Z","iopub.status.idle":"2022-07-12T14:10:55.223751Z","shell.execute_reply.started":"2022-07-12T14:10:55.197749Z","shell.execute_reply":"2022-07-12T14:10:55.222588Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = px.bar(df_role, color=dst_value, y=dst_value, x=dst_key, width=1000 , text=dst_value)\nfig.update_xaxes(tickangle=-45)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-12T14:11:00.492024Z","iopub.execute_input":"2022-07-12T14:11:00.492422Z","iopub.status.idle":"2022-07-12T14:11:00.566079Z","shell.execute_reply.started":"2022-07-12T14:11:00.492391Z","shell.execute_reply":"2022-07-12T14:11:00.564843Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Q7 : What programming languages do you use on a regular basis? (Select all that apply)","metadata":{}},{"cell_type":"code","source":"src_key = 'Q7'\ndst_key = 'language'\ndst_value = 'ratio'\nlist_of_Q = [(df[f'{src_key}_Part_{x}'][1:]) for x in range(1,13) ]\nlist_of_Q.append(df[f'{src_key}_OTHER'][1:])\n\n\ndf_basic_progrm_lang = make_df_multichoice(df, True, (src_key, dst_key, dst_value, convert_type, list_of_Q))\ndf_basic_progrm_lang","metadata":{"execution":{"iopub.status.busy":"2022-07-12T14:11:26.037433Z","iopub.execute_input":"2022-07-12T14:11:26.037813Z","iopub.status.idle":"2022-07-12T14:11:26.095420Z","shell.execute_reply.started":"2022-07-12T14:11:26.037784Z","shell.execute_reply":"2022-07-12T14:11:26.094135Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = px.bar(df_basic_progrm_lang, color=dst_value, y=dst_value, x=dst_key, width=1000 , text=dst_value)\nfig.update_xaxes(tickangle=-45)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-12T14:11:32.307495Z","iopub.execute_input":"2022-07-12T14:11:32.307892Z","iopub.status.idle":"2022-07-12T14:11:32.387890Z","shell.execute_reply.started":"2022-07-12T14:11:32.307860Z","shell.execute_reply":"2022-07-12T14:11:32.386758Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Q8 : What programming language would you recommend an aspiring data scientist to learn first?","metadata":{}},{"cell_type":"code","source":"src_key = 'Q8'\ndst_key = 'recommand_lang'\ndst_value = 'ratio'\nconvert_type = 'string'\n\ndf_recommand_lang = make_df_countplot(df, True, src_key, dst_key, dst_value, convert_type)\ndf_recommand_lang","metadata":{"execution":{"iopub.status.busy":"2022-07-12T14:12:24.133723Z","iopub.execute_input":"2022-07-12T14:12:24.134131Z","iopub.status.idle":"2022-07-12T14:12:24.160359Z","shell.execute_reply.started":"2022-07-12T14:12:24.134101Z","shell.execute_reply":"2022-07-12T14:12:24.159188Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = px.bar(df_recommand_lang, color=dst_value, y=dst_value, x=dst_key, width=1000 , text=dst_value)\nfig.update_xaxes(tickangle=-45)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-12T14:12:26.247294Z","iopub.execute_input":"2022-07-12T14:12:26.247752Z","iopub.status.idle":"2022-07-12T14:12:26.333561Z","shell.execute_reply.started":"2022-07-12T14:12:26.247717Z","shell.execute_reply":"2022-07-12T14:12:26.332391Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Basic language vs Recommand language","metadata":{}},{"cell_type":"code","source":"df_lang_diff = pd.concat([df_basic_progrm_lang, df_recommand_lang], axis = 1)\ndf_lang_diff.columns = ['basic_lang', 'basic_ratio', 'recommand_lang', 'recommand_ratio']\ndf_lang_diff","metadata":{"execution":{"iopub.status.busy":"2022-07-12T14:13:17.516639Z","iopub.execute_input":"2022-07-12T14:13:17.517103Z","iopub.status.idle":"2022-07-12T14:13:17.534125Z","shell.execute_reply.started":"2022-07-12T14:13:17.517066Z","shell.execute_reply":"2022-07-12T14:13:17.532658Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"trace_basic_lang = go.Bar(x=df_lang_diff['basic_lang'],  y=df_lang_diff['basic_ratio'], name = 'basic_lang', text=df_lang_diff['basic_ratio'])\ntrace_basic_recommand = go.Bar(x=df_lang_diff['recommand_lang'],  y=df_lang_diff['recommand_ratio'], name = 'recommand_lang', text=df_lang_diff['recommand_ratio'] )\n\nfig = go.Figure([trace_basic_lang, trace_basic_recommand])\nfig.update_layout(xaxis={'categoryorder':'total descending'} , title = 'Basic Language vs Recommand Language') # add only this line\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-12T14:13:21.763504Z","iopub.execute_input":"2022-07-12T14:13:21.763904Z","iopub.status.idle":"2022-07-12T14:13:21.784392Z","shell.execute_reply.started":"2022-07-12T14:13:21.763873Z","shell.execute_reply":"2022-07-12T14:13:21.781977Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Q9 : Which of the following integrated development environments (IDE's) do you use on a regular basis? (Select all that apply)","metadata":{}},{"cell_type":"code","source":"src_key = 'Q9'\ndst_key = 'IDE'\ndst_value = 'ratio'\nlist_of_Q = [(df[f'{src_key}_Part_{x}'][1:]) for x in range(1,13) ]\nlist_of_Q.append(df[f'{src_key}_OTHER'][1:])\n\n\ndf_ide = make_df_multichoice(df, True, (src_key, dst_key, dst_value, convert_type, list_of_Q))\ndf_ide","metadata":{"execution":{"iopub.status.busy":"2022-07-12T14:14:00.331174Z","iopub.execute_input":"2022-07-12T14:14:00.331995Z","iopub.status.idle":"2022-07-12T14:14:00.391276Z","shell.execute_reply.started":"2022-07-12T14:14:00.331948Z","shell.execute_reply":"2022-07-12T14:14:00.389980Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = px.bar(df_ide, color=dst_value, y=dst_value, x=dst_key, width=1000 , text=dst_value)\nfig.update_xaxes(tickangle=-45)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-12T14:14:09.913344Z","iopub.execute_input":"2022-07-12T14:14:09.913727Z","iopub.status.idle":"2022-07-12T14:14:09.986012Z","shell.execute_reply.started":"2022-07-12T14:14:09.913696Z","shell.execute_reply":"2022-07-12T14:14:09.984747Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Q14 : What data visualization libraries or tools do you use on a regular basis?","metadata":{}},{"cell_type":"code","source":"src_key = 'Q14'\ndst_key = 'visualization lib'\ndst_value = 'ratio'\nlist_of_Q = [(df[f'{src_key}_Part_{x}'][1:]) for x in range(1,12) ]\nlist_of_Q.append(df[f'{src_key}_OTHER'][1:])\n\n\ndf_visual_lib = make_df_multichoice(df, True, (src_key, dst_key, dst_value, convert_type, list_of_Q))\ndf_visual_lib","metadata":{"execution":{"iopub.status.busy":"2022-07-12T14:14:32.638676Z","iopub.execute_input":"2022-07-12T14:14:32.639514Z","iopub.status.idle":"2022-07-12T14:14:32.694274Z","shell.execute_reply.started":"2022-07-12T14:14:32.639470Z","shell.execute_reply":"2022-07-12T14:14:32.693403Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = px.bar(df_visual_lib, color=dst_value, y=dst_value, x=dst_key, width=1000 , text=dst_value)\nfig.update_xaxes(tickangle=-45)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-12T14:14:38.640971Z","iopub.execute_input":"2022-07-12T14:14:38.641377Z","iopub.status.idle":"2022-07-12T14:14:38.718542Z","shell.execute_reply.started":"2022-07-12T14:14:38.641331Z","shell.execute_reply":"2022-07-12T14:14:38.717381Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Q15 : For how many years have you used machine learning methods?","metadata":{}},{"cell_type":"code","source":"src_key = 'Q15'\ndst_key = 'leaning years'\ndst_value = 'ratio'\nconvert_type = 'string'\n\ndf_learning_years = make_df_countplot(df, True, src_key, dst_key, dst_value, convert_type)\ndf_learning_years","metadata":{"execution":{"iopub.status.busy":"2022-07-12T14:14:58.394649Z","iopub.execute_input":"2022-07-12T14:14:58.395031Z","iopub.status.idle":"2022-07-12T14:14:58.418627Z","shell.execute_reply.started":"2022-07-12T14:14:58.395002Z","shell.execute_reply":"2022-07-12T14:14:58.417579Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = px.bar(df_learning_years, color=dst_value, y=dst_value, x=dst_key, width=1000 , text=dst_value)\nfig.update_xaxes(tickangle=-45)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-12T14:15:04.665892Z","iopub.execute_input":"2022-07-12T14:15:04.666360Z","iopub.status.idle":"2022-07-12T14:15:04.743499Z","shell.execute_reply.started":"2022-07-12T14:15:04.666317Z","shell.execute_reply":"2022-07-12T14:15:04.742393Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Q16 : Which of the following machine learning frameworks do you use on a regular basis?","metadata":{}},{"cell_type":"code","source":"src_key = 'Q16'\ndst_key = 'ML lib'\ndst_value = 'ratio'\nlist_of_Q = [(df[f'{src_key}_Part_{x}'][1:]) for x in range(1,18) ]\nlist_of_Q.append(df[f'{src_key}_OTHER'][1:])\n\n\ndf_visual_lib = make_df_multichoice(df, True, (src_key, dst_key, dst_value, convert_type, list_of_Q))\ndf_visual_lib","metadata":{"execution":{"iopub.status.busy":"2022-07-12T14:15:26.776291Z","iopub.execute_input":"2022-07-12T14:15:26.776822Z","iopub.status.idle":"2022-07-12T14:15:26.859689Z","shell.execute_reply.started":"2022-07-12T14:15:26.776775Z","shell.execute_reply":"2022-07-12T14:15:26.858623Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = px.bar(df_visual_lib, color=dst_value, y=dst_value, x=dst_key, width=1000 , text=dst_value)\nfig.update_xaxes(tickangle=-45)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-12T14:15:34.171904Z","iopub.execute_input":"2022-07-12T14:15:34.172457Z","iopub.status.idle":"2022-07-12T14:15:34.255155Z","shell.execute_reply.started":"2022-07-12T14:15:34.172408Z","shell.execute_reply":"2022-07-12T14:15:34.254196Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Q17 : Which of the following ML algorithms do you use on a regular basis?","metadata":{}},{"cell_type":"code","source":"src_key = 'Q17'\ndst_key = 'ML algorithm'\ndst_value = 'ratio'\nlist_of_Q = [(df[f'{src_key}_Part_{x}'][1:]) for x in range(1,12) ]\nlist_of_Q.append(df[f'{src_key}_OTHER'][1:])\n\n\ndf_ml_algo = make_df_multichoice(df, True, (src_key, dst_key, dst_value, convert_type, list_of_Q))\ndf_ml_algo","metadata":{"execution":{"iopub.status.busy":"2022-07-12T14:16:07.531581Z","iopub.execute_input":"2022-07-12T14:16:07.532066Z","iopub.status.idle":"2022-07-12T14:16:07.597294Z","shell.execute_reply.started":"2022-07-12T14:16:07.532026Z","shell.execute_reply":"2022-07-12T14:16:07.596105Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = px.bar(df_ml_algo, color=dst_value, y=dst_value, x=dst_key, width=1000 , text=dst_value)\nfig.update_xaxes(tickangle=-45)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-12T14:16:09.741039Z","iopub.execute_input":"2022-07-12T14:16:09.741437Z","iopub.status.idle":"2022-07-12T14:16:09.816138Z","shell.execute_reply.started":"2022-07-12T14:16:09.741404Z","shell.execute_reply":"2022-07-12T14:16:09.814841Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Q18 : Which categories of computer vision methods do you use on a regular basis?","metadata":{}},{"cell_type":"code","source":"src_key = 'Q18'\ndst_key = 'Vision Method'\ndst_value = 'ratio'\nlist_of_Q = [(df[f'{src_key}_Part_{x}'][1:]) for x in range(1,7) ]\nlist_of_Q.append(df[f'{src_key}_OTHER'][1:])\n\n\ndf_vision_method = make_df_multichoice(df, True, (src_key, dst_key, dst_value, convert_type, list_of_Q))\ndf_vision_method\n","metadata":{"execution":{"iopub.status.busy":"2022-07-12T14:17:44.343872Z","iopub.execute_input":"2022-07-12T14:17:44.344284Z","iopub.status.idle":"2022-07-12T14:17:44.382555Z","shell.execute_reply.started":"2022-07-12T14:17:44.344254Z","shell.execute_reply":"2022-07-12T14:17:44.381299Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = px.bar(df_vision_method, color=dst_value, y=dst_key, x=dst_value, width=1000 , text=dst_value)\nfig.update_xaxes(tickangle=-45)\nfig.update_layout(yaxis={'categoryorder':'total ascending'} , title = 'Vison Method on a regular basis') # add only this line\n\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-12T14:17:48.512584Z","iopub.execute_input":"2022-07-12T14:17:48.513280Z","iopub.status.idle":"2022-07-12T14:17:48.589456Z","shell.execute_reply.started":"2022-07-12T14:17:48.513239Z","shell.execute_reply":"2022-07-12T14:17:48.588282Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Q20 : In what industry is your current employer/contract (or your most recent employer if retired)?","metadata":{}},{"cell_type":"code","source":"src_key = 'Q20'\ndst_key = 'Industry'\ndst_value = 'ratio'\nconvert_type = 'string'\n\ndf_industry = make_df_countplot(df, True, src_key, dst_key, dst_value, convert_type)\ndf_industry","metadata":{"execution":{"iopub.status.busy":"2022-07-12T14:18:13.727647Z","iopub.execute_input":"2022-07-12T14:18:13.728038Z","iopub.status.idle":"2022-07-12T14:18:13.754764Z","shell.execute_reply.started":"2022-07-12T14:18:13.728009Z","shell.execute_reply":"2022-07-12T14:18:13.753635Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = px.bar(df_industry, color=dst_value, y=dst_value, x=dst_key, width=1000 , text=dst_value)\nfig.update_xaxes(tickangle=-45)\nfig.update_layout(yaxis={'categoryorder':'total ascending'} , title = 'industry') # add only this line\n\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-12T14:18:23.941913Z","iopub.execute_input":"2022-07-12T14:18:23.942461Z","iopub.status.idle":"2022-07-12T14:18:24.293601Z","shell.execute_reply.started":"2022-07-12T14:18:23.942412Z","shell.execute_reply":"2022-07-12T14:18:24.292360Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Q23 : Does your current employer incorporate machine learning methods into their business?","metadata":{}},{"cell_type":"code","source":"src_key = 'Q23'\ndst_key = 'ML methods into the business'\ndst_value = 'ratio'\nconvert_type = 'string'\n\ndf_ml_into_business = make_df_countplot(df, True, src_key, dst_key, dst_value, convert_type)\ndf_ml_into_business","metadata":{"execution":{"iopub.status.busy":"2022-07-12T14:18:43.892061Z","iopub.execute_input":"2022-07-12T14:18:43.892607Z","iopub.status.idle":"2022-07-12T14:18:43.915237Z","shell.execute_reply.started":"2022-07-12T14:18:43.892562Z","shell.execute_reply":"2022-07-12T14:18:43.914058Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = px.bar(df_ml_into_business, color=dst_value, y=dst_key, x=dst_value, width=1000, text=dst_value)\nfig.update_xaxes(tickangle=-45)\nfig.update_layout(yaxis={'categoryorder':'total ascending'} , title = 'industry') # add only this line\nfig.update_layout()","metadata":{"execution":{"iopub.status.busy":"2022-07-12T14:18:52.082260Z","iopub.execute_input":"2022-07-12T14:18:52.083176Z","iopub.status.idle":"2022-07-12T14:18:52.159998Z","shell.execute_reply.started":"2022-07-12T14:18:52.083131Z","shell.execute_reply":"2022-07-12T14:18:52.158898Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Q27-A : Which of the following cloud computing platforms do you use on a regular basis?","metadata":{}},{"cell_type":"code","source":"src_key = 'Q27_A'\ndst_key = 'cloud'\ndst_value = 'ratio'\nlist_of_Q = [(df[f'{src_key}_Part_{x}'][1:]) for x in range(1,12) ]\nlist_of_Q.append(df[f'{src_key}_OTHER'][1:])\n\n\ndf_cloud = make_df_multichoice(df, True, (src_key, dst_key, dst_value, convert_type, list_of_Q))\ndf_cloud","metadata":{"execution":{"iopub.status.busy":"2022-07-12T14:19:28.346488Z","iopub.execute_input":"2022-07-12T14:19:28.346887Z","iopub.status.idle":"2022-07-12T14:19:28.390626Z","shell.execute_reply.started":"2022-07-12T14:19:28.346857Z","shell.execute_reply":"2022-07-12T14:19:28.389402Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = px.bar(df_cloud, color=dst_value, x=dst_key, y=dst_value, text=dst_value, title=\"Cloud computing platform\")\nfig.update_layout(yaxis={'categoryorder':'total ascending'})\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-12T14:19:30.383824Z","iopub.execute_input":"2022-07-12T14:19:30.384657Z","iopub.status.idle":"2022-07-12T14:19:30.459958Z","shell.execute_reply.started":"2022-07-12T14:19:30.384604Z","shell.execute_reply":"2022-07-12T14:19:30.458651Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Google Cloud vs AWS Cloud by Industry","metadata":{}},{"cell_type":"code","source":"df_aws = df[['Q20', 'Q27_A_Part_1']][1:].dropna()\ndf_google = df[['Q20', 'Q27_A_Part_2']][1:].dropna()\n\ndf_aws.columns = ['industry', 'cloud']\ndf_google.columns = ['industry', 'cloud']\ndf_clouds = pd.concat([df_aws, df_google], axis=0)\n\ngrouped = df_clouds.groupby('cloud')\nfor name, group in grouped:\n  if name == 'Amazon Web Services (AWS)':\n    df_aws = group\n  if name == 'Microsoft Azure':\n    df_google = group\n\nsrc_key = 'industry'\ndst_key = 'Industry'\ndst_value = 'count'\nconvert_type = 'string'\ndf_aws_industry = make_df_countplot(df_aws, False, src_key, dst_key, dst_value, convert_type)\ndf_google_industry = make_df_countplot(df_google, False, src_key, dst_key, dst_value, convert_type)\n\ntrace_aws = go.Bar(x=df_aws_industry['Industry'],  y=df_aws_industry[dst_value], name = 'AWS Cloud')\ntrace_google = go.Bar(x=df_google_industry['Industry'],  y=df_google_industry[dst_value], name = 'Google Cloud' )\n\nfig = go.Figure([trace_aws, trace_google])\nfig.update_layout(xaxis={'categoryorder':'total descending'} , title = 'Google Cloud vs AWS Cloud by Industry') # add only this line\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-12T14:19:51.712295Z","iopub.execute_input":"2022-07-12T14:19:51.712681Z","iopub.status.idle":"2022-07-12T14:19:51.762675Z","shell.execute_reply.started":"2022-07-12T14:19:51.712653Z","shell.execute_reply":"2022-07-12T14:19:51.761458Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Q34-A : Which of the following business intelligence tools do you use on a regular basis?","metadata":{}},{"cell_type":"code","source":"src_key = 'Q35'\ndst_key = 'BI tool'\ndst_value = 'ratio'\nconvert_type = 'string'\n\n\ndf_BI = make_df_countplot(df, True, src_key, dst_key, dst_value, convert_type)\ndf_BI","metadata":{"execution":{"iopub.status.busy":"2022-07-12T14:20:24.982491Z","iopub.execute_input":"2022-07-12T14:20:24.982872Z","iopub.status.idle":"2022-07-12T14:20:25.006123Z","shell.execute_reply.started":"2022-07-12T14:20:24.982844Z","shell.execute_reply":"2022-07-12T14:20:25.004935Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = px.bar(df_BI, color=dst_value, x=dst_key, y=dst_value, text=dst_value, title=\"BI tool\")\nfig.update_layout(yaxis={'categoryorder':'total ascending'})\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-12T14:20:32.059055Z","iopub.execute_input":"2022-07-12T14:20:32.059428Z","iopub.status.idle":"2022-07-12T14:20:32.136853Z","shell.execute_reply.started":"2022-07-12T14:20:32.059399Z","shell.execute_reply":"2022-07-12T14:20:32.136041Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# BI analysis by count","metadata":{}},{"cell_type":"code","source":"df_BI = df[['Q3', 'Q35', 'Q20']][1:].dropna()\ndf_BI.columns = ['Country', 'BI', 'Industry']\ndf_BI = df_BI[ ((df_BI['Industry'] == 'Accounting/Finance') | (df_BI['Industry'] == 'Computers/Technology')| (df_BI['Industry'] == 'Academics/Education') | (df_BI['Industry'] == 'Retail/Sales') |(df_BI['Industry'] == 'Manufacturing/Fabrications') | (df_BI['Industry'] == 'Medical/Pharmaceutical')) &\n                ((df_BI['Country'] == 'South Korea') | (df_BI['Country'] == 'India') | (df_BI['Country'] == 'Japan') | (df_BI['Country'] == 'United States of America'))]\nfig = px.scatter(df_BI, x=\"Industry\", y=\"BI\", color=\"BI\" , facet_col='Country')\nfig.update_xaxes(tickangle=-45)\n\nfig.show()\n","metadata":{"execution":{"iopub.status.busy":"2022-07-12T14:20:52.067975Z","iopub.execute_input":"2022-07-12T14:20:52.068398Z","iopub.status.idle":"2022-07-12T14:20:52.458018Z","shell.execute_reply.started":"2022-07-12T14:20:52.068364Z","shell.execute_reply":"2022-07-12T14:20:52.456740Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Q39 : Where do you publicly share or deploy your data analysis or machine learning applications?","metadata":{}},{"cell_type":"code","source":"src_key = 'Q39'\ndst_key = 'Deploy Method'\ndst_value = 'ratio'\nlist_of_Q = [(df[f'{src_key}_Part_{x}'][1:]) for x in range(1,10) ]\nlist_of_Q.append(df[f'{src_key}_OTHER'][1:])\n\n\ndf_deploy = make_df_multichoice(df, True, (src_key, dst_key, dst_value, convert_type, list_of_Q))\ndf_deploy","metadata":{"execution":{"iopub.status.busy":"2022-07-12T14:21:33.678113Z","iopub.execute_input":"2022-07-12T14:21:33.678534Z","iopub.status.idle":"2022-07-12T14:21:33.721350Z","shell.execute_reply.started":"2022-07-12T14:21:33.678497Z","shell.execute_reply":"2022-07-12T14:21:33.720097Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = px.bar(df_deploy, color=dst_value, x=dst_key, y=dst_value, text=dst_value, title=\"Deploy Method\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-12T14:21:38.000617Z","iopub.execute_input":"2022-07-12T14:21:38.001662Z","iopub.status.idle":"2022-07-12T14:21:38.074152Z","shell.execute_reply.started":"2022-07-12T14:21:38.001625Z","shell.execute_reply":"2022-07-12T14:21:38.073067Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Q40 : On which platforms have you begun or completed data science courses? ","metadata":{}},{"cell_type":"code","source":"src_key = 'Q40'\ndst_key = 'Learning Course'\ndst_value = 'ratio'\nlist_of_Q = [(df[f'{src_key}_Part_{x}'][1:]) for x in range(1,12) ]\nlist_of_Q.append(df[f'{src_key}_OTHER'][1:])\n\n\ndf_course = make_df_multichoice(df, True, (src_key, dst_key, dst_value, convert_type, list_of_Q))\ndf_course = make_df_multichoice(df, True, (src_key, dst_key, dst_value, convert_type, list_of_Q))\ndf_course","metadata":{"execution":{"iopub.status.busy":"2022-07-12T14:22:03.473343Z","iopub.execute_input":"2022-07-12T14:22:03.473775Z","iopub.status.idle":"2022-07-12T14:22:03.565644Z","shell.execute_reply.started":"2022-07-12T14:22:03.473739Z","shell.execute_reply":"2022-07-12T14:22:03.564577Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = px.bar(df_course, color=dst_value, x=dst_key, y=dst_value, text=dst_value, title=\"Learning Course\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-12T14:22:05.600763Z","iopub.execute_input":"2022-07-12T14:22:05.601153Z","iopub.status.idle":"2022-07-12T14:22:05.675074Z","shell.execute_reply.started":"2022-07-12T14:22:05.601121Z","shell.execute_reply":"2022-07-12T14:22:05.673835Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"","metadata":{}}]}