{"cells":[{"metadata":{"_uuid":"a880bd7fdccaf4dee2200aa1f9540480c5888bd3","_execution_state":"idle","trusted":true,"_kg_hide-output":true},"cell_type":"code","source":"## Importing packages\nlibrary(tidyverse) # metapackage with lots of helpful functions\nlibrary(data.table)\nlibrary(stringi)\nlibrary(h2o)\n\n\ntrain_df <- fread('../input/train.csv')\ntest_df <- fread('../input/test.csv')\ntest_df$target<-as.integer(1)\n\nprint(\"Data and libraries Loaded\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"487b4d192364205de544232808404d1e0ad2d3fb"},"cell_type":"markdown","source":"Lets see some brief information about the train & test data sets especially on the columns, total number of rows."},{"metadata":{"trusted":true,"_uuid":"25c1868c3d5939f281be06a904a0add3c7f2965e"},"cell_type":"code","source":"str(train_df)\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ab02a88d0e036e5142a2e5747b14d825eff3b72d"},"cell_type":"code","source":"str(test_df)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"81ddbe1b70c2ecc7c9f65dfbb420b11959a7b931"},"cell_type":"markdown","source":"We have **1306122** rows in the training data set and **56370** rows in the testing data set.\n"},{"metadata":{"_uuid":"48d36ca8a758328f8aea64ca18d33cbbd7dc4eb8"},"cell_type":"markdown","source":"What is the frequency of the target variable in the training data set. As mentioned in the contest notes \"*target - a question labeled \"insincere\" has a value of 1, otherwise 0*\"\n"},{"metadata":{"trusted":true,"_uuid":"fadb8af5c184f2b26e6c0ded806473770125a3f9"},"cell_type":"code","source":"train_df%>%\ncount(target)%>%\nmutate(prop=prop.table(n)*100)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"70112eb03783b7a3b4e2cb8c1da99ec60474163f"},"cell_type":"markdown","source":"Just 6.1% have insincere comments while rest all are otherwise. Would have liked a higher percentage so that predictions could be more robust.\n\nI had read an interesting kernel at <[Link](https://www.kaggle.com/tunguz/just-some-simple-eda)> and it had details on creating new features that can be used for more EDA and later prediction.\n\n**Features**\n\n* Number of words in the text\n* Number of unique words in the text\n* Number of characters in the text\n* Number of stopwords\n* Number of punctuations\n* Number of upper case words\n* Number of title case words\n* Average length of the words\n\n\n"},{"metadata":{"trusted":true,"_uuid":"4923140cd30b70dc191f58d7e9a64fc9d7d28800"},"cell_type":"code","source":"# lets get some new variables added to the dataframe\ntrain_df%>%\n  mutate(ques_length=stri_length(question_text))->train_df\n\ntrain_df%>%\n  mutate(num_words=stri_count_regex(question_text,'\\\\w+'))->train_df\n\ntrain_df%>%\n  mutate(num_unique_words=lengths(lapply(str_split(question_text,'\\\\s+'), unique)))->train_df\n\ntrain_df%>%\n  mutate(avg_word_len=as.numeric(lapply(lapply(str_split(question_text,'\\\\s+'), str_length),mean)))->train_df\n\ntrain_df%>%\n  mutate(ucase_word_count=lengths(str_match_all(question_text, \"\\\\b[A-Z]{2,}\\\\b\")))->train_df\n\ntrain_df%>%\n  mutate(number_word_count=lengths(str_match_all(question_text, \"\\\\b[0-9]{2,}\\\\b\")))->train_df\n\ntrain_df%>%\n  mutate(nonalpha_word_count=lengths(str_match_all(question_text, \"[[:punct:]]{2,}\")))->train_df\n\nprint(\"New variables calculated in Train dataframe\")\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c11e48b9d1599e4a9bf1e4c8458fc78ff4e78e40"},"cell_type":"code","source":"str(train_df)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"8e689535555ae73b6441b514084adf37a21b6d90"},"cell_type":"code","source":"plot_facet<-function(in_col){\n  in_col<-enquo(in_col)\n  plot1<-train_df%>%\n    ggplot(aes(x = !!in_col)) +\n    geom_density() +\n    facet_wrap( ~ target, scales = \"free\")\n  \n  return(plot1)\n\n}","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"aeadcb4d596135cc33320c29fad0bc2a520dc360"},"cell_type":"markdown","source":"Lets do some comparision via density plots and see how the sincere and insincere relate."},{"metadata":{"trusted":true,"_uuid":"95d75b9025fbd5f6c9bf5cd5ef4b4f6d60b223fa"},"cell_type":"code","source":"options(repr.plot.width=4, repr.plot.height=3)\nplot_facet(ques_length)\nplot_facet(num_words)\nplot_facet(num_unique_words)\nplot_facet(avg_word_len)\nplot_facet(ucase_word_count)\nplot_facet(number_word_count)\nplot_facet(nonalpha_word_count)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d3c0d3868a976bfe5725a9297cc5c7dbe0cb9278"},"cell_type":"markdown","source":"Lets say that questions that have >3 words in upper case or >2 non alpha numeric chars we will consider it to be insincere else sincere.\n\nCalculate the new columns for the test_df"},{"metadata":{"trusted":true,"_uuid":"81645f3da694ef0ce9181f68f91e089f04726d7a"},"cell_type":"code","source":"test_df%>%\n  mutate(ques_length=stri_length(question_text))->test_df\n\ntest_df%>%\n  mutate(num_words=stri_count_regex(question_text,'\\\\w+'))->test_df\n\n\ntest_df%>%\n  mutate(num_unique_words=lengths(lapply(str_split(question_text,'\\\\s+'), unique)))->test_df\n\n\ntest_df%>%\n  mutate(avg_word_len=as.numeric(lapply(lapply(str_split(question_text,'\\\\s+'), str_length),mean)))->test_df\n\ntest_df%>%\n  mutate(ucase_word_count=lengths(str_match_all(question_text, \"\\\\b[A-Z]{2,}\\\\b\")))->test_df\n\ntest_df%>%\n  mutate(number_word_count=lengths(str_match_all(question_text, \"\\\\b[0-9]{2,}\\\\b\")))->test_df\n\ntest_df%>%\n  mutate(nonalpha_word_count=lengths(str_match_all(question_text, \"[[:punct:]]{2,}\")))->test_df\n\nprint(\"New variables calculated in Test dataframe\")\n\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"86602b59da17c0002d875568edf4486f30fe1eb1"},"cell_type":"code","source":"# test_df%>%\n# mutate(target=ifelse(ucase_word_count>=3 | nonalpha_word_count >2,1,0))->test_df\n\n# lets save out submission file\n\n# sub_df<- test_df%>%\n# select(qid,target)%>%\n# rename(prediction=target)\n# write.csv(sub_df,\"submission.csv\",row.names=F)\n## this got us to rank of 530 !!!\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"48b645966d8d440dd8bf89ea75431dd450e2b4ad"},"cell_type":"markdown","source":"Lets start the h2o engine and load our train & test dataframe onto it."},{"metadata":{"trusted":true,"_uuid":"64d72c501886cf8132163cc62ea5640e3ac87f04"},"cell_type":"code","source":"localH2O <- h2o.init(nthreads = -1)\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ad25c04da78887825199567e56f8bee27ff1d846"},"cell_type":"code","source":"h2o.init()\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"fa5ea77f240a9426c19b1aaf0249f6c3f1a8beaf"},"cell_type":"markdown","source":"Lets run the prediction by training on the new EDA columns were created earlier and then execute the model on the \"test_df\".\nYou will need to review the model information to gain insights and see if our independent variables are explaining the required dependent variable i.e \"target\""},{"metadata":{"trusted":true,"_uuid":"35df393fa21ecc583ed0695a98028bc646e66a88"},"cell_type":"code","source":"\n#data to h2o cluster\ntrain_df$target<- as.integer(train_df$target) \ntest_df$target<- as.integer(test_df$target) \n\n# create the h20 objects for test and train\ntrain.h2o <- as.h2o(train_df%>%select(-question_text))\ntest.h2o <- as.h2o(test_df%>%select(-question_text))\n\n## make the target variable as factor so that our prediction also returns it as 0/1. I see that in keeping the dtypes as numeric makes our predictions decimals !\n# yet to figure out the nuances of h2o and modelling !!\n\ntrain.h2o$target = as.factor(train.h2o$target) \n#dependent variable (target)\ny.dep <- 2\n\n#independent variables (dropping ID variables)\nx.indep <- c(3:9)\n\n\n\n#below is good.. need train_df as factor\nregression.model<-h2o.gbm(\n  training_frame = train.h2o,      ## H2O frame holding the training data\n  x=x.indep,                 ## this can be names or column numbers\n  y=\"target\",                   ## target: using the logged variable created earlier\n  model_id=\"gbm1\",              ## internal H2O name for model\n  ntrees = 200,                  ## use fewer trees than default (50) to speed up training\n  learn_rate = 0.1,             ## lower learn_rate is better, but use high rate to offset few trees\n  score_tree_interval = 3,      ## score every 3 trees\n  sample_rate = 0.5,            ## use half the rows each scoring round\n  col_sample_rate = 0.8        ## use 4/5 the columns to decide each split decision\n)\n\npredict.reg <- as.data.frame(h2o.predict(regression.model, test.h2o))\nprint(\"Prediction completed !\")\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"312eb70bcd48a48e4c917c9ca67dd638a4697566"},"cell_type":"markdown","source":"Create submission dataframe and write to the csv file as per submission guidelines. \nEnsure you keep the columns and their format as per the guidlines."},{"metadata":{"trusted":true,"_uuid":"000dca180232a47d3a3a098d29cc3f90e5cb70c5"},"cell_type":"code","source":"sub_df<- cbind(test_df%>%select(qid),predict.reg%>%select(predict)%>%rename(prediction=predict))\nwrite_excel_csv(sub_df,\"submission.csv\")","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"R","language":"R","name":"ir"},"language_info":{"mimetype":"text/x-r-source","name":"R","pygments_lexer":"r","version":"3.4.2","file_extension":".r","codemirror_mode":"r"}},"nbformat":4,"nbformat_minor":1}