{"cells":[{"metadata":{"_uuid":"3ea5b23ea6b838f804540afee33d9f81c3e2bbbf","_execution_state":"idle"},"cell_type":"markdown","source":"## Text mining on description :\n\n**This Kernel is dedicated to a code to transform the description in features usable for our model. **\n\nThe code is coming mainly from [**Datasciencedojo**](https://datasciencedojo.com/), a very good website to learn more on Data Science.\n\n![Imgur](https://i.imgur.com/tqfwhDv.png)\n\nAs most of the challenger, these data haven't brought much value on my final model, but I still think that the process is interesting and might be useful for other project :)"},{"metadata":{"trusted":true,"_uuid":"1fdb5c3e0bebbbe4a2bd9942e4b65ccf46475d89"},"cell_type":"code","source":"library(data.table)\ntrain<-fread(\"../input/train/train.csv\")\ntest<-fread(\"../input/test/test.csv\")\n\ndtb <- rbindlist(list(test, train), use.names = TRUE, fill = TRUE)\n\ndtb_text<-dtb[,c(\"PetID\",\"Description\")]\n\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a4dc613593d48d675f49883a6792caaf1bf406ed"},"cell_type":"markdown","source":"### Tokenize the description with the library quanteda :"},{"metadata":{"trusted":true,"_uuid":"212e6b47335ceee29ed41f80bfaf85bb2256a8c7"},"cell_type":"code","source":"library(quanteda)\n#lower all cases\ndtb_text$Description<-tolower(dtb_text$Description)\n# Tokenization\ndtb.tokens <- tokens(dtb_text$Description, what = \"word\", \n                       remove_numbers = TRUE, remove_punct = TRUE,\n                       remove_symbols = TRUE, remove_hyphens = TRUE)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7192fba08d45d5771e8006b0a17beea2579f5732"},"cell_type":"markdown","source":"**Cleaning the tokens :**\n* Remove stop words\n* Apply stemming to keep only the root of the words"},{"metadata":{"trusted":true,"_uuid":"aeb15ebe4c798a584fce8939f1d38b4b18fff314"},"cell_type":"code","source":"dtb.tokens <- tokens_select(dtb.tokens, stopwords(),selection = \"remove\")\n# Perform stemming on the tokens.                             \ndtb.tokens <- tokens_wordstem(dtb.tokens, language = \"english\")\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d65ed6f94035c4d187fea19f159d1f5d816314f2"},"cell_type":"markdown","source":"**Bag of words :**"},{"metadata":{"trusted":true,"_uuid":"f16907fcd2db3f2fa36d8641512845f5d302badb"},"cell_type":"code","source":"dtb.tokens.dfm <- dfm(dtb.tokens, tolower = FALSE)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"44f1e0dccf350d11ea65ad9b77e32baadb51ed35"},"cell_type":"markdown","source":"The bag of words can already be used as input in your model."},{"metadata":{"_uuid":"cebe961305e73c9429a5fa27865dc6e26a74d74f"},"cell_type":"markdown","source":"### TF-IDF\nTo go further we will perform TF-IDF. For more details about how it works, everything is available on Datasciencedojo's videos"},{"metadata":{"trusted":true,"_uuid":"29d6b789087091ece619a7b81ec440f852f16be1"},"cell_type":"code","source":"dtb.tokens.matrix <- as.matrix(dtb.tokens.dfm[,1:10000])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"19a3c64ea514ada83aa00992991cadd5f62881eb"},"cell_type":"markdown","source":"I keep only the first 10000 words for a matter of computing performance"},{"metadata":{"trusted":true,"_uuid":"26cf7a32f8c6d89e305b10315dfc997aa1c685a0"},"cell_type":"code","source":"\nterm.frequency <- function(row) {\n  row / sum(row)\n}\n\ninverse.doc.freq <- function(col) {\n  corpus.size <- length(col)\n  doc.count <- length(which(col > 0))\n  log10(corpus.size / doc.count)\n}\n\ntf.idf <- function(x, idf) {\n  x * idf\n}\n\ndtb.tokens.df <- apply(dtb.tokens.matrix, 1, term.frequency)\ndtb.tokens.idf <- apply(dtb.tokens.matrix, 2, inverse.doc.freq)\ndtb.tokens.tfidf <-  apply(dtb.tokens.df, 2, tf.idf, idf = dtb.tokens.idf)\ndtb.tokens.tfidf <- t(dtb.tokens.tfidf)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"9048b3ab5b7c221604a9cf39322f397b5b60f0dc"},"cell_type":"markdown","source":"The new database can be used for the model but it might be very heavy to run it. You may have to reduce the number of variables that you include."},{"metadata":{"_uuid":"1d8784908af98528231d66b6ea39834eeb129db2"},"cell_type":"markdown","source":"### SVD with irlba:\n\nNow we use SVD to summarize TF-IDF and bring more information in the model."},{"metadata":{"trusted":true,"_uuid":"db528524ef14cbc5382a82112724c432b74bd66b"},"cell_type":"code","source":"incomplete.cases <- which(!complete.cases(dtb.tokens.tfidf))\n# Fix incomplete cases\ndtb.tokens.tfidf[incomplete.cases,] <- rep(0.0, ncol(dtb.tokens.tfidf))\n\nlibrary(irlba)\n\n#  Reduce dimensionality down to 300 columns\ndtb.irlba <- irlba(t(dtb.tokens.tfidf), nv = 300, maxit = 1000)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"20bce7355d4c3da9a48923c7f747f467c85c4cef"},"cell_type":"code","source":"dtb.svd <- data.frame(dtb$PetID, dtb.irlba$v)\n\nwrite.csv(dtb.svd, \"SVD_all_300_v2.csv\",row.names=FALSE)","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"R","language":"R","name":"ir"},"language_info":{"mimetype":"text/x-r-source","name":"R","pygments_lexer":"r","version":"3.4.2","file_extension":".r","codemirror_mode":"r"}},"nbformat":4,"nbformat_minor":1}