{
  "id": 59790,
  "title": "Several questions on the method used",
  "url": "/competitions/avito-demand-prediction/discussion/59790",
  "author_name": "Ethan Sukhyun Hong",
  "post_date": "2018-06-27T05:54:49.571000",
  "votes": 0,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hi, it's quite late for the new discussion / or question since we only have few hours left for this competition, but still, I want to ask something regarding data processing methods that lots of people use in this competition. </p>\n\n<p>Since I am a beginner, I referred to many of others' work, and tried to adopt some of those with my own method. Therefore, there are several methods that I used / got improvement, but I do not thoroughly understand. So, even the competition is near to it's end, I want to ask question on several method to share knowledge (And I do believe there are many beginners refer to others' kernal but does not thoroughly understand everything :))</p>\n\n<ol>\n<li>About image processing\nThis was the first time I've dealt with Image processing. There were many suggested ways to extract image features that load image, and extract features with cv2.</li>\n</ol>\n\n<p>However, I learned that we could extract features without loading the image, but rather within the zip file directly. The code I used is:</p>\n\n<pre><code>def get_blurrness(file):\nexfile = zipped.read(file)\narr = np.frombuffer(exfile, np.uint8)\nif arr.size &amp;gt; 0:   # exclude dirs and blanks\n    imz = cv2.imdecode(arr, flags=cv2.COLOR_BGR2GRAY)\n    fm = cv2.Laplacian(imz, cv2.CV_64F).var()\nelse: \n    fm = -1\nreturn fm\n</code></pre>\n\n<p>This somehow gave me the image blurrness score, and I tried same method with extracting numper of key points with this code:</p>\n\n<pre><code>def keyp(img):\ntry:        \n    img = image_path + str(img) + \".jpg\"\n    exfile = zipped.read(img)\n    arr = np.frombuffer(exfile, np.uint8)\n\n    imz = cv2.imdecode(arr, 1)\n    fast = cv2.FastFeatureDetector_create()\n\n# find and draw the keypoints\n    kp = fast.detect(imz,None)\n    kp =len(kp)\n    return kp\nexcept:\n    return 0\n</code></pre>\n\n<p>However, I could not thoroughly understand how this process work, and if it really extract valid features. I want to hear any opinions / explanation in this process.</p>\n\n<ol>\n<li>About Tf - Idf processing</li>\n</ol>\n\n<p>Most of the kernals use TF - IDF processing for text features - title and description. As I understand, TF - IDF score shows statistical weight and importance of each words. However, I could not exactly understand how TF-IDF vectorizor works. <a href=\"https://www.kaggle.com/him4318/avito-lightgbm-with-ridge-feature-v-2-0/code\">This kernal</a> is the one I referred to a lot regarding TFIDF vectorization. The major process is: </p>\n\n<pre><code>print(\"\\n[TF-IDF] Term Frequency Inverse Document Frequency Stage\")\nrussian_stop = set(stopwords.words('russian'))\n\ntfidf_para = {\n    \"stop_words\": russian_stop,\n    \"analyzer\": 'word',\n    \"token_pattern\": r'\\w{1,}',\n    \"sublinear_tf\": True,\n    \"dtype\": np.float32,\n    \"norm\": 'l2',\n    #\"min_df\":5,\n    #\"max_df\":.9,\n    \"smooth_idf\":False\n}\n\n\ndef get_col(col_name): return lambda x: x[col_name]\n##I added to the max_features of the description. It did not change my score much but it may be worth investigating\nvectorizer = FeatureUnion([\n        ('description',TfidfVectorizer(\n            ngram_range=(1, 2),\n            max_features=17000,\n            **tfidf_para,\n            preprocessor=get_col('description'))),\n        ('title',CountVectorizer(\n            ngram_range=(1, 2),\n            stop_words = russian_stop,\n            #max_features=7000,\n            preprocessor=get_col('title')))\n    ])\n\nstart_vect=time.time()\n\n#Fit my vectorizer on the entire dataset instead of the training rows\n#Score improved by .0001\nvectorizer.fit(df.to_dict('records'))\n\nready_df = vectorizer.transform(df.to_dict('records'))\ntfvocab = vectorizer.get_feature_names()\n</code></pre>\n\n<p>Is what this process basically do : split all the words - give TFIDF score to all words - add each words to the columns - and state if each words are contained in each row (like one hot coding) / TFIDF score of the word? I wonder if I understood it right</p>",
  "messages": [
    {
      "id": 348690,
      "postDate": "2018-06-27T05:54:49.573Z",
      "content": "<p>Hi, it's quite late for the new discussion / or question since we only have few hours left for this competition, but still, I want to ask something regarding data processing methods that lots of people use in this competition. </p>\n\n<p>Since I am a beginner, I referred to many of others' work, and tried to adopt some of those with my own method. Therefore, there are several methods that I used / got improvement, but I do not thoroughly understand. So, even the competition is near to it's end, I want to ask question on several method to share knowledge (And I do believe there are many beginners refer to others' kernal but does not thoroughly understand everything :))</p>\n\n<ol>\n<li>About image processing\nThis was the first time I've dealt with Image processing. There were many suggested ways to extract image features that load image, and extract features with cv2.</li>\n</ol>\n\n<p>However, I learned that we could extract features without loading the image, but rather within the zip file directly. The code I used is:</p>\n\n<pre><code>def get_blurrness(file):\nexfile = zipped.read(file)\narr = np.frombuffer(exfile, np.uint8)\nif arr.size &amp;gt; 0:   # exclude dirs and blanks\n    imz = cv2.imdecode(arr, flags=cv2.COLOR_BGR2GRAY)\n    fm = cv2.Laplacian(imz, cv2.CV_64F).var()\nelse: \n    fm = -1\nreturn fm\n</code></pre>\n\n<p>This somehow gave me the image blurrness score, and I tried same method with extracting numper of key points with this code:</p>\n\n<pre><code>def keyp(img):\ntry:        \n    img = image_path + str(img) + \".jpg\"\n    exfile = zipped.read(img)\n    arr = np.frombuffer(exfile, np.uint8)\n\n    imz = cv2.imdecode(arr, 1)\n    fast = cv2.FastFeatureDetector_create()\n\n# find and draw the keypoints\n    kp = fast.detect(imz,None)\n    kp =len(kp)\n    return kp\nexcept:\n    return 0\n</code></pre>\n\n<p>However, I could not thoroughly understand how this process work, and if it really extract valid features. I want to hear any opinions / explanation in this process.</p>\n\n<ol>\n<li>About Tf - Idf processing</li>\n</ol>\n\n<p>Most of the kernals use TF - IDF processing for text features - title and description. As I understand, TF - IDF score shows statistical weight and importance of each words. However, I could not exactly understand how TF-IDF vectorizor works. <a href=\"https://www.kaggle.com/him4318/avito-lightgbm-with-ridge-feature-v-2-0/code\">This kernal</a> is the one I referred to a lot regarding TFIDF vectorization. The major process is: </p>\n\n<pre><code>print(\"\\n[TF-IDF] Term Frequency Inverse Document Frequency Stage\")\nrussian_stop = set(stopwords.words('russian'))\n\ntfidf_para = {\n    \"stop_words\": russian_stop,\n    \"analyzer\": 'word',\n    \"token_pattern\": r'\\w{1,}',\n    \"sublinear_tf\": True,\n    \"dtype\": np.float32,\n    \"norm\": 'l2',\n    #\"min_df\":5,\n    #\"max_df\":.9,\n    \"smooth_idf\":False\n}\n\n\ndef get_col(col_name): return lambda x: x[col_name]\n##I added to the max_features of the description. It did not change my score much but it may be worth investigating\nvectorizer = FeatureUnion([\n        ('description',TfidfVectorizer(\n            ngram_range=(1, 2),\n            max_features=17000,\n            **tfidf_para,\n            preprocessor=get_col('description'))),\n        ('title',CountVectorizer(\n            ngram_range=(1, 2),\n            stop_words = russian_stop,\n            #max_features=7000,\n            preprocessor=get_col('title')))\n    ])\n\nstart_vect=time.time()\n\n#Fit my vectorizer on the entire dataset instead of the training rows\n#Score improved by .0001\nvectorizer.fit(df.to_dict('records'))\n\nready_df = vectorizer.transform(df.to_dict('records'))\ntfvocab = vectorizer.get_feature_names()\n</code></pre>\n\n<p>Is what this process basically do : split all the words - give TFIDF score to all words - add each words to the columns - and state if each words are contained in each row (like one hot coding) / TFIDF score of the word? I wonder if I understood it right</p>",
      "rawMarkdown": "Hi, it's quite late for the new discussion / or question since we only have few hours left for this competition, but still, I want to ask something regarding data processing methods that lots of people use in this competition. \n\nSince I am a beginner, I referred to many of others' work, and tried to adopt some of those with my own method. Therefore, there are several methods that I used / got improvement, but I do not thoroughly understand. So, even the competition is near to it's end, I want to ask question on several method to share knowledge (And I do believe there are many beginners refer to others' kernal but does not thoroughly understand everything :))\n\n1. About image processing\nThis was the first time I've dealt with Image processing. There were many suggested ways to extract image features that load image, and extract features with cv2.\n\nHowever, I learned that we could extract features without loading the image, but rather within the zip file directly. The code I used is:\n\n    def get_blurrness(file):\n    exfile = zipped.read(file)\n    arr = np.frombuffer(exfile, np.uint8)\n    if arr.size &gt; 0:   # exclude dirs and blanks\n        imz = cv2.imdecode(arr, flags=cv2.COLOR_BGR2GRAY)\n        fm = cv2.Laplacian(imz, cv2.CV_64F).var()\n    else: \n        fm = -1\n    return fm\n\nThis somehow gave me the image blurrness score, and I tried same method with extracting numper of key points with this code:\n\n    def keyp(img):\n    try:        \n        img = image_path + str(img) + \".jpg\"\n        exfile = zipped.read(img)\n        arr = np.frombuffer(exfile, np.uint8)\n\n        imz = cv2.imdecode(arr, 1)\n        fast = cv2.FastFeatureDetector_create()\n\n    # find and draw the keypoints\n        kp = fast.detect(imz,None)\n        kp =len(kp)\n        return kp\n    except:\n        return 0\n\nHowever, I could not thoroughly understand how this process work, and if it really extract valid features. I want to hear any opinions / explanation in this process.\n\n2. About Tf - Idf processing\n\nMost of the kernals use TF - IDF processing for text features - title and description. As I understand, TF - IDF score shows statistical weight and importance of each words. However, I could not exactly understand how TF-IDF vectorizor works. [This kernal][1] is the one I referred to a lot regarding TFIDF vectorization. The major process is: \n\n    print(\"\\n[TF-IDF] Term Frequency Inverse Document Frequency Stage\")\n    russian_stop = set(stopwords.words('russian'))\n    \n    tfidf_para = {\n        \"stop_words\": russian_stop,\n        \"analyzer\": 'word',\n        \"token_pattern\": r'\\w{1,}',\n        \"sublinear_tf\": True,\n        \"dtype\": np.float32,\n        \"norm\": 'l2',\n        #\"min_df\":5,\n        #\"max_df\":.9,\n        \"smooth_idf\":False\n    }\n    \n    \n    def get_col(col_name): return lambda x: x[col_name]\n    ##I added to the max_features of the description. It did not change my score much but it may be worth investigating\n    vectorizer = FeatureUnion([\n            ('description',TfidfVectorizer(\n                ngram_range=(1, 2),\n                max_features=17000,\n                **tfidf_para,\n                preprocessor=get_col('description'))),\n            ('title',CountVectorizer(\n                ngram_range=(1, 2),\n                stop_words = russian_stop,\n                #max_features=7000,\n                preprocessor=get_col('title')))\n        ])\n        \n    start_vect=time.time()\n    \n    #Fit my vectorizer on the entire dataset instead of the training rows\n    #Score improved by .0001\n    vectorizer.fit(df.to_dict('records'))\n    \n    ready_df = vectorizer.transform(df.to_dict('records'))\n    tfvocab = vectorizer.get_feature_names()\n\nIs what this process basically do : split all the words - give TFIDF score to all words - add each words to the columns - and state if each words are contained in each row (like one hot coding) / TFIDF score of the word? I wonder if I understood it right\n\n\n  [1]: https://www.kaggle.com/him4318/avito-lightgbm-with-ridge-feature-v-2-0/code"
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "348690": "Hi, it's quite late for the new discussion / or question since we only have few hours left for this competition, but still, I want to ask something regarding data processing methods that lots of people use in this competition. \n\nSince I am a beginner, I referred to many of others' work, and tried to adopt some of those with my own method. Therefore, there are several methods that I used / got improvement, but I do not thoroughly understand. So, even the competition is near to it's end, I want to ask question on several method to share knowledge (And I do believe there are many beginners refer to others' kernal but does not thoroughly understand everything :))\n\n1. About image processing\nThis was the first time I've dealt with Image processing. There were many suggested ways to extract image features that load image, and extract features with cv2.\n\nHowever, I learned that we could extract features without loading the image, but rather within the zip file directly. The code I used is:\n\n    def get_blurrness(file):\n    exfile = zipped.read(file)\n    arr = np.frombuffer(exfile, np.uint8)\n    if arr.size &gt; 0:   # exclude dirs and blanks\n        imz = cv2.imdecode(arr, flags=cv2.COLOR_BGR2GRAY)\n        fm = cv2.Laplacian(imz, cv2.CV_64F).var()\n    else: \n        fm = -1\n    return fm\n\nThis somehow gave me the image blurrness score, and I tried same method with extracting numper of key points with this code:\n\n    def keyp(img):\n    try:        \n        img = image_path + str(img) + \".jpg\"\n        exfile = zipped.read(img)\n        arr = np.frombuffer(exfile, np.uint8)\n\n        imz = cv2.imdecode(arr, 1)\n        fast = cv2.FastFeatureDetector_create()\n\n    # find and draw the keypoints\n        kp = fast.detect(imz,None)\n        kp =len(kp)\n        return kp\n    except:\n        return 0\n\nHowever, I could not thoroughly understand how this process work, and if it really extract valid features. I want to hear any opinions / explanation in this process.\n\n2. About Tf - Idf processing\n\nMost of the kernals use TF - IDF processing for text features - title and description. As I understand, TF - IDF score shows statistical weight and importance of each words. However, I could not exactly understand how TF-IDF vectorizor works. [This kernal][1] is the one I referred to a lot regarding TFIDF vectorization. The major process is: \n\n    print(\"\\n[TF-IDF] Term Frequency Inverse Document Frequency Stage\")\n    russian_stop = set(stopwords.words('russian'))\n    \n    tfidf_para = {\n        \"stop_words\": russian_stop,\n        \"analyzer\": 'word',\n        \"token_pattern\": r'\\w{1,}',\n        \"sublinear_tf\": True,\n        \"dtype\": np.float32,\n        \"norm\": 'l2',\n        #\"min_df\":5,\n        #\"max_df\":.9,\n        \"smooth_idf\":False\n    }\n    \n    \n    def get_col(col_name): return lambda x: x[col_name]\n    ##I added to the max_features of the description. It did not change my score much but it may be worth investigating\n    vectorizer = FeatureUnion([\n            ('description',TfidfVectorizer(\n                ngram_range=(1, 2),\n                max_features=17000,\n                **tfidf_para,\n                preprocessor=get_col('description'))),\n            ('title',CountVectorizer(\n                ngram_range=(1, 2),\n                stop_words = russian_stop,\n                #max_features=7000,\n                preprocessor=get_col('title')))\n        ])\n        \n    start_vect=time.time()\n    \n    #Fit my vectorizer on the entire dataset instead of the training rows\n    #Score improved by .0001\n    vectorizer.fit(df.to_dict('records'))\n    \n    ready_df = vectorizer.transform(df.to_dict('records'))\n    tfvocab = vectorizer.get_feature_names()\n\nIs what this process basically do : split all the words - give TFIDF score to all words - add each words to the columns - and state if each words are contained in each row (like one hot coding) / TFIDF score of the word? I wonder if I understood it right\n\n\n  [1]: https://www.kaggle.com/him4318/avito-lightgbm-with-ridge-feature-v-2-0/code"
  }
}