{
  "id": 20793,
  "title": "How to use the attrsJSON column",
  "url": "/competitions/avito-duplicate-ads-detection/discussion/20793",
  "author_name": "",
  "post_date": "2016-05-07T20:28:22.530Z",
  "votes": null,
  "comment_count": 6,
  "views": 1355,
  "content": "<p>This column contains a list of dictionaries representing additional parameters of the ad in JSON format. My question is how would be possible ways of using this data as features. One possible way could be to count the number of common keys and common values between pairs of dictionaries. Any other idea?</p>",
  "messages": [
    {
      "id": "119177",
      "postDate": "05/07/2016 20:28:22",
      "content": "<p>This column contains a list of dictionaries representing additional parameters of the ad in JSON format. My question is how would be possible ways of using this data as features. One possible way could be to count the number of common keys and common values between pairs of dictionaries. Any other idea?</p>",
      "rawMarkdown": "This column contains a list of dictionaries representing additional parameters of the ad in JSON format. My question is how would be possible ways of using this data as features. One possible way could be to count the number of common keys and common values between pairs of dictionaries. Any other idea?",
      "votes": null
    },
    {
      "id": "119727",
      "postDate": "05/12/2016 14:10:00",
      "content": "<p>I used the same approach you mentioned. I counted the number of matches in the key-value pairs.</p>",
      "rawMarkdown": "I used the same approach you mentioned. I counted the number of matches in the key-value pairs.",
      "votes": null
    },
    {
      "id": "124220",
      "postDate": "06/16/2016 10:17:55",
      "content": "<p>I am doing the same, but using relative frequencies rather than absolute counts. It is 1 in perfect match and 0 if no match. What is known as <a href=\"https://en.wikipedia.org/wiki/Jaccard_index\">Jaccard similarity</a>. (In essence, I think it is equivalent to building a TF dictionary and then comparing vectors using dot-product, just much faster.)</p>\n\n<p>You can then do smarter things on top of that. Extract features when #keys &gt; x, otherwise zero it out, because small json do not convey much information. If you use look-ahead decision trees, you can add a complementary feature like number of keys. You can also only match numbers which seem the most discriminative json feature. You just have to throw as much features as you can at it... It's easier to remove them afterwards on your feature selection process.</p>",
      "rawMarkdown": "I am doing the same, but using relative frequencies rather than absolute counts. It is 1 in perfect match and 0 if no match. What is known as [Jaccard similarity][1]. (In essence, I think it is equivalent to building a TF dictionary and then comparing vectors using dot-product, just much faster.)\r\n\r\nYou can then do smarter things on top of that. Extract features when #keys > x, otherwise zero it out, because small json do not convey much information. If you use look-ahead decision trees, you can add a complementary feature like number of keys. You can also only match numbers which seem the most discriminative json feature. You just have to throw as much features as you can at it... It's easier to remove them afterwards on your feature selection process.\r\n\r\n\r\n  [1]: https://en.wikipedia.org/wiki/Jaccard_index",
      "votes": null
    },
    {
      "id": "125582",
      "postDate": "06/30/2016 15:16:04",
      "content": "<p>Can you tell me are you guys including the count of common keys/values as a feature?  I am counting the number of common keys and have that number against that key in a different column</p>",
      "rawMarkdown": "Can you tell me are you guys including the count of common keys/values as a feature?  I am counting the number of common keys and have that number against that key in a different column",
      "votes": null
    },
    {
      "id": "125584",
      "postDate": "06/30/2016 16:07:22",
      "content": "<p>Anon, you should &quot;invent&quot; as many features as you can , not only one feature. There is no &quot;better way&quot;. Every feature you remember is good. It is better to select features afterwards than before because you never know what works.</p>\n\n<p>That being said, I would try also with relative frequencies. The problem with absolute frequencies, as you did, is that the number of attributes depends on the category. Try with both and see which your model prefers.</p>\n\n<p>Personally, beating the avito benchmark was good enough for me, so my model only had about 20 features. Anyhow, some suggestions: do relative frequencies but only for numbers (those seem to be the best predictor), and also try to create features that are zero if the number of categories is small so that the only non-zero values represent a strong signal. Each duplicate belongs to the same category, so you could study avito.com and add a metric based on the number of form attributes that both did not filled (notice those json come from the forms that users manually insert).</p>",
      "rawMarkdown": "Anon, you should \"invent\" as many features as you can , not only one feature. There is no \"better way\". Every feature you remember is good. It is better to select features afterwards than before because you never know what works.\r\n\r\nThat being said, I would try also with relative frequencies. The problem with absolute frequencies, as you did, is that the number of attributes depends on the category. Try with both and see which your model prefers.\r\n\r\nPersonally, beating the avito benchmark was good enough for me, so my model only had about 20 features. Anyhow, some suggestions: do relative frequencies but only for numbers (those seem to be the best predictor), and also try to create features that are zero if the number of categories is small so that the only non-zero values represent a strong signal. Each duplicate belongs to the same category, so you could study avito.com and add a metric based on the number of form attributes that both did not filled (notice those json come from the forms that users manually insert).",
      "votes": null
    },
    {
      "id": "125585",
      "postDate": "06/30/2016 16:09:52",
      "content": "<p>Thank you Ricardo for the detailed reply..</p>",
      "rawMarkdown": "Thank you Ricardo for the detailed reply..",
      "votes": null
    },
    {
      "id": "126183",
      "postDate": "07/07/2016 00:38:19",
      "content": "<p>I tried splitting the keys and values of the Json objects and encoded them using scikit-learn's label encoder. After merging the key-value pairs with the item pairs from the train file, it turns out that almost all of the itemID pairs have the same Json key-value pairs.  Has anyone else encountered this problem?</p>",
      "rawMarkdown": "I tried splitting the keys and values of the Json objects and encoded them using scikit-learn's label encoder. After merging the key-value pairs with the item pairs from the train file, it turns out that almost all of the itemID pairs have the same Json key-value pairs.  Has anyone else encountered this problem?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 119727,
      "author_name": "axelgoblet",
      "author_url": "",
      "post_date": "05/12/2016 14:10:00",
      "content": "<p>I used the same approach you mentioned. I counted the number of matches in the key-value pairs.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 124220,
      "author_name": "rpmcruz",
      "author_url": "",
      "post_date": "06/16/2016 10:17:55",
      "content": "<p>I am doing the same, but using relative frequencies rather than absolute counts. It is 1 in perfect match and 0 if no match. What is known as <a href=\"https://en.wikipedia.org/wiki/Jaccard_index\">Jaccard similarity</a>. (In essence, I think it is equivalent to building a TF dictionary and then comparing vectors using dot-product, just much faster.)</p>\n\n<p>You can then do smarter things on top of that. Extract features when #keys &gt; x, otherwise zero it out, because small json do not convey much information. If you use look-ahead decision trees, you can add a complementary feature like number of keys. You can also only match numbers which seem the most discriminative json feature. You just have to throw as much features as you can at it... It's easier to remove them afterwards on your feature selection process.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 125582,
      "author_name": "anwarsadique",
      "author_url": "",
      "post_date": "06/30/2016 15:16:04",
      "content": "<p>Can you tell me are you guys including the count of common keys/values as a feature?  I am counting the number of common keys and have that number against that key in a different column</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 125584,
      "author_name": "rpmcruz",
      "author_url": "",
      "post_date": "06/30/2016 16:07:22",
      "content": "<p>Anon, you should &quot;invent&quot; as many features as you can , not only one feature. There is no &quot;better way&quot;. Every feature you remember is good. It is better to select features afterwards than before because you never know what works.</p>\n\n<p>That being said, I would try also with relative frequencies. The problem with absolute frequencies, as you did, is that the number of attributes depends on the category. Try with both and see which your model prefers.</p>\n\n<p>Personally, beating the avito benchmark was good enough for me, so my model only had about 20 features. Anyhow, some suggestions: do relative frequencies but only for numbers (those seem to be the best predictor), and also try to create features that are zero if the number of categories is small so that the only non-zero values represent a strong signal. Each duplicate belongs to the same category, so you could study avito.com and add a metric based on the number of form attributes that both did not filled (notice those json come from the forms that users manually insert).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 125585,
      "author_name": "anwarsadique",
      "author_url": "",
      "post_date": "06/30/2016 16:09:52",
      "content": "<p>Thank you Ricardo for the detailed reply..</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 126183,
      "author_name": "kevinkirchhoff",
      "author_url": "",
      "post_date": "07/07/2016 00:38:19",
      "content": "<p>I tried splitting the keys and values of the Json objects and encoded them using scikit-learn's label encoder. After merging the key-value pairs with the item pairs from the train file, it turns out that almost all of the itemID pairs have the same Json key-value pairs.  Has anyone else encountered this problem?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "119177": "This column contains a list of dictionaries representing additional parameters of the ad in JSON format. My question is how would be possible ways of using this data as features. One possible way could be to count the number of common keys and common values between pairs of dictionaries. Any other idea?",
    "119727": "I used the same approach you mentioned. I counted the number of matches in the key-value pairs.",
    "124220": "I am doing the same, but using relative frequencies rather than absolute counts. It is 1 in perfect match and 0 if no match. What is known as [Jaccard similarity][1]. (In essence, I think it is equivalent to building a TF dictionary and then comparing vectors using dot-product, just much faster.)\r\n\r\nYou can then do smarter things on top of that. Extract features when #keys > x, otherwise zero it out, because small json do not convey much information. If you use look-ahead decision trees, you can add a complementary feature like number of keys. You can also only match numbers which seem the most discriminative json feature. You just have to throw as much features as you can at it... It's easier to remove them afterwards on your feature selection process.\r\n\r\n\r\n  [1]: https://en.wikipedia.org/wiki/Jaccard_index",
    "125582": "Can you tell me are you guys including the count of common keys/values as a feature?  I am counting the number of common keys and have that number against that key in a different column",
    "125584": "Anon, you should \"invent\" as many features as you can , not only one feature. There is no \"better way\". Every feature you remember is good. It is better to select features afterwards than before because you never know what works.\r\n\r\nThat being said, I would try also with relative frequencies. The problem with absolute frequencies, as you did, is that the number of attributes depends on the category. Try with both and see which your model prefers.\r\n\r\nPersonally, beating the avito benchmark was good enough for me, so my model only had about 20 features. Anyhow, some suggestions: do relative frequencies but only for numbers (those seem to be the best predictor), and also try to create features that are zero if the number of categories is small so that the only non-zero values represent a strong signal. Each duplicate belongs to the same category, so you could study avito.com and add a metric based on the number of form attributes that both did not filled (notice those json come from the forms that users manually insert).",
    "125585": "Thank you Ricardo for the detailed reply..",
    "126183": "I tried splitting the keys and values of the Json objects and encoded them using scikit-learn's label encoder. After merging the key-value pairs with the item pairs from the train file, it turns out that almost all of the itemID pairs have the same Json key-value pairs.  Has anyone else encountered this problem?"
  },
  "source": "meta"
}