2013年6月20日 星期四

Aspect Extraction

Mukherjee and Liu (2012). Aspect Extraction through Semi-Supervised Modeling. ACL.

  1. A key task of the framework is to extract aspects of entities that have been commented in opinion documents.
  2. Two main types:
    • The first type only extracts aspect terms without grouping them;
    • The second type uses statistical topic models to extract aspects and group them.
  3. This paper that given some seeds in the user interested categories.
  4. The models are related to the DFLDA model in (Andrzejewski et al., 2009), while DF-LDA is only for topics/aspects.
  5. There are many existing works on aspect extraction
    • to find frequent noun terms and possibly with the help of dependency relations 
    • to use supervised sequence labeling
  6. Aspect and sentiment extraction using topic modeling come in two flavors:
    • discovering aspect words sentiment wise (放在一起表示)
    • separately discovering both aspects and sentiments (used Maximum-Entropy, Mei
      et al., 2007; Zhao et al., 2010)
    • 思考上述兩種方法的優缺點,改進的空間
  7. One problem with these existing models is that many discovered aspects are not understandable / meaningful to users.
  8. Standard LDA and existing aspect and sentiment models based on document level, so many “non-specific” terms being pulled and clustered
  9. Aspect terms tend to be nouns or noun phrases and sentiment terms tend to be adjectives, adverbs
Zhao et al., (2010). jointly modeling aspects and opinions with a mazEnt-LDA Hybrid. EMNLP.

  1. Separateing aspects and opinion words can be very useful.
    • can be used to construct a domain-dependent sentiment lexicon and applied to tasks such as sentiment classification. 
  2. Global topic models may not be suitable for detecing rateable aspects.
Bagheri et al., (2013). An Unsupervised Aspect Detection Model for Sentiment Analysis of Reviews. NLDB.
  1. Aspects are important because without knowing them, the opinions expressed in a sentence or a review are of limited use.

2013年6月19日 星期三

multi-aspect sentence

Many sentences in real reviews often involve two or more aspects.
The first sentence contains three single-aspect segments: an environment-segment (环境不错/ the environment is nice), a food-segment (菜品一般/ the quality of food is so so), and a charge-segment (很贵/ the food is very expensive)

2013年6月17日 星期一

Terminology

  • topic: a multinomial distribution over words that represents a coherent concept in text.
  • aspect: a multinomial distribution over words that represents a more speci c topic in reviews, for example,"lens" in camera reviews.
  • senti-aspect: a multinomial distribution over words that represents a pair of aspect and sentiment, for example, "screen, positive" in a laptop review.
  • affective word: a word that expresses a feeling, for example "satisfied", "disappointed".
  • evaluative word: a word that expresses sentiment by evaluating an aspect, for example, "excellent", "nice".
  • general evaluative word: an evaluative word that expresses a consistent sentiment every time it is used, for example, "good", "bad".
  • aspect-specific evaluative word: an evaluative word that may express di erent sentiments depending on the aspect, for example, a "small" font size on a monitor that is hard to read vs. a "small" vacuum that is portable.
  • sentiment word: a word that conveys sentiment. It is either an a ective word, general evaluative word, or aspect-speci c evaluative word.

source: Jo and Oh, WSDM'11.

Gold-standard lexicon

The gold-standard lexicon mentioned in the former case is obtained through one
of the following ways:
a) by manually tagging words from a domain corpus;
b) by one or more domain experts choosing aspects and keywords without the use of a
corpus; or
c) using review sets that have already been annotated with aspects and
keywords by the original reviewers

2011年8月9日 星期二

2011年8月4日 星期四

Relationship between topic modeling and documents clustering

I am learning to use topic modeling for documents clustering. I would like to clarify whether my understaning of the relationship between latent dirichlet allocation (LDA) and the generic task of document clustering is
correct or not?

The LDA analysis tends to output the topic proportions for each document. This is not the direct result of document clustering. However, we can treat this probability proportions as a feature reprsentation for each document. Afterwards, we can invoke other established clustering method, like K-means, to cluster documents based on the feature configurations generated by LDA analysis.

The best metric we found for computing the semantic similarity of topics was a pairwise topic coherence, using the coherence metric from "Automatic Evaluation of Topic Coherence," by Newman et al., NAACL 2010.

2011年8月3日 星期三

關於鍵盤異常

有時候鍵盤會發生特殊情形,如 T 鍵變win + T,F鍵變成win + F,等等…,這個問題常發生在使用虛擬機或遠端登入後,鍵盤的Fn鍵被開啟,這時只要再按一下 win + 任意功能鍵,就會恢復正常了

Types of Bots: An Overview

Learn more about all the different varieties of bots, and what they can do for you http://botnerds.com/types-of-bots/ In this articl...