This work attempts to extract a viewer’s emotion from the three modalities of a movie: audio, visual and text. A major obstacle for emotion research has been the lack of appropriately annotated databases, limiting the potential of supervised algorithms. To that end we develop and present a database of movie affect, annotated in continuous time, on a continuous valence-arousal scale. Supervised learning methods are proposed to model the continuous affective response using hidden Markov Models and low-level audio-visual features and classify each video frame into one of seven discrete categories (in each dimension); the discrete-valued curves are then converted to continuous values via spline interpolation. A variety of audio-visual features are investigated and an optimal feature set is selected. The potential of the method is verified on twelve 30-minute movie clips with good precision at a macroscopic level. This method proves not suitable to process subtitle information, so we explore the creation of a textual affective model, starting with a fully automated algorithm for expanding an affective lexicon with new entries. Continuous valence ratings are estimated for unseen words under the assumption that semantic similarity implies affective similarity. Starting from a set of manually annotated words, a linear model is trained using the least mean squares algorithm. The semantic similarity between the selected features and the unseen words is computed with various similarity metrics, and used to compute the valence of unseen words. The proposed algorithm performs very well on reproducing the valence ratings of the Affective Norms for English Words (ANEW) and General Inquirer datasets. We then use three simple fusion schemes to combine lexical valence scores into sentence-level scores, producing state-of-the-art results on the sentence rating task of the SemEval 2007 corpus.